Skip to content

Research / Preprint

Measurement Validity / AI Agent TelemetryPublished 30 July 2026

Precise Records, Unstable Meanings: Measurement Validity and Unsupported Claims Derived from AI Agent Telemetry

Rolando Bosch · Hermes Labs

Read the PDF DOI 10.5281/zenodo.21652317
Preprint · 2026

Agent systems are increasingly evaluated and modified through operational telemetry. Yet precise records do not necessarily identify the sessions, tasks, outcomes, or behaviors named by later metrics. When those interpretations guide routing, memory, autonomy, evaluation, or redesign, measurement validity becomes an architectural concern.

We report a naturalistic measurement-validity audit of one coding-agent orchestration-harness deployment observed over twelve weeks. The frozen corpus comprised 41,495 transcript files and six auxiliary telemetry streams. Each candidate claim was examined for producer provenance, analytical unit, construct, population, observation window, denominator, validation evidence, and sensitivity to alternative treatments of inter-event gaps.

Across the audited mappings, valid record-level quantities did not support several broader construct-level interpretations. Transcript-file event spans produced a file-level distribution but no validated measure of logical-session or task duration; 46 of 87 files spanning more than eight hours fell below that threshold after removal of their largest internal gap. Of 28,055 claims-ledger rows, 27,938 were automatic retrieval-reflex events; the remaining 117 heterogeneous entries did not establish completed tasks or correct outcomes. An anomaly instrument flagged approximately 65% of scored turns, but without independent behavioral labels this was an instrument-positive rate, not an estimate of failure prevalence.

The study provides an empirical account of how record-level measures change at the claim layer and proposes a Telemetry-to-Claim Gate for documenting producer, unit, construct, population, window, denominator, validation, and sensitivity before telemetry is used in evaluation, governance, or system redesign.

Companion paper: The Generative Horizon: Applied Hermeneutics, Linguistic Attractors, and the Limits of Model Self-Report. The papers address distinct questions; neither validates the other.

AuthorRolando BoschAffiliationHermes LabsPublishedDOI10.5281/zenodo.21652317Full textPDFLicenseCC BY 4.0

Keywords: AI agent telemetry · measurement validity · construct validity · agent evaluation · operational telemetry · telemetry provenance · agent observability · coding agents · Telemetry-to-Claim Gate · agent traces · trajectory-level evaluation · evaluation validity