Telemetry to claim validity
Your agent logs are precise. That is not the same as their meaning what the dashboard says they mean.
1 paper · 1 proposed gate · 1 tool · 1 term
Claim 01
A twelve-week audit of one agent deployment found that file event spans supported no validated measure of session or task duration.
The frozen corpus was 41,495 transcript files and six auxiliary telemetry streams from a single coding-agent orchestration harness. Transcript-file event spans produce a real, well-defined distribution — but a file-level distribution, not a measure of how long a logical session or a task took.
The sensitivity check is the sharp part. Of 87 files spanning more than eight hours, 46 fell below that threshold once their single largest internal gap was removed. The records did not change; the treatment of idle time did.
What this does not show
This is one orchestration-harness deployment observed over twelve weeks. It does not establish that the same gap between record and construct appears in other harnesses, other telemetry stacks, or other organizations.
Claim 02
A claims ledger with 28,055 rows contained 117 rows that were not automatic reflex events, and none of them established a completed task.
27,938 of the 28,055 rows were automatic retrieval-reflex events — the instrument recording its own activity. The residual 117 entries were heterogeneous, and reading them did not yield completed tasks or correct outcomes.
A ledger that large invites being counted. Counting it produces a number that is arithmetically correct and answers a question nobody asked.
What this does not show
This is a record-level finding about one specific claims ledger and its producer. It neither certifies nor refutes the design of any other ledger.
Claim 03
An anomaly instrument flagged roughly 65% of scored turns, and without independent behavioral labels that figure is an instrument-positive rate, not a failure rate.
The distinction is the whole finding. 65% of turns being flagged tells you what the instrument does. It tells you nothing about how often the agent actually failed, because no independent labels exist to compare against.
A number of that size, read as a failure rate, would justify redesigning the system. Read correctly, it justifies validating the instrument first.
What this does not show
The 65% figure is scoped to the audited deployment and its specific instrument. It must not be read as a prevalence estimate for agent failure anywhere else.
Claim 04
The paper proposes a gate: name the producer, unit, construct, population, window, denominator, validation, and sensitivity before a metric is allowed to drive a decision.
Each item on that list is a question that has an answer or does not. Which process emitted this? What is one row? What construct is this supposed to measure? Over whom, over what window, divided by what? What validates the mapping, and how much does the number move under a different reasonable treatment?
Nothing about it is exotic. It is the documentation a measurement would need before anyone reorganized an agent architecture around it.
What this does not show
The gate is proposed from a single study. It has not been shown to produce consistent results across analysts, and no evidence yet shows that applying it improves decisions.
Claim 05
The same evidentiary discipline can be applied to review scores, by binding each score to the evidence cited for it.
hermes-rubric records structured review scores together with the evidence each score was based on, and keeps reproducibility receipts so a score can be re-derived rather than re-argued.
It is the telemetry problem in a smaller frame: a number that looks authoritative, whose meaning depends entirely on what produced it.
What this does not show
hermes-rubric is advisory and not a binary release gate. Its published agreement result still needs independent reproduction, so its own scoring consistency is not yet externally established.
Every artifact on this path
| Artifact | Kind | What it establishes | Status |
|---|---|---|---|
| Precise Records, Unstable Meanings | Paper | Where record-level measures stop supporting construct-level claims | preprint |
| Telemetry-to-Claim Gate | Term | Eight items to document before telemetry drives a decision | proposed |
| hermes-rubric | Tool | Review scores bound to cited evidence, with receipts | PyPI 1.0.0 |
If this is the failure you are seeing, tell us the system and the symptom. Book a 30-minute call, or send a note and get a written read instead.