Agent reliability
The failures that matter in production are the ones that leave no error. Two of them are now fixed in frameworks you already run.
Four arguments, assembled from work that already exists. Each path makes numbered claims, and every claim links to the paper, tool, merged upstream fix, case study, or defined term it rests on — then states, in its own words, what it does not show.
4 paths18 claimsEvery claim states its limit
The failures that matter in production are the ones that leave no error. Two of them are now fixed in frameworks you already run.
Compression is where meaning gets lost, and almost nothing in the stack records that it happened.
Your agent logs are precise. That is not the same as their meaning what the dashboard says they mean.
Language is not only what an agent reasons about. It is part of what the agent reasons through — and that has engineering consequences.
Long-form reference pages on the problems this work addresses, written in the field’s own language for readers arriving fresh. Living pages: each carries its revision date.
What actually determines an agent's behavior at runtime: assembled state, not weights — and why context integrity is the property that state has to keep.
How a system can fail while returning a convincing result — the patterns (null-result omission, false completion, instruction relaxation) and the controls that make absence visible.
Which prompt-checking tool does what — static linters, evaluation suites, validators, runtime guards — and how to choose a stack starting from the failure, not the tool.
The papers themselves are on research. The full evidence ledger, including the tools’ present evidence boundaries and the patent filings, is on proof.