Agent reliability
The failures that matter in production are the ones that leave no error. Two of them are now fixed in frameworks you already run.
Four arguments, assembled from work that already exists. Each path makes numbered claims, and every claim links to the paper, tool, merged upstream fix, case study, or defined term it rests on — then states, in its own words, what it does not show.
4 paths18 claimsEvery claim states its limit
The failures that matter in production are the ones that leave no error. Two of them are now fixed in frameworks you already run.
Compression is where meaning gets lost, and almost nothing in the stack records that it happened.
Your agent logs are precise. That is not the same as their meaning what the dashboard says they mean.
Language is not only what an agent reasons about. It is part of what the agent reasons through — and that has engineering consequences.
The papers themselves are on research. The full evidence ledger, including the tools’ present evidence boundaries and the patent filings, is on proof.