Agent reliability
Agent failures can appear as rejected requests, lost instructions, invalid schemas, or misleading evaluation errors. Each has a verified upstream fix.
Four arguments, assembled from work that already exists. Each path makes numbered claims, and every claim links to the paper, tool, merged upstream fix, case study, or defined term it rests on — then states, in its own words, what it does not show.
4 paths18 claimsEvery claim states its limit
Agent failures can appear as rejected requests, lost instructions, invalid schemas, or misleading evaluation errors. Each has a verified upstream fix.
Compression is where meaning gets lost, and almost nothing in the stack records that it happened.
Your agent logs are precise. That is not the same as their meaning what the dashboard says they mean.
Language is not only what an agent reasons about. It is part of what the agent reasons through — and that has engineering consequences.
Living technical guides that answer practitioner problems in the field’s own language and connect the answer to evidence. Each guide carries its revision date.
What actually determines an agent's behavior at runtime: assembled state, not weights — and why context integrity is the property that state has to keep.
How a system can fail while returning a convincing result — the patterns (null-result omission, false completion, instruction relaxation) and the controls that make absence visible.
How an AI system can keep responding fluently while a retrieved document, compressed summary, or lost correction changes the question it is actually answering.
Which prompt-checking tool does what — static linters, evaluation suites, validators, runtime guards — and how to choose a stack starting from the failure, not the tool.
The papers themselves are on research. The full evidence ledger, including the tools’ present evidence boundaries and the patent filings, is on proof.