Less history.
Missing instruction.
The reducer could remove the instruction that was meant to govern the conversation.
Find the failing layer
Not every agent failure is an AI failure. Sometimes it’s a system failure. Prompts, history reducers, tool bindings and evaluation code can change what an agent receives or does. We investigate the failing path, from model behavior to the systems around it.
We reproduce failures, identify the layer the evidence supports, and engineer fixes with regression evidence. Our public work includes maintainer-merged fixes in Microsoft Semantic Kernel, LangChain and DSPy, a vulnerability-record correction acknowledged by CISA, and open-source tools and research you can inspect.
Paid diagnosis & engineering · Scope agreed before work starts
Public upstream work · Merged 2026-03-19
This exhibit covers the truncation fix. It does not establish deployment or operational use.
Read the case and limits Microsoft Semantic Kernel PR #13610Public upstream work · Merged 2026-03-05
Root cause credited to the community reporter. Tool calls are no longer guaranteed on those requests; auto and thinking-off bindings are unchanged.
Read the case and limits LangChain PR #35544Public upstream work · Merged 2026-08-26
No score is invented for an evaluation that did not run. Merge evidence does not identify an installed DSPy version.
Read the case and limits DSPy PR #9978Public vulnerability record · Corrected
An autonomous audit found the mismatch. This is one record correction, not a new vulnerability or an error rate.
Read the case and limits CISA acknowledgement on issue #333One AI fix / Microsoft Semantic KernelSystem prompt deleted First system message preserved.Public upstream work · Maintainer-merged PR #13610
Inspect the before & afterWe found a scoring error in a public vulnerability record. CISA acknowledged it and corrected the record.
Error. The CISA enrichment for CVE-2026-14216 stored a CVSS score of 5.3 for a CVSS 3.1 vector that calculates to 6.5.
Fix. CISA republished the score as 6.5; the vector and the Medium severity did not change.
How. A Hermes Labs autonomous system audited 5,324 recent CISA ADP records, including 1,060 CVSS 3.x scores checked against their vectors, reproduced the one mismatch and prepared the report.
CISA's reply on issue #333 · CVE correction commit
Related engineering: Making LangChain’s agent tool binding compatible with Claude’s extended thinking · PR #35544 · merged 2026-03-05
Inside the work / Microsoft Semantic Kernel
History truncation could delete the system or developer message. We reproduced the deletion, ported Microsoft’s documented .NET behavior, and added regression tests.
The reducer could remove the instruction that was meant to govern the conversation.
The merged fix preserves the first system or developer message during truncation.
Public upstream workPR #13610 · Maintainer-merged, 19 March 2026
Read the fix and regression evidenceAnother layer / LangChain
A 20-line guard handles forced tool choice when extended thinking is enabled, with two test functions covering eight scenarios. Community reporter’s root-cause credit is preserved.
These are public examples of our method. Paid work starts with your system, your symptom, and an agreed scope.
Put the lab on your problemPaid expert work
A written diagnosis when the cause is unclear. Engineering when the failing layer is known. Scope and terms are agreed before work starts.
Bring a symptom, a trace, or a reproduction. We inspect the prompts, tools, memory, and configuration on the path that fails.
A written diagnosis: reproduction status, supported findings, the recommended fix and what it would touch, or the next evidence needed.
When implementation is useful, we change the relevant layer in your Python, JavaScript, or TypeScript stack.
A scoped change with regression evidence. These public fixes show how a specific finding becomes an inspectable engineering result.
The tools and research are free. Use the 9-project open-source catalog, papers, and public receipts without hiring us.
Hermes Labs does AI reliability engineering. We find where AI agents and the systems they run on fail before production. Founded by Rolando Bosch (Roli), it builds and maintains LintLang, static analysis for AI agent instructions, and Little Canary, prompt-injection sensing for untrusted input.
The instruction is still configured, but a truncation or summary step silently dropped it from the effective context sent to the model several turns ago. The agent stops following the rule, and nothing logs that it lapsed.
Fix · mergedWe named the model-behavior failure in our taxonomy, then fixed a framework defect with the same operational failure shape in Microsoft Semantic Kernel. How we found it.
A user corrects the system, but the correction disappears into a long chat log. The next response repeats the same drift as if the repair never happened.
ToolOur open-source hermeneutic mines correction episodes from chat logs, retrieves relevant prior guidance for similar tasks, and checks outgoing drafts against fixed English risk patterns. Explore the research.
Your summary or memory layer compresses context to save tokens and quietly changes what it meant: a dropped qualifier, a paraphrase that is not what you wrote.
ToolOur open-source Fidelis Memory returns your context verbatim instead of paraphrasing it. Explore the research.
Something went wrong in an autonomous run, and you cannot show what the agent actually did, in what order, or why.
ScopeStart by naming the actions that need evidence and the records your system actually retains. Discuss the audit gap.
Diagnosis and engineering on your system are paid engagements; the tools, papers and case-study receipts are free. Match a production problem to its receipt →
For enterprise AI teams
Tell us the system and the symptom. Send a note below or discuss it with us. We will say what the symptom points to, or what we would need to see before naming a cause. Scope and terms are agreed before paid diagnosis or engineering starts.
Discuss your system (Calendly, opens in a new tab) →