Less history.
Missing instruction.
The reducer could remove the instruction that was meant to govern the conversation.
Work with the lab
Not every agent failure is an AI failure. Sometimes it’s a system failure. Prompts, history reducers, tool bindings and evaluation code can change what an agent receives or does. We investigate the failing path, from model behavior to the systems around it.
We reproduce failures, identify the layer the evidence supports, and engineer fixes with regression evidence. Our public work includes maintainer-merged fixes in Microsoft Semantic Kernel, LangChain and DSPy, a vulnerability-record correction acknowledged by CISA, and open-source tools and research you can inspect.
Paid diagnosis & engineering · Scope agreed before work starts
Public upstream work · Merged 2026-03-19
This exhibit covers the truncation fix. It does not establish deployment or operational use.
Read the case and limits Microsoft Semantic Kernel PR #13610Public upstream work · Merged 2026-03-05
Root cause credited to the community reporter. Tool calls are no longer guaranteed on those requests; auto and thinking-off bindings are unchanged.
Read the case and limits LangChain PR #35544Public upstream work · Merged 2026-08-26
No score is invented for an evaluation that did not run. Merge evidence does not identify an installed DSPy version.
Read the case and limits DSPy PR #9978Public vulnerability record · Corrected
An autonomous audit found the mismatch. This is one record correction, not a new vulnerability or an error rate.
Read the case and limits CISA acknowledgement on issue #333One AI fix / Microsoft Semantic KernelSystem prompt deleted First system message preserved.Public upstream work · Maintainer-merged PR #13610
Inspect the before & afterWe found a scoring error in a public vulnerability record. CISA acknowledged it and corrected the record.
Error. The CISA enrichment for CVE-2026-14216 stored a CVSS score of 5.3 for a CVSS 3.1 vector that calculates to 6.5.
Fix. CISA republished the score as 6.5; the vector and the Medium severity did not change.
How. A Hermes Labs autonomous system audited 5,324 recent CISA ADP records, including 1,060 CVSS 3.x scores checked against their vectors, reproduced the one mismatch and prepared the report.
CISA's reply on issue #333 · CVE correction commit
Related engineering: Making LangChain’s agent tool binding compatible with Claude’s extended thinking · PR #35544 · merged 2026-03-05
Inside the work / Microsoft Semantic Kernel
History truncation could delete the system or developer message. We reproduced the deletion, ported Microsoft’s documented .NET behavior, and added regression tests.
The reducer could remove the instruction that was meant to govern the conversation.
The merged fix preserves the first system or developer message during truncation.
Public upstream workPR #13610 · Maintainer-merged, 19 March 2026
Read the fix and regression evidenceAnother layer / LangChain
A 20-line guard handles forced tool choice when extended thinking is enabled, with two test functions covering eight scenarios. Community reporter’s root-cause credit is preserved.
These are public examples of our method. Paid work starts with your system, your symptom, and an agreed scope.
Put the lab on your problemPaid expert work
A written diagnosis when the cause is unclear. Engineering when the failing layer is known. Scope and terms are agreed before work starts.
Bring a symptom, a trace, or a reproduction. We inspect the prompts, tools, memory, and configuration on the path that fails.
A written diagnosis: reproduction status, supported findings, the recommended fix and what it would touch, or the next evidence needed.
When implementation is useful, we change the relevant layer in your Python, JavaScript, or TypeScript stack.
A scoped change with regression evidence. These public fixes show how a specific finding becomes an inspectable engineering result.
Agent reliability engineering · Memory and context integrity · Evaluation and auditability · Runtime controls
The tools and research are free. Use the 9-project open-source catalog, papers, and public receipts without hiring us.
Find your starting point
Choose the symptom closest to yours. Inspect the public work, then discuss a diagnosis, or use a free tool where one fits.
Scope is set on the call. The handoff does not change: what you send us, what we send back, and what we say when the symptom will not reproduce.
Pick a single system and the path through it that misbehaves. Tell us the behavior you expected and the behavior you observed. Send a reproduction if you have one, or the nearest failing trace, log, or session record if you do not. Include the prompts, tool descriptions, memory or retrieval configuration, and settings that sit on that path.
Please don’t include credentials, personal data, or unredacted production secrets; describe the symptom and share redacted evidence first.
The write-up states whether we reproduced the symptom and what the evidence supports. When the evidence supports a finding, it identifies which layer is failing, the fix we recommend, and what implementing that fix would touch. Otherwise, it names the next evidence required. It is written to be read by an engineer who was not on the call.
Some symptoms do not reproduce from what is available, and some evidence is too thin to place a cause. We say that plainly, state what the evidence does and does not support, and name the next trace, log, or configuration we would need. We do not turn an unreproduced symptom into a confident finding.
Tell us the system and the symptom. We will tell you what the symptom points to, or what we would need to see before naming a cause, and whether we are the right people to fix it.
Data and partnership inquiries. Bring the data you are proposing, the use case it would support, and the partnership shape you have in mind, on the 20-minute call or by email at roli@hermes-labs.ai. Hermes will say directly whether it fits and what it would need to see next.