Skip to content

Inside the work / Microsoft Semantic Kernel

The conversation kept going.
The system prompt didn’t.

History truncation could delete the system or developer message. We reproduced the deletion, ported Microsoft’s documented .NET behavior, and added regression tests.

Before / the defect

Less history.
Missing instruction.

System / developer message removedRecent conversation retained

The reducer could remove the instruction that was meant to govern the conversation.

After / the merged fix

Less history.
Instruction preserved.

First system / developer message preservedRecent conversation retained

The merged fix preserves the first system or developer message during truncation.

Public upstream workPR #13610 · Maintainer-merged, 19 March 2026

Read the fix and regression evidence

Another layer / LangChain

Making LangChain’s agent tool binding compatible with Claude’s extended thinking

A 20-line guard handles forced tool choice when extended thinking is enabled, with two test functions covering eight scenarios. Community reporter’s root-cause credit is preserved.

See PR #35544 and the tradeoff →

These are public examples of our method. Paid work starts with your system, your symptom, and an agreed scope.

Put the lab on your problem

Paid expert work

Hire the lab for
the part you need.

A written diagnosis when the cause is unclear. Engineering when the failing layer is known. Scope and terms are agreed on the free call.

01 / DiagnosisPaid

Find the failing layer.

Bring a symptom, a trace, or a reproduction. We inspect the prompts, tools, memory, and configuration on the path that fails.

You get

A written diagnosis: reproduction status, supported findings, the recommended fix and what it would touch—or the next evidence needed.

Explore diagnosis
02 / EngineeringPaid

Patch the right layer.

When implementation is useful, we change the relevant layer in your Python, JavaScript, or TypeScript stack.

You get

A scoped change with regression evidence. These public fixes show how a specific finding becomes an inspectable engineering result.

Explore engineering

The tools and research are free. Use the 10-project open-source catalog, papers, and public receipts without hiring us.

Get the tools →Read the research →

Hermes Labs is a reliability lab for production AI and agent systems: it diagnoses failures, engineers fixes, and publishes the tools, papers and receipts behind that work. The company is building the reliability layer for autonomous systems. Founded by Rolando Bosch (Roli), it builds and maintains LintLang, static analysis for AI agent instructions, and Little Canary, prompt-injection sensing for untrusted input.

01

A guardrail disappears mid-conversation

The instruction is still configured, but a truncation or summary step silently dropped it from the effective context sent to the model several turns ago. The agent stops following the rule, and nothing logs that it lapsed.

Fix · mergedWe named the model-behavior failure in our taxonomy, then fixed a framework defect with the same operational failure shape in Microsoft Semantic Kernel. How we found it.

02

The same mistake returns

Corrections that do not persist

A user corrects the system, but the correction disappears into a long chat log. The next response repeats the same drift as if the repair never happened.

ToolOur open-source hermeneutic mines correction triples from chat logs and gates the next response before the same drift ships twice. Explore the research.

03

Compression changes the record

Your summary or memory layer compresses context to save tokens and quietly changes what it meant: a dropped qualifier, a paraphrase that is not what you wrote.

ToolOur open-source Fidelis Memory returns your context verbatim instead of paraphrasing it. Explore the research.

04

The run leaves no usable audit trail

Unprovable agent actions

Something went wrong in an autonomous run, and you cannot show what the agent actually did, in what order, or why.

ScopeStart by naming the actions that need evidence and the records your system actually retains. Discuss the audit gap.

Diagnosis and engineering on your system are paid engagements; the tools, papers and case-study receipts are free. Match a production problem to its receipt →

Inspect the research and engineering record
1,461
controlled experiments, conducted primarily on GPT-4o, behind the taxonomy of epistemic failure modes. The taxonomy names the failure classes the audit methodology looks for.
in the taxonomy paper
Zenodo 19042468
4 fixes
merged AI/framework fixes: forced tool choice under Anthropic thinking in LangChain, system-prompt preservation and duplicate-null schema handling in Semantic Kernel, and empty-devset validation in DSPy. 96 published external pull requests merged in 2026, including 49 code, workflow, typing or maintenance contributions across open-source infrastructure.
6 papers
on Zenodo with permanent DOIs: three preprints, one working paper, and two technical notes. Their distinct evidence roles include a seven-mode corpus taxonomy, a matched-vignette experiment, a twelve-week telemetry audit, a conceptual paper on model self-report, a bounded static analysis of tool descriptions, and a prompt-injection sensing architecture.
10-project catalog
a curated public path from project setup and preflight checks to runtime observation, evaluation, and local memory. Each tool names its present evidence boundary.
Case study
a falsification-first diagnosis of a closed-source Claude Code resume hang, with measured cause classes and the evidence boundary stated explicitly.

For enterprise AI teams

Tell us what’s breaking.

Book a free 20-minute call and tell us the system and the symptom, or send a note below. Either way you get a specific read on what the symptom points to, or what we would need to see before naming a cause, and whether we are the right people to fix it, not a pitch.

Book a 20-min call (Calendly, opens in a new tab) →

Or send a note · describe the symptom

Please don’t include credentials, personal data, or unredacted production secrets; describe the symptom and share redacted evidence first.

How we handle your note: Privacy.