Skip to content

Inside the work / Microsoft Semantic Kernel

The conversation kept going.
The system prompt didn’t.

History truncation could delete the system or developer message. We reproduced the deletion, ported Microsoft’s documented .NET behavior, and added regression tests.

Before / the defect

Less history.
Missing instruction.

System / developer message removedRecent conversation retained

The reducer could remove the instruction that was meant to govern the conversation.

After / the merged fix

Less history.
Instruction preserved.

First system / developer message preservedRecent conversation retained

The merged fix preserves the first system or developer message during truncation.

Public upstream workPR #13610 · Maintainer-merged, 19 March 2026

Read the fix and regression evidence

Another layer / LangChain

Making LangChain’s agent tool binding compatible with Claude’s extended thinking

A 20-line guard handles forced tool choice when extended thinking is enabled, with two test functions covering eight scenarios. Community reporter’s root-cause credit is preserved.

See PR #35544 and the tradeoff →

These are public examples of our method. Paid work starts with your system, your symptom, and an agreed scope.

Put the lab on your problem

Paid expert work

Hire the lab for
the part you need.

A written diagnosis when the cause is unclear. Engineering when the failing layer is known. Scope and terms are agreed before work starts.

01 / DiagnosisPaid

Find the failing layer.

Bring a symptom, a trace, or a reproduction. We inspect the prompts, tools, memory, and configuration on the path that fails.

You get

A written diagnosis: reproduction status, supported findings, the recommended fix and what it would touch, or the next evidence needed.

Discuss a diagnosis (Calendly, opens in a new tab)
02 / EngineeringPaid

Patch the right layer.

When implementation is useful, we change the relevant layer in your Python, JavaScript, or TypeScript stack.

You get

A scoped change with regression evidence. These public fixes show how a specific finding becomes an inspectable engineering result.

Discuss engineering (Calendly, opens in a new tab)

Agent reliability engineering · Memory and context integrity · Evaluation and auditability · Runtime controls

The tools and research are free. Use the 9-project open-source catalog, papers, and public receipts without hiring us.

Get the tools →Read the research →

Find your starting point

Choose the symptom closest to yours. Inspect the public work, then discuss a diagnosis, or use a free tool where one fits.

01

A published record contradicts its own inputs

Paid diagnosis

ProofCISA CVE-2026-14216 receipt

NextDiscuss your system (Calendly, opens in a new tab)

02

The model stops following its system prompt in long conversations

Paid diagnosis

ProofSemantic Kernel #13610 receipt

NextTell us what's breaking

03

Enabling a model feature breaks your agent's tool calls

Paid diagnosis

ProofLangChain #35544 receipt

NextDiscuss your system (Calendly, opens in a new tab)

04

An evaluation fails with the wrong error, or reports a result you cannot trust

Paid diagnosis

ProofDSPy #9978 receipt

NextDiscuss your system (Calendly, opens in a new tab)

05

Generated tool schemas come out malformed

Paid diagnosis

ProofSemantic Kernel #13635 receipt

NextDiscuss your system (Calendly, opens in a new tab)

06

Retrieval scores look plausible but rank the wrong memories

Paid diagnosis

Proofmem0 receipt

NextDiscuss your system (Calendly, opens in a new tab)

07

A closed-source AI tool fails and you cannot see inside it

Paid diagnosis

ProofClaude Code #55241 receipt

NextDiscuss your system (Calendly, opens in a new tab)

08

Instruction and tool-definition defects you want caught before merge

Free · self-serve

ProofLintLang

NextRun LintLang

Commanduvx --from lintlang==0.8.2 lintlang scan --discover .

09

Untrusted input reaching an agent before it acts

Free · self-serve

ProofLittle Canary

NextTry Little Canary (experimental)

What to bring and what comes back

Scope is set on the call. The handoff does not change: what you send us, what we send back, and what we say when the symptom will not reproduce.

Bring

One system, one symptom, and the artifacts around it.

Pick a single system and the path through it that misbehaves. Tell us the behavior you expected and the behavior you observed. Send a reproduction if you have one, or the nearest failing trace, log, or session record if you do not. Include the prompts, tool descriptions, memory or retrieval configuration, and settings that sit on that path.

Please don’t include credentials, personal data, or unredacted production secrets; describe the symptom and share redacted evidence first.

Receive

A written diagnosis your engineers can act on.

The write-up states whether we reproduced the symptom and what the evidence supports. When the evidence supports a finding, it identifies which layer is failing, the fix we recommend, and what implementing that fix would touch. Otherwise, it names the next evidence required. It is written to be read by an engineer who was not on the call.

If not

If we cannot reproduce it, we say so.

Some symptoms do not reproduce from what is available, and some evidence is too thin to place a cause. We say that plainly, state what the evidence does and does not support, and name the next trace, log, or configuration we would need. We do not turn an unreproduced symptom into a confident finding.

Explore all case studies →Inspect merged upstream contributions →

Tell us the system and the symptom. We will tell you what the symptom points to, or what we would need to see before naming a cause, and whether we are the right people to fix it.

Data and partnership inquiries. Bring the data you are proposing, the use case it would support, and the partnership shape you have in mind, on the 20-minute call or by email at roli@hermes-labs.ai. Hermes will say directly whether it fits and what it would need to see next.

Tell us what's breaking →Discuss your system →