Research / Working paper
The Asymmetric Burden of Proof: LLMs Show a Null-Result Asymmetry in a Matched-Vignette Benchmark
Rolando Bosch Rodriguez · Hermes Labs
Abstract
We report matched-pair experiments testing whether large language models apply symmetric evidential standards to positive and null scientific claims. Three models (GPT-4o, GPT-5.2 Thinking, Claude Haiku 4.5) evaluated fictional scientific vignettes in which evidence quality was held constant while only the conclusion direction was reversed. Across all six model-format conditions, models allocated less conclusion-consistent probability mass to null claims than to matched positive claims (gaps of 19.6–56.7 percentage points; all bootstrap 95% CIs excluding zero). The asymmetry was directionally consistent in 23 of 24 pair-condition cells and persisted even when discrete classification labels collapsed entirely, surfacing through probability allocation rather than categorical commitment. We characterize this as an asymmetric burden of proof: models treat non-detection as more provisional than matched detection claims, with implications for evidence synthesis, safety assessment, and decision-support pipelines that rely on LLM-generated confidence scores.
Keywords: large language models · LLM evaluation · sycophancy · Artificial Intelligence · Null-result bias · Evidence synthesis · LLM failure modes · Silent failure modes · LLM behavior · Production LLMs · AI safety
How to cite
@misc{boschrodriguez2026asymmetric,
author = {Bosch Rodriguez, Rolando},
title = {The Asymmetric Burden of Proof: LLMs Show a Null-Result
Asymmetry in a Matched-Vignette Benchmark},
year = {2026},
publisher = {Zenodo},
doi = {10.5281/zenodo.18867694},
url = {https://doi.org/10.5281/zenodo.18867694},
note = {Working paper}
}APABosch Rodriguez, R. (2026). The Asymmetric Burden of Proof: LLMs Show a Null-Result Asymmetry in a Matched-Vignette Benchmark. Zenodo. https://doi.org/10.5281/zenodo.18867694