Research / Preprint
A Taxonomy of Epistemic Failure Modes in Large Language Models
Rolando Bosch Rodriguez · Hermes Labs
Abstract
Large language models (LLMs) exhibit systematic distortions in how they evaluate, represent, and communicate epistemic states. These distortions are distinct from hallucination—they concern cases where the model’s expressed confidence, standards of scrutiny, or assignment of causality diverge from what the evidence warrants, even when the factual content is roughly correct. Through bottom-up analysis of 1,461 controlled experiments conducted primarily on GPT-4o, we identify seven structural failure modes: (1) Null-Result Asymmetry, (2) Source-Status Credibility Bias, (3) Agency Dissolution, (4) Performative Hedging, (5) Constraint Evasion, (6) Silent Instruction Relaxation, and (7) Controversy-Truth Conflation. Each mode is defined mechanistically, illustrated with experimental evidence, and assessed for downstream consequences. Across the seven modes, a common pattern emerges: models track surface-level signals—prestige markers, hedging vocabulary, controversy language, banned word lists—rather than the semantic content those signals are supposed to index.
Keywords: AI · Artificial intelligence · Taxonomy · Silent AI failure modes · AI failure modes · Epistemic Failure Modes in LLMs · Epistemic engineering · AI assurance · model behavior · AI behavioral testing
How to cite
@misc{boschrodriguez2026taxonomy,
author = {Bosch Rodriguez, Rolando},
title = {A Taxonomy of Epistemic Failure Modes in Large Language
Models},
year = {2026},
publisher = {Zenodo},
doi = {10.5281/zenodo.19042469},
url = {https://doi.org/10.5281/zenodo.19042469},
note = {Preprint}
}APABosch Rodriguez, R. (2026). A Taxonomy of Epistemic Failure Modes in Large Language Models. Zenodo. https://doi.org/10.5281/zenodo.19042469