Trust the Monitor, Not the Model: Epistemic Capability Control for LLM Agents
Abstract
As LLM agents turn uncertain claims into memory, tool use, and external actions, hypotheses can silently acquire operational authority while uncertainty labels re- main entrusted to the same fallible models. We propose Epistemic Capability Con- trol (ECC), a structured runtime-authorization framework that separates semantic formalisation from deterministic verification. Natural-language policies and raw workflow context must first be mapped onto a common structured representation. Given these structured representations, ECC explicitly represents provenance, sensitivity, capability, epistemic status, evidence, support relations, and promo- tion history. A deterministic reference monitor mediates consequential actions, prevents ordinary model transformations from increasing authority, and permits authority increases only through admissible trusted promotions such as indepen- dent verification or human approval. ECC’s guarantees are therefore conditional on the structured policy and runtime state supplied to the monitor, rather than on the correctness of semantic interpretation itself. Our theory establishes non- expansion and structural-authorization invariants, shows that epistemic soundness is impossible from untrusted labels alone, and bounds residual over-authorization through requirement-relative threshold-crossing events. ECC thereby identifies a central reliability boundary: semantic formalisation remains learning-dependent, whereas formal authorization can be deterministic, reproducible, and auditable once the relevant state has been structured. On 480 held-out agent traces, ECC- Ref achieves 85.8% authorization accuracy with no decision flips under unsup- ported authority perturbations, while the learned formalisation interface exhibits a 26.25% authorization-relevant error rate.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.