acceptodds
Under review as a conference paper at ICLR 2027

Inference-Time Guardrails: Reducing Hallucinations via Structured Invariants

Abstract

Large language models (LLMs) and vision-language models (VLMs) often produce outputs that are internally inconsistent: a high-level hypothesis conflicts with finer-grained fields that should determine it under explicit domain policy. Such failures persist under chain-of-thought prompting and under tool- or solver-mediated scaffolding when the model must itself nondeterministically generate fragile code or specifications at inference time. We propose *Inference-Time Guardrails* (ITG), a post-inference scaffolding framework that requires no weight updates. Given that LLM/VLM performs a single structured generation that jointly predicts the answer to a high-level hypothesis and answers to a fixed set of auxiliary probes, lightweight, human-authored predicates then check consistency, and a deterministic correction repairs violations when possible. This decouples probabilistic extraction from auditable policy enforcement. We implement ITG across a diverse set of domains. ITG is compared with structured baselines without guardrails, hypothesis-only structured outputs, unconstrained chain-of-thought, and with related work (when applicable) such as program-style (PAL) and solver-style (SatLM) scaffolding. The evaluation is performed across multiple generalist and specialist LLMs and VLMs. ITG often yields large accuracy gains over these baselines. Moreover, for smaller models, ITG sometimes closes the performance gap with larger models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.