ScaffoldShift: Evaluating the Impact of Agent Scaffolds on Hallucination
Abstract
Agent scaffolds extend language models with tools and execution workflows, allowing them to act on external information and respond to execution feedback. These capabilities raise a fundamental question about reliability: how does operating within a scaffold change the hallucination behavior of a language model? Existing evaluations focus on hallucination in agents as complete systems, without assessing the underlying language model’s hallucination behavior on the same tasks as a baseline. However, evaluating an agent alone does not establish whether its scaffold amplifies or suppresses hallucination relative to the underlying model. What remains insufficiently understood is the change associated with the scaffold itself. To address this gap, we make two contributions. First, we introduce the study of scaffold-associated hallucination shift: whether hallucination increases, decreases, or shifts across task types when the same language model operates within an agent scaffold. Studying this shift requires paired evaluations on corresponding tasks with shared assessment criteria. Second, we introduce ScaffoldShift, a benchmark containing 2,360 cases across six task families that operationalizes this study. Each case pairs agent-controlled execution with base-model replay of a corresponding reference trajectory. Both views use shared ground truth and automatic scoring of structured outputs and execution evidence, without human or model-based judges. Together, the proposed study and benchmark enable paired measurement of whether scaffolds amplify, suppress, or redistribute hallucination relative to the underlying model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.