acceptodds
Under review as a conference paper at ICLR 2027

Do More Loops Follow the Evidence?

Abstract

More inference-time computation can improve looped language models, but the same performance gain can reflect different roles for added computation. Does another recurrent step make the model begin using facts in the prompt, or merely widen the score gap around a choice already based on them? To answer this question, we introduce a framework that redirects the relational path supporting one candidate toward the other while keeping the question, entities, relation names, and answer pair fixed. We track which candidate wins and the score gap between candidates at successive outputs. We evaluate the framework across several recurrent systems using controlled fictional graphs and benchmark-derived multi-hop questions. Controlled fictional graphs reveal three observed profiles: acquisition, where choice control emerges at later outputs; sharpening, where control is already present and remains stable while the margin grows; and margin-control divergence, where the margin grows even as choice control declines. Repair and structural controls confirm that these effects follow the path that determines the answer rather than unrelated prompt changes. Yet control over a fixed answer pair does not guarantee reliable free-form generation. Our results show that performance gains alone cannot explain what additional recurrent computation contributes. Evaluating choice and score together reveals how evidence control changes across outputs rather than treating a larger score gap as stronger control.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.