acceptodds
Under review as a conference paper at ICLR 2027

Action, Not Belief: What Makes Hypothesis Prompting Work in Debugger Agents

Abstract

Recent work has introduced hypothesis-driven debugging into LLM-based program repair, prompting agents to formulate hypotheses, predict observations, and test them with debugging tools. However, these methods combine hypothesis generation with guidance for subsequent actions, making it unclear which component drives the repair gain. We disentangle these components through controlled comparisons under a fixed debugger interface. To examine whether prediction–observation agreement tracks repair performance, we introduce HypoCal, a measure of agreement between stated runtime predictions and subsequent debugger observations. Our experiments show that injecting hypotheses into the next action prompt improves repair, whereas eliciting hypotheses alone does not. The gain disappears when the injected hypothesis concerns another bug or misidentifies the fault location. HypoCal remains low across prompting conditions and does not track repair gains. An analysis of tool use shows that hypothesis injection shifts tool use from source inspection toward runtime probing. Across four models, repair effects have the same rank ordering as baseline structured-probing rates, while injection redirects the models toward different tools. Together, these findings characterize hypothesis prompting as action shaping: its repair effect varies with the model's existing tool-use behavior.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.