When Acceptable Answers Share Parts, Feedback Helps More
Abstract
A correct observation about one part of an answer can help with many acceptable answers, or with very few. We study whether this shared structure changes the value of component-level feedback. Our controlled task asks language models to recover five terms from a forty-term catalogue. We vary overlap among acceptable answers while holding their number, the numerical data, report format, and verifier allowance fixed. Across 24 families and 11 models, the graded-recovery advantage of the full over a profile-substituted report increases with overlap: the interaction is +0.419 [+0.323, +0.508] per unit of a chance-corrected overlap index. Correct-minus-no-report recovery changes from −0.012 at zero overlap to +0.366 at high overlap. A calibrated mechanical reader, defined by an explicit expected-overlap objective, has a graded interaction of +0.440 and a larger exact-recovery interaction than the models. Fixed-report and proposal-conditioned studies on external SAT formulas show related positive associations; the fixed-report slope is +0.169 [+0.086, +0.251] after answer-count control. The construction, controls, and reader comparison show how a task property changes the usefulness of a specified report. They also separate recovering useful parts from assembling an acceptable whole. Code and data: https://anonymous.4open.science/r/verified-feedback-discovery-ready-7381/README.md.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.