acceptodds
Under review as a conference paper at ICLR 2027

RECAP: From Requirement Recovery to Reliable Execution in Context Learning

Abstract

Large language models are increasingly used in long-context settings where the context defines temporary knowledge, rules, procedures, and constraints. Many failures are near misses: models capture the main task but miss some low-salience requirements. Yet improving requirement recall alone can produce redundant and fragmented representations that are difficult to execute, making context learning a joint problem of broad requirement recovery and executable representation. We study whether such representations can be discovered from empirical feedback rather than designed entirely by hand. To this end, we perform feedback-driven search over a structured space of executable inference pipelines, treating each edit as a falsifiable hypothesis and retaining regressions, ineffective edits, and non-transferring changes as constraints for subsequent search. Within this structured search space, feedback-driven evaluation supports broad requirement enumeration followed by consolidation into a compact representation before execution. This design yields RECAP, which constructs such a representation at inference time and conditions generation on it. For weaker models, we further introduce RECAP-RL, which learns to execute self-constructed requirement representations using rubric-level rewards. On CL-bench, RECAP improves Claude Opus-4.8 from 21.09% to 27.04% strict accuracy and Qwen3.5-9B from 11.74% to 15.60%, while RECAP-RL further raises Qwen3.5-9B to 17.96%. The approach also transfers to complex instruction following and multimodal context learning. These results suggest a general recipe for context learning: discover an effective requirement representation, construct it explicitly, and learn to execute it.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.