Fix-Test Collapse: Diagnosing and Resolving Self-Confirming Verification in Code Agents
Abstract
Self-testing lets code agents inspect execution feedback and repair errors, yet successful local validation can precede failure on held-out tests. We call this failure Fix–Test Collapse and distinguish tests that miss bugs from misinterpretation of outputs that expose them. We propose Falsification-Guided Self-Testing, which combines protocol SFT, action-level reinforcement learning, and trajectory rewards under a single-trajectory interface that withholds reference outputs and hidden verdicts from the policy. Offline scorers reward diagnostic tests on contest tasks, and an issue-grounded judge, updated with on-policy hard negatives, supplies the diagnostic signal on repository tasks. Across three model families, pass rates improve by up to 34.7% on CodeContests and 19.5% on LiveCodeBench relative to untrained tool-enabled policies, and Qwen3.5-35B reaches 48.3% resolution on 731 SWE-bench Pro tasks and 28.3% on the complete 113-task DeepSWE split. In a matched control starting from an identical SFT checkpoint, falsification RL outperforms final-only RL on all three paired training seeds, by 7.7 points on Pro and 10.6 points on DeepSWE on average. Training also raises self-test coverage and lowers collapse among locally validated submissions; this conditional reliability measure is distinct from an all-instance failure rate.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.