Failure Buys a Turn: Correct Answers Are Its Hidden Casualties
Abstract
Many LLM evaluations retry only the answers that a validator, such as unit tests, rejects, so neither scores nor logs show what that turn would do to answers that pass. We prove that such logs leave the value of this policy anywhere between zero and the first-attempt score when the validator is the task’s own scorer. In unscored shadow runs that give accepted answers the same turn, the largest losses are outputdelivery failures: the correct answer is still in context, but the next reply omits it. On HumanEval, “Your submission has been received” makes Gemma-3-12B drop 129 of its 137 correct answers, while “Your submission was incorrect” on the same conversations drops five. When the reply replaces the first answer, eight of ten open models have a message that destroys 19% to 94% of their correct answers. Adding sentences reveals two routes to recovery: after an acknowledgement, a request to submit restores nearly every answer, but after our review request Qwen3-8B still loses 80 of 139 until the exact output format is requested. Without a validator, deleting the final request to restate the answer raises Llama-2-70b-chat’s GSM8K self-correction loss from 12 to 36 points, while deleting only its box changes little. Retry evaluations should report their continuation message and audit accepted answers.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.