acceptodds
Under review as a conference paper at ICLR 2027

Layer-Edit Coverage Beyond Sampling: From Candidates to Selected Answers

Abstract

Can layer-execution edits supply correct answers that ordinary sampling misses, and can a selector recover them? We study skipping, repetition, and substitutions that combine the two in frozen language models. Our grid contains approximately 14.2 million correctness evaluations over 1,261 paths, five mathematical difficulty strata, and four prompt conditions on Qwen2.5-3B-Instruct, with a Qwen3-4B comparison. Across three model–difficulty groups of 565 questions each, the exhaustive 1,260-edit pool covers 15–56 questions missed by all 57 unedited-model samples. Yet question-held-out path ranking and direct 57-path voting correctly answer only 0–3 of these additional questions and trail 57-sample self-consistency by 23–57 correct answers per group. Shortlists exclude many available answers, and voting rarely recovers those retained. The edit grid also has reproducible structure: average substitution performance tracks constituent effects, while 45–51% of discovered substitution events with failed constituents succeed again on the rewrite prompt. Tested pre-generation rankers show no improvement over a structural prior, despite informative post-generation correctness signals. These findings separate additional candidate coverage from the ability to select a correct answer. They use original-checker scores and previously inspected questions and outputs; unequal coverage pools and unmeasured end-to-end costs preclude an efficiency claim.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.