FALSIFICATION-TIME COMPUTE: SCALING INFER-ENCE BY SEARCHING FOR COUNTEREXAMPLES
Abstract
Additional inference-time compute can change the evidence available for choos-ing an answer, rather than generate another answer. Falsiffcation-Time Compute (FTC) actively acquires executable worlds—witnesses—that distinguish surviving candidate answers. It separates two questions: can search ffnd a useful disagree-ment, and what semantic authority justiffes acting on it? On 100 Spider tasks with frozen model-generated candidate pools constructed to pass weak public ex-ecution tests, FTC improves selection from 66% to 73% under true Test Suite Accuracy (TSA), against an 83% pool oracle, closing 41.2% of the measured gap. Controlled and natural-pool comparisons show that candidate conditioning strengthens weak proposals, while sufffciently strong shared search can also ffnd useful distinctions. Open witness search retains positive selection value under TSA without explicit hand-written failure-family routing. Yet ffnding evidence does not ensure its correct use: candidate-blind learned adjudication introduces an additional regression in a fully local SQL conffguration and loses both avail-able search recoveries in a small code diagnostic. FTC treats executable evidence discovery and the semantic authority needed to use it as distinct inference-time computational problems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.