acceptodds
Under review as a conference paper at ICLR 2027

When Output Budgets Masquerade as Representation Failures in Program-Based Reasoning

Abstract

Comparisons between structured output formats usually give every format the same generation-token cap. When two formats serialize the same computation at different lengths, an equal cap is not an equal opportunity to finish. We show how to identify this confound inside one fine-tuned program representation with paired cross-budget trajectories, and what it hides. In a program-based mathematical-reasoning benchmark whose 512-token evaluation appeared to show flat JSON programs collapsing on five-call problems while compositional plans did not, we re-decode the same three flat checkpoints at 512 and 1536 tokens, with and without JSON-schema constraints, in 24 arms with 19,752 predictions. On the five-call slice the 512-token cap ends 65.7% of skeleton-disjoint and 68.5% of augmented outputs by length; at 1536 tokens no output is cut, parse rates rise from 35.2% and 31.7% to 100%, execution by 60.0 and 63.2 percentage points, and accuracy by 19.0 and 22.8 points. Three tests locate the mechanism: truncating the saved 1536-token trajectories at 512 tokens reproduces the short-cap five-call accuracy exactly and 98.6% of row-level outcomes; exact-prefix pairs carry 247 of the 253 net correct gains; and offline estimates for 640, 768, and 1024 tokens, registered before decoding, match 97.9% of decoded correctness outcomes. Tripling the cap adds 14% generated tokens on the five-call slices. Validity recovers in every held-out five-call skeleton and on all 368 length-finished four- and six-call rows, while the accuracy gain is carried by one skeleton: truncation had turned wrong computations into validity errors. The same censoring appears in every seed on all 1,319 GSM8K test problems, including 206 whose audited reference computation has no training structure, and in a second model family, Llama-3.1-8B-Instruct, with offline truncation predicting at least 98.0% of row outcomes. Executable and semantic audits find 43 test rows whose labels contradict the text and 110 whose quantities contradict a premise of the question; on the rows that pass both, the five-call accuracy gains are 27.8 and 43.7 points. A semantics-preserving inliner shows that the historical converter changed the returned value of 6.0% of training targets. Retraining Qwen3.5-9B on matched targets in both formats and decoding each to saturation leaves a compositional five-call advantage of 10.5 and 9.2 points over the inliner's flat serialization. It is positive in every seed, under exact matching, on average over audited rows, and with any skeleton or source left out; on the GSM8K test split the advantage is 4.1 points. Format comparisons therefore need per-format saturation checks before any representation-level interpretation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.