acceptodds
Under review as a conference paper at ICLR 2027

Same Task, Different Representation: When LLM Success Fails to Transfer

Abstract

Large language models (LLMs) often solve tasks using familiar output representations, such as implementing Boolean logic in Verilog/Python or expressing chess moves in standard notation. Yet users may need the same task solved in a different representation, requiring models to apply acquired skills under new constraints. Does demonstrated success transfer when the required output representation changes, and can additional reasoning recover lost success? We thus introduce Paired Representation Robustness (PRR), which measures the fraction of problems a model already solves that remain correct under an alternate output representation. The semantic task stays fixed, and the alternate representation’s rules are explicitly supplied. We evaluate five LLMs on 500 Boolean logic problems (Python code → Minecraft Redstone circuits) and 1,000 chess puzzles (standard move notation → a synthetic cipher notation, Syn-Chess). With CoT, and additional reasoning disabled, GPT-5.6-Sol solves 98.8% of logic problems in Python but retains only 1.01% of those successes in Redstone; it identifies the correct first move in 86.6% of chess puzzles but retains only 0.69% under Syn-Chess. On a held-out subset of 100 Syn-Chess puzzles, adaptive reasoning raises first-move PRR from below 4% to approximately 79% for both GPT-5.6-Sol and GPT-5.6-Luna, at higher inference cost. These findings reveal a barrier to applying demonstrated skills under new output requirements: transfer can collapse even with explicit CoT prompting, while adaptive reasoning recovers much of the lost success at over 10× the baseline inference cost. Output representation thus emerges as a consequential factor in both the reliability and computational cost of LLM problem solving.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.