acceptodds
Under review as a conference paper at ICLR 2027

LangSelect: Cost-Aware Target-Language Routing for LLM Code Generation

Abstract

Code generators usually write in a programming language fixed before decoding. For function-level tasks that several languages can solve and the same unit tests can check, we study target-language routing under a verifier with LANGSELECT, a loop in which a policy orders candidate languages, a failed attempt falls back to the next language, and the code tokens of every attempt are counted. We build MultiLang-Bench, 3,000 tasks with verified solutions in eight languages. We show that the final pass rate depends only on which languages are tried and that, under a mild independence assumption, adjacent languages should be ordered by cost per success conditional on earlier failures. With GPT-5, a fixed language order emits 50.3% fewer code tokens than default-prompt Python on 450 held-out tasks (final pass rate 92.9% vs. 92.0%) and 48.2% fewer on a second, 400-task mix, where its pass rate is 2.8 points lower. A supervised selector and a LinUCB bandit that recover about 87% of the headroom in replay emit 3.7% and at least 13.9% more tokens than Python-only in live generation. A chain that ends in Python keeps the expected pass rate of Python-only; recomposed from logs, it saves 39% on the data used to choose it and 13% (95% CI−18 to 38%) in a fresh run on 60 of these tasks. Against a terse Python prompt, a cross-campaign estimate leaves the fixed order 11–15% and this chain no saving. With DeepSeek-Coder-6.7B, bandit routing emits 2.3 times as many tokens as Python-only. For GPT-5, generation time increases under routing, and reasoning tokens are not observed.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.