acceptodds
Under review as a conference paper at ICLR 2027

Calibrated Tool Use via Learned Adapters to Compiled Neural Programs

Abstract

Language models typically call a tool with a single set of arguments, even when several readings of the input remain plausible. This discards useful information, and as a result, reinforcement learning for tool use must estimate credit by sampling many trajectories. We propose SPACE-LLM: Sampling-free Probabilistic Adapters for Calibrated Execution in Large Language Models. Learned adapters over LLM hidden states predict distributions over the arguments of a compiled neural program, a symbolic program written as tensor operations that execute inside the model's own forward pass. Retaining these argument belief distributions allows inference to combine support for an answer across ambiguous readings and enables training to compute analytic gradients from a ground truth answer. We show that maintaining belief distributions at inference, rather than committing to one reading, raises accuracy and improves calibration without a calibration-specific loss. Using only answer-level supervision, we observe comparable or stronger out-of-distribution generalization than fully supervised finetuning and a comparable Brier score, from a single pass without sampling or a recalibration step. On identical training sets, SPACE matches or exceeds GRPO accuracy with a lower Brier score more than 50 times faster, and the tool runs inside the model's forward pass. These advantages are limited to domains in which maintaining the factored belief state is tractable, the answer carries enough information to credit each argument, and the order of operations is known in advance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.