acceptodds
Under review as a conference paper at ICLR 2027

Task Calibrated Decoding: Improving Black-Box LLMs with In-Domain Data

Abstract

When using Large Language Models (LLMs) in downstream tasks, labeled in-domain examples provide a natural source of supervision. However, finetuning the underlying LLM can be costly or intractable, particularly when access is restricted to only sampling responses in a black-box manner. We propose Task-Calibrated Decoding, a lightweight procedure that uses labeled examples to improve decoding in a frozen, black-box LLM. We represent sampled free-form responses as a distribution over their task-relevant semantics, such as numerical ratings or class labels. Using in-domain labels, we calibrate the full distribution and apply Minimum Bayes Risk Decoding to select the response minimizing expected task loss. The procedure requires no access to model weights or token probabilities, and is provably optimal among decoding strategies based on the LLM’s output distribution under ideal recalibration. Empirically, it improves generation quality across baselines for different tasks and models, and we present qualitative examples showing how it reduces systematic biases. Our results demonstrate that task-calibrated decoding is a theoretically motivated, lightweight, and broadly applicable strategy to exploit in-domain supervision by adapting downstream decisions while keeping the underlying model fixed

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.