acceptodds
Under review as a conference paper at ICLR 2027

When to Constrain, Guide, or Defer: Adapting Frozen Language Models to Private Inventories

Abstract

In this paper, we explore how to adapt a frozen, served language model to a private, user-specific inventory, such as a model hub, a tool catalog, or a database schema, without modifying its weights. Focusing on the sampling space, we study two complementary mechanisms for inference-time adaptation: constrained decoding and a small expert trained on user-side usage logs. We analyze when each is beneficial, at what granularity it should be applied, and how they should be composed. In our experiments, constrained decoding yields gains only where the base model already possesses partial knowledge of the relevant names (+6 on the primary model where it does, down to where it does not; the boundary holds across the 46 host–catalog pairs we can probe), while the neural expert helps only within its training scope (+21 on its catalog, down to off it). We further find that the signal for deciding when an expert should act is weak in thresholded token-level statistics, only partly recovered by a learned per-token head, and present in a representation of the full request. We propose a prefill-based router that reads the expert's own prefill state to either apply adaptations or default to the base model. The same request-level verdict can activate both mechanisms at once: when both apply, the mask adds its guarantee that no invented name is emitted. We observe that the routed, unmasked coupling adds 16 points of accuracy on its target dataset and returns the base model exactly on unrelated math questions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.