Sampled Pragmatic Adaptation: Zero-shot Pragmatic Inference for Language Models
Abstract
Large language models (LLMs) may assign high probability to various responses which are plausible for a query but are also compatible with nearby interpretations of the user’s intent. We introduce Sampled Pragmatic Adaptation (SPA), a zero-shot inference-time adaptation method that use this ambiguity to improve its output distribution by generating a local neighborhood of query-specific alternative intents. We formulate the adaptation through a minimum-intervention principle: among distributions satisfying query-local marginal constraints, we seek the one that minimally departs from the original distribution in KL divergence. The resulting correction connects pragmatic reasoning under the Rational Speech Act framework, iterative proportional fitting, and prior-shift calibration. We derive utterance- and token-level SPA variants that require neither training nor task-level calibration data, making the approach applicable to open-ended sequential generation. We study the method on tool invocation, a controlled setting in which models must select and construct function calls from conversational context and tool specifications, and introduce StaTI, a static benchmark assembled from existing datasets to evaluate tool invocation independently of downstream execution.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.