acceptodds
Under review as a conference paper at ICLR 2027

Amortizing Cross-Example Responses from Large Language Models

Abstract

An example placed in context can change how a large language model (LLM) responds to other inputs. We show that this cross-example response can provide useful supervision when outcome labels are scarce. For each tabular record, we place the same record in context with each possible label and measure the resulting responses on a set of unlabeled probes. A small student learns to predict this response from the original features without outcome labels; the scarce labels enter only through the final classifier. The predictor distilled from the signed response reaches 63.61% balanced accuracy and outperforms raw logistic regression, Extra Trees, TabPFN-3.5, and TabICLv2. Matched controls show that the gain comes from the signed change induced by the hypothetical label. Altering the learned predictor while leaving its predictions on every labeled example unchanged substantially reduces held-out accuracy. The response data therefore supply predictive structure that the labeled examples alone do not determine. On a distinct low-label task with a different LLM backbone, the signed-response student improves balanced accuracy by 7.11 points over raw logistic regression and attains the highest point estimate among the tested methods, including XGBoost. In network-flow classification, useful information instead remains distributed across individual probe responses and is largely lost by scalar averaging. The LLM is used only offline to produce the response supervision; once the student is trained, new records require no LLM calls.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.