Universal In-Context Learning over Foundation Model Embeddings
Abstract
Foundation-model embeddings are widely used across many domains, yet using them for prediction usually requires training a separate classifier for each dataset. This adds engineering work, makes results sensitive to hyperparameter choices, and increases the risk of overfitting when labeled data are scarce. It also means that embedding benchmarks measure classifier optimization as well as representation quality. Meanwhile, tabular foundation models have shown that transformers can classify arbitrary tabular data in context, in a single forward pass and without task-level training. We adapt a tabular foundation model into a single in-context predictor that works across embedding spaces. Instead of training a separate head for each task, we meta-train one shared model across tasks and encoders. At every training step, we apply a freshly sampled random orthogonal projection that changes the coordinate basis of the embeddings while preserving their dimension and geometry, which prevents the model from relying on any fixed coordinate orientation. At inference, the model averages its predictions over multiple random orthogonal projections of the same inputs and requires no parameter updates. We curate a benchmark spanning 7 domains, 29 encoders, 75 tasks, 308 task–encoder pairs, and 169,554 samples, and evaluate it under in-distribution, novel-task, novel-encoder, and joint task–encoder transfer. Averaged over all 308 pairs, our model reaches 0.8886 ROC-AUC, improving on the strongest baseline by 0.0198 (95% CI ) without a single gradient step at inference. We release all code, benchmark data, and trained model checkpoints.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.