SuffiKV: Learning Sufficient Task States via Layer-wise KV Memories for Scalable Tabular Foundation Models
Abstract
Tabular foundation models (TFMs) enable supervised prediction by conditioning a frozen model on a labeled support table, but this makes the table part of every inference computation. As support sets grow, full-context inference becomes a latency and memory bottleneck, and additional raw rows may be redundant, noisy, or poorly aligned with the frozen attention mechanism. We introduce SuffiKV, a test-time adaptation framework that compiles a support table into a compact layer-wise key-value task state for a frozen TFM. SuffiKV separates the table used as evidence from the read state used at prediction time: a write stage constructs context-derived KV anchors, a bounded trust-region residual refines them, and a read stage predicts future queries using only the cached KV state. The state is optimized by anchored operator distillation, using full-context predictions and attention effects as early functional guidance while held-out pseudo-query labels refine the state toward supervised task risk. This turns tabular in-context learning into task-state inference: more data can improve the written state without increasing read-time context length. On a long-context tabular benchmark, SuffiKV improves the accuracy–latency–memory Pareto frontier, reducing subsequent inference latency by 89.9 and peak memory by 66.2% while improving AUC by 1.5%. Code and models will be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.