Compiled Task Memory for Efficient Inference with Frozen Tabular Foundation Models
Abstract
Context optimization offers a way to reduce the inference cost of tabular foundation models by replacing large training contexts with compact learned representations. Building on this approach, we study compiled task memory: task-specific context vectors optimized offline and reused through a frozen predictor. Slots are initialized from stratified frozen row encodings and optimized using supervised loss and out-of-fold teacher guidance, with all backbone weights held fixed. Our study evaluates accuracy, retained state, and cached inference cost on TabDPT and TabICLv2 across 35 classification and 13 regression tasks, using complete training and test splits and three seeds. Compared with the same unoptimized 256-slot initializations, optimization improves mean classification accuracy by 2.14 and 2.55 percentage points and reduces geometric-mean RMSE by 19.2% and 17.6%, respectively, with essentially unchanged compact inference cost. On classification tasks, cached 256-slot inference achieves geometric-mean speedups of and over full cached context, with mean accuracy losses of 0.76 and 0.78 percentage points. One-time preparation costs are reported separately. Construction and teacher-weight ablations characterize how memory learning affects quality at fixed capacity. Across –1024, larger memories improve regression at higher serving and storage cost, while classification gains beyond 256 slots remain uncertain. These findings characterize the practical accuracy–cost tradeoffs of context optimization.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.