Exact Expert-State Reuse for Efficient Training-Plan Search
Abstract
In multi-expert adaptation, comparing assignments of training examples to experts is costly because each candidate requires a sequence of parameter updates. Different assignments can give an expert the same ordered training history, even after their global plans diverge. Factorized Counterfactual Training (FCT) exploits this repetition by computing each required expert update once and retaining both parameters and optimizer state. With independent expert updates and a fixed numerical program, it preserves candidate states, scores, and deterministic search decisions for the same requests. A controlled four-expert workload reduces updates from 3,124 to 124, with a warm speedup over whole-prefix reuse. On SciQ and ARC-Challenge across four language models, FCT avoids 36.8–46.9% of prefix-cached updates and preserves every test prediction. All 24 paired runs are faster, with median speedups of to including graph preparation and full testing on a resident model. A separate 120-second search on Qwen3-0.6B more than doubles mean candidate counts and improves mean test accuracy by 0.93 and 0.20 percentage points, respectively. These savings make it possible to compare more training plans within a fixed search budget.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.