VERDI: Verified Contradictions for Continual World-Model Optimization
Abstract
World models serve different domains and are tuned to different metrics, from visual fidelity to physical consistency and efficiency, so automating their optimization is valuable. Their optimization operators—reusable training, representation, sampling, or inference changes—look portable, yet models agreeing on architecture class, size, and baseline score can respond to the same operator in opposite directions, so what repairs one degrades its nearest neighbor. Existing methods supply the search but not the answer: they rank candidates by observational descriptors, assume these predict the same intervention response, and discard failed transfers as noise. We argue that an operator's effect sign is fixed by the model's response along the direction the operator modifies, not by any descriptor, so similarity is not a sufficient statistic and no fixed description can repair itself. The repair signal is a verified contradiction—nearby models whose opposing outcomes for the same operator are established by probing and by paired experiments under a verifier the search loop cannot edit. VERDI turns this into recursive self-optimization through shared-probe Optimization Fingerprints, IRG-guided retrieval, frozen-verifier settlement, and probe admission gated by held-out replay. Across nine backbones, 27 leave-one-backbone-out campaigns, and 216 transfer decisions, VERDI predicts effect signs with 83% accuracy, reaches the first verified positive in 3.6 trials and 96 GPU-h (68% and 69% below cold start), and lowers negative-transfer risk from 0.34 to 0.06 at 47% coverage. Ablations confirm that the geometry, not permissive acceptance, does the work, and a small-scale Franka Panda study gains 17.5 percentage points on two tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.