acceptodds
Under review as a conference paper at ICLR 2027

MoDiff: Learning What Is Hard for Spatial Vision-Language Models

Abstract

Spatial VLMs are often trained with curricula whose difficulty is fixed by relation count, viewpoint change, or task stage. Such scores describe the task, but not necessarily the current learner: a simple depth relation may expose a visual blind spot, whereas a complex question may admit a language shortcut. We introduce MoDiff, a posterior curriculum correction framework for this mismatch. MoDiff maps static difficulty to an item prior and updates a shared multi-skill ability state and item difficulty from ordinary target-likelihood evidence. A posterior-averaged local information score guides sampling without extra VLM forward passes. Under a 25% token budget, MoDiff reaches 62.4% average accuracy across five spatial benchmarks, 3.1 points above static scheduling and 0.6 points above the strongest adaptive baseline in three-seed means, with 3.8% wall-clock overhead. These results foreground sample-efficient spatial fine-tuning through online correction of a fixed curriculum.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.