MoDiff: Learning What Is Hard for Spatial Vision-Language Models
Abstract
Spatial VLMs are often trained with curricula whose difficulty is fixed by relation count, viewpoint change, or task stage. Such scores describe the task, but not necessarily the current learner: a simple depth relation may expose a visual blind spot, whereas a complex question may admit a language shortcut. We introduce MoDiff, a posterior curriculum correction framework for this mismatch. MoDiff maps static difficulty to an item prior and updates a shared multi-skill ability state and item difficulty from ordinary target-likelihood evidence. A posterior-averaged local information score guides sampling without extra VLM forward passes. Under a 25% token budget, MoDiff reaches 62.4% average accuracy across five spatial benchmarks, 3.1 points above static scheduling and 0.6 points above the strongest adaptive baseline in three-seed means, with 3.8% wall-clock overhead. These results foreground sample-efficient spatial fine-tuning through online correction of a fixed curriculum.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.