DRA: State-Conditioned Evidence Allocation for Missing-Modality Fine-Tuning
Abstract
When a modality disappears, the value of evidence still available but unread can change. We study state-conditioned evidence allocation: using observed modality availability to direct a limited budget of additional text encoding. A Bayes-risk characterization relates this choice to information substitutes and complements:the preferred state can reverse with the loss, and at an interior budget, state-blind allocation is optimal only at equal value per cost. DRA (Decide, Read, Adapt)implements the allocation with a state-gated residual reader that reuses an adapted CLIP text tower; we retain shared adaptation and validation thresholds as interfaces. On MM-IMDb, same-movie image removal with symmetric reader training gives a positive log-loss interaction of 0.00839 nats per label (movie-cluster95% CI [0.00767, 0.00905]). With the reader and extra-window budget fixed, targeting image-missing rows improves macro-F1 by +0.18 and AP by +0.33 over state-blind allocation. In retrained controls, DRA comes within 0.19 F1 of all-row reading using 53% of the extra windows. Under published evaluation rules across MM-IMDb, UPMC-Food101 and Hateful Memes, DRA exceeds reported MoRA in 23 of 27 benchmark settings
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.