Multiple-Choice Supervision for Embodied Vision-Language Models: A Data-Centric Fine-Tuning Framework
Abstract
Supervised fine-tuning (SFT) can improve the perception and reasoning of embodied vision-language models (VLMs), yet how supervision should be formulated to promote transfer across embodied tasks remains unclear. In particular, comparisons between choice and open-ended training conflate the information supplied by candidate answers with the form of the supervised target. To disentangle these factors, we introduce Embodied Choice Tuning (ECT), a data-centric framework that jointly designs capability coverage, candidate inputs, distractors, and output targets. ECT retains candidate sets with confusable distractors in the input and supervises the correct option with a compact letter target. We instantiate ECT with a corpus of 10,000 examples spanning eight embodied capabilities and 102 question types. From the same visual observations, questions, and correct answers, we derive four supervision formats that isolate candidate availability, distractor semantics, and output targets. We evaluate the resulting models through Embodied Arena on nine established benchmarks in their native task formats. Across three model families, ECT yields the highest aggregate performance among the four formats. On Qwen3.5-4B, this advantage amounts to 3.89 points over the base and 2.27 points over open-ended supervision. Controlled comparisons further show that candidate inputs improve transfer when answer targets are fixed. A controlled distractor intervention shows that confusable distractors outperform clearly mismatched alternatives for all three choice targets. With candidates fixed, supervising answer text, either alone or alongside the letter, does not outperform letter-only supervision. Together, these results support ECT as a practical approach for data-limited embodied SFT: retain confusable semantic alternatives in the input and supervise the final decision with a compact letter target.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.