ComboBench: From Semantic Goals to Device-Command Plans
Abstract
Translating a goal into device commands requires more than naming relevant actions: a plan must specify the right controller, operation, parameters, and sequence. We introduce ComboBench, a text-only benchmark of device-command planning drawn from four virtual-reality games. Goals and controller conventions are mapped to ordered manipulation steps, allowing command specification to be studied separately from visual perception and online control. The twelve-model evaluation reveals complementary strengths rather than a single best system. Five demonstrations increase reference-prefix recovery for every model in the reported aggregate, yet strict full-sequence matching remains at most 11.7%. A 50-scenario study with additional valid reference paths examines the ambiguity of single-path scoring, alongside a human text baseline and game-level analyses. We further investigate what these scores measure: a semantic matcher accepts 263 of 374 constructed command-attribute mismatches, and snapshot-wide normalization can give an empty plan a score of 0.867. A 68-scenario reanalysis supplies paired uncertainty estimates. Together, the benchmark and diagnostics make device-command specificity measurable while distinguishing reference agreement, format validity, and execution success. The artifact is available at [https://anonymous.4open.science/r/ComboBench-623B/](https://anonymous.4open.science/r/ComboBench-623B/).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.