PLANSPACE: SEPARATING REFERENCE REPRODUCTION, PLAN VALIDITY, AND INTERFACE COMPLIANCE IN SYMBOLIC PLANNING
Abstract
A household task can have several correct plans: groceries may go into different allowed cabinets, and independent placements may occur in either order. Scoring only against one recorded sequence rejects such alternatives. It also leaves a low score hard to interpret: did the model format its answer incorrectly, choose an action it could not execute, or fail to reach the goal? We introduce PLANSPACE, an evaluation protocol that answers these questions separately. It checks whether a submitted plan can be interpreted, whether it matches a reference or the dependencies of a known plan family, and whether its actions achieve the goal under an explicit symbolic model. Across six open-weight instruction-tuned language models on 171 supported BEHAVIOR-1K tasks, exact matching misses successful plans for every model. For Qwen3-8B, the discrepancy is especially large when tasks permit alternative goal choices or require longer plans. Other low scores arise before planning can be assessed: a small formatting correction exposes successful plans in one model’s existing outputs but leaves the other models unchanged. These findings show why reference agreement and task success should be reported separately, alongside the interface checks needed to interpret a plan.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.