acceptodds
Under review as a conference paper at ICLR 2027

Sensitivity Is Not Validity: LLM Plan Judges Predict Execution Worse Than Overlap

Abstract

We ask whether plan evaluators predict what happens when the plan is executed – in AI2-THOR, under an oracle low-level controller that removes perception and actuation as causes of failure. Plan-level evaluation promises to separate an embodied agent's reasoning from its perception and control, but only if its scores predict something real. We meta-evaluate deterministic, symbolic, embedding-based and rubric-guided LLM-judge evaluators by how well each predicts execution of the same plans. Across 24,615 rollouts of 1,641 ALFRED-derived episodes, a plain overlap metric predicts execution better than EMBER – our own rubric judge, a contestant here rather than a proposal (primary pool, n=6,564: AUROC 0.890 vs. 0.786; preregistered difference -0.104, 95% CI [-0.125,-0.082]). The deficit holds for all fifteen planners we executed, eleven of them in four families scoring the same episodes, with a mean of -0.114 over those eleven. One family is architecturally disjoint from the one we developed on, and one planner is multimodal, its image moving plans but not execution. A frontier judge given the same inputs still trails (0.857). The judge's usual defence – that it credits valid alternatives overlap penalizes – fails on natural output and holds only weakly under control: on execution-verified rewrites of gold plans it beats LCS by a preregistered +0.102, but by +0.079 once no rewrite names a new object, it is near chance on precondition violations a symbolic checker catches, and an order-free overlap metric beats it throughout. A second preregistered test finds no localization gain over a single overlap scalar. The ordering holds for decisions: choosing one of eleven planners' plans per instruction, overlap picks an executable plan 50.2% of the time, the judge 36.0%. What the judge adds with a reference is small (exploratory): +0.016 AUROC over overlap, reordering plans below the top rather than changing which is chosen, and +0.005 over a fitted composite of cheap text features, the strongest scorer we have (0.947). Without a reference, a frontier judge is the best option we tested. An evaluator's validity has to be established for the decision it serves. We contribute the protocol, a validated executor, and these execution-anchored results.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.