acceptodds
Under review as a conference paper at ICLR 2027

Language-Conditioned Executable World Hypotheses for UAVs: Causal Evaluation and Certified Resolution

Abstract

Language-conditioned UAV systems can emit syntactically valid, executable actions whose behavior, referents, or evidential support are nevertheless wrong. Existing evaluations usually isolate interface validity, semantic accuracy, or trajectory error, obscuring whether a model’s complete decision is both justified and executable. We formalize the executable world hypothesis, a contract-bound claim that couples an intended operation with typed grounding, epistemic support, and deterministic rollout, and define Executable-World-Hypothesis Correctness (EWHC) as its end-to-end criterion. We instantiate this problem in AeroWorld32K, a benchmark independent of model family with 4,096 executable UAV scenes and 32,768 causally paired queries spanning supported and open world conditions. We further introduce Certificate-guided Epistemic Resolution for Executable States (CERES), a deterministic resolver that requires no training and audits one frozen model completion against a typed public evidence representation, certifying a uniquely grounded and supported interpretation or compiling an explicit abstention to the protocol defined safe hold. Across ten frozen languagemodel configurations and six deterministic comparators, unassisted models remain limited particularly at the open world support boundary, and a deterministic symbolic resolver outperforms every unassisted language model. Giving five representative models the same public information in a typed representation closes a mean 9.44% of their model-wise Direct-to-CERES gaps. CERES yields an average absolute EWHC gain of 35.07% across all model configurations, repairs 114,908 Direct errors, and causes no observed EWHC harm, harm on supported cases, or unsafe determinate output on the complete evaluation population. Preregistered closed loop simulations further show that small ADE/FDE can coexist with an incorrect executable world hypothesis. These results establish evidence grounded resolution as a practical complement to language conditioned action generation under open world uncertainty.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.