TextMani: A Constraint-Based Framework for Evaluating and Repairing LLM Reasoning from Language to Robotic Action
Abstract
Translating large language model (LLM) reasoning into robotic actions requires an interface between semantics and execution. This interface should make intermediate decisions inspectable and editable.Moreover, task success rates alone provide limited insight into task progress and points of failure throughout this process. We introduce TextMani, a unified framework for language-to-action execution, evaluation, and repair that uses explicit constraints as a shared intermediate representation. Given natural language instructions and scene features, TextMani generates execution targets through subgoal decomposition, feature selection, constraint formulation, and mathematical solving. Planning and control modules then translate these targets into robot actions. By linking intermediate artifacts to execution evidence, the framework supports process-level analysis based on physical subgoal progress and enables targeted repair following failures. Guided by execution feedback, it selects stages or constraints for modification, preserves unaffected content, and updates dependent downstream artifacts before re-execution. We evaluate 15 models across 18 simulated tasks and 10 models across six real-world robotic tasks. On real robots, TextMani improves task success from 27.2% with direct numerical-target prediction to 36.3%. Feedback-driven repair is also effective: Task-specialist repair recovers 47.7% of failed tasks compared with 24.5% for self-repair. Blind diagnosis reveals that failure recognition is the primary bottleneck.These results show that explicit executable specifications provide a common interface not only for comparing heterogeneous foundation models, but also for exposing where execution fails and enabling targeted physical recovery.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.