CoT-Faith: Magnitude Scores Do Not Measure Faithfulness in Manipulation VLAs
Abstract
Vision-language-action models (VLAs) that emit chain-of-thought (CoT) are proposed as a route to interpretable robotics, on the premise that the emitted reasoning drives the action. The standard test scores an edit to the reasoning by how often it moves the action. We introduce CoT-Faith, a paired-intervention benchmark with 13 edit families, size-matched nulls and closed-loop evaluation, and apply it to 11 token-decoded VLA configurations from two model lineages. We show that the test’s verdict depends on which meaning-preserving edit serves as its floor: under teacher-forced scoring, two judge-validated floors give opposite verdicts on 10 of 11 configurations, including the two competent policies. The score mostly records whether the action changes at all, which depends on what an edit touches rather than what it means. Nulls matched to each edit in size and tag restore a consistent sign, with direction flips beating their null on 11 of 11, but a control fine-tuned without a CoT target passes them too. Requiring the action to point the implied direction moves the top-ranked model from first of eight to seventh. On two competent LIBERO policies the action is sensitive to the reasoning at each step, yet editing or even emptying it leaves closed-loop success unchanged: sensitivity is not reliance. Edit-based scores should be read against matched nulls, a direction-aware score and closed-loop success before they license a claim.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.