Seeing is Believing: An Embodied Benchmark for Trust Calibration in Multimodal Agents
Abstract
Embodied agents receive instructions and reports on channels that are not equally reliable, and only their own perception can settle what is true. We ask whether current omni-modal models know whom to trust, and whether they can earn that knowledge by looking. We introduce , an embodied benchmark over dynamic household scenes in which a spoken and a written informant carry hidden, persistent reliability profiles, and the agent's own camera is the only means of verification, at a cost in navigation steps. Six evaluation levels, from perceiving the scene to acting on unverified claims, each isolate one capability with one primary metric, so that a failure is attributed to a capability rather than folded into one score. The capstone metric counts only test informants and only decisions the agent chose not to verify, so it cannot be farmed by trusting, declining, or verifying everything, and a ground-truth perception variant separates perception failures from trust failures. Models are evaluated without parameter updates, with the agent's own verification history as the only state carried across episodes. Across six open omni-modal models, the levels rank the models differently, so a single score would hide where each model fails. Several models favor whichever informant is labeled as speaking, the model with the highest memory score is not the one that acts best without checking, and only one model beats trusting every claim by a margin that does not survive our robustness checks. thus provides a farm-resistant testbed for embodied trust and a map of where current agents fail to earn it.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.