acceptodds
Under review as a conference paper at ICLR 2027

TRACE: Target-State Risk-Aligned Conditional Execution for Embodied Tracking

Abstract

Embodied visual tracking requires a robot to follow a language-specified person under occlusion, look-alike distractors, and ambiguous instructions. We show that improving tracking persistence can simultaneously increase safety violations. Using a compact B vision-language-action (VLA) tracker with a frozen backbone, we update only the planner head with proximal policy optimization (PPO). PPO improves success and following rates across all three EVT-Bench settings, but increases the collision rate on distracted tracking because episodes that previously ended in target loss can instead end in contact. We introduce , an execution-time target-state risk gate that preserves the learned policy and intervenes selectively. A lightweight risk head processes cues, including candidate boxes, identity margins, a learned range-and-bearing probe, and the policy's commands. When the calibrated risk exceeds a threshold, rescales the translational component of the next waypoint while leaving yaw unchanged. The executed action is identical to PPO's on more than of steps. In paired evaluations on identical episodes, reduces the distracted-tracking collision rate by percentage points, while further improving success and following rates. Matched slowdown controls without target-state timing do not reproduce this improvement. Concurrent evaluation reproduces a collision reduction of points, compared with points for identical-arm null pairs, indicating that the effect exceeds evaluation variability. These results show that target-state-aware execution can recover the safety cost of policy post-training with minimal intervention.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.