acceptodds
Under review as a conference paper at ICLR 2027

Severity Targets Can Reward Corrupted Labels: Diagnosing Sample Weights for Control

Abstract

A difficult trajectory is not necessarily a reliable one. We show that severity-based sample weighting can conflate these properties and reward corrupted action labels. With 20% of Push-T training episodes corrupted, applying a severity target directly gives corrupted windows 1.89× the clean-window weight; a learned weighting network retains the same preference. A direct-target control and a feature × target ablation locate the failure in the target rather than the network. Replacing severity-only supervision with a consistency-aware target changes the corrupted-versus-clean ordering, while an amplitude-matched control that preserves the original ordering does not recover performance. A jerk- and path-matched negative control confirms the consistency-aware target preserves genuinely hard but correct maneuvers while collapsing a matched corrupted twin, so it reads label reliability rather than dynamical roughness. Across four tested synthetic corruption modes, the effect follows the target components: corruptions that inflate action jerk or path length are up-weighted, whereas a constant offset that leaves them intact is nearly neutral. Five-seed Push-T experiments further show that target correlation and closed-loop utility need not agree. USV experiments provide complementary calibration diagnostics and the same separation between target agreement and task utility. We therefore recommend evaluating weight ordering, dispersion, corruption preference, and matched task outcomes separately.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.