acceptodds
Under review as a conference paper at ICLR 2027

Learning Safety-Preserving Task Degradation for Fault-Tolerant Robotic Manipulation

Abstract

Learning one secondary-task coefficient offers a narrow interface for fault-tolerant manipulation with given fault reports and priority control. We distinguish changes to a trained policy from useful constrained control. In Isaac Sim, six PPO actors deploy constant full retention despite stochastic training, as established by guarded final-layer logit bounds. A paired intervention removes only their initial full-retention bias, retaining architecture, geometry, optimization, budgets, and the selection rule. Training and selection change the action map and produce variable deployed behavior, but all prespecified constrained-control criteria fail. Across 992 quarter-tolerance tasks per neutral fit, policies have 6-12 primary-crossing cases; a selected margin rule has three and weakly improves recovery count and secondary cost over every fit. A subsequent same-input diagnostic establishes a material replacement of the trained categorical distribution by its mode. A same-weight native deployment intervention then executes both choices: sampling improves some costs and recoveries but introduces 46 primary-crossing episodes across twelve sampled arms, versus zero in their six modal counterparts. None of the three prespecified fixed-penalty fits meets the useful-decoder criterion. Complementary selection and force-budget interventions establish bounded scheduling opportunity; failed gradient, residual, and fixed-probe screens retain their stopping conclusions. Removing constant deployment through initialization is insufficient for constrained-control gains. These findings neither identify a complete explanation of the remaining loss nor establish RL impossibility or a physical safety certificate. Two archived hardware trials from a separately deployed PPO policy document scalar execution.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.