acceptodds
Under review as a conference paper at ICLR 2027

StableBFM: Extending Behavior Foundation Models to Robust Behavior

Abstract

Pretrained behavior foundation models (BFMs) provide a reusable foundation for whole-body humanoid control by conditioning on large-scale reference motions and decoding actions, enabling high-fidelity tracking of diverse behaviors. However, real-world deployment also requires reliable closed-loop recovery from off-nominal states such as falls. How to extend the capabilities of a pretrained BFM to recovery while preserving its existing tracking ability and keeping training overhead low remains a key challenge. Given the current robot state and a reference motion, we define the state–reference discrepancy as the deviation between the current state and the target state implied by the reference. Standard tracking pretraining mainly constrains behavior in the small-discrepancy regime, whereas fall recovery requires the decoder to start from large-discrepancy states, reduce the discrepancy through closed-loop control, and return to the nominal tracking region. From this perspective, we formulate recovery as extending the effective closed-loop control range of a pretrained BFM, and unify tracking and recovery as the same state–reference conditional decoding problem across discrepancy regimes. To this end, we treat recovery as a capability extension of a pretrained motion-tracking BFM and perform simple, low-overhead post-training without adding model parameters. We retain the original tracking PPO objective, allocate a subset of parallel environments to recovery, and apply DAgger-style expert supervision to states visited by the learner, thereby extending recovery supervision to the large-discrepancy regime. We further design corresponding recovery-injection mechanisms for two decoder parameterizations. Experiments show that the proposed framework significantly improves recovery while maintaining tracking performance, and supports deployment on a real humanoid robot.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.