acceptodds
Under review as a conference paper at ICLR 2027

FIRE: Failure-Informed REcovery Post-Training for Real-Time Visual Combat

Abstract

Real-time visual combat exposes a mismatch between how human demonstrations are collected and where a learned policy needs supervision. Expert data are typically collected open loop and emphasize successful trajectories, leaving policy-induced failure states underrepresented. As the policy evolves, these deficiencies also shift, making additional generic demonstrations increasingly inefficient. We argue that post-training should instead close the loop between policy experience and human data acquisition, using rollout failures to determine what should be demonstrated next. We introduce FIRE: Failure-Informed REcovery Post-Training, a closed-loop framework that turns the policy's evolving failures into a curriculum for human supervision. Rather than collecting more generic expert play, FIRE uses recurring rollout failures to identify missing recovery behaviors, acquires demonstrations targeted to these deficiencies, and feeds them back into subsequent policy updates. A multimodal coach translates recurring failures into actionable recovery requests. The resulting demonstrations are organized in a recovery library and retrieved according to the encountered failure modes during subsequent supervised updates, while rollout feedback supports critic-free policy optimization and selective replay. Across seven Black Myth: Wukong bosses and two character configurations, FIRE improves normalized Boss HP Reduction over supervised fine-tuning (SFT) on all 14 boss-configuration pairs and Win Rate on 13. Under an equal additional demonstration-video budget, FIRE achieves higher final HP Reduction on all three controlled encounters than either additional successful demonstrations or Static recovery targeting only the initial SFT policy's failures. Supplementary videos are available at https://fire-anonymous-supplement.pages.dev/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.