acceptodds
Under review as a conference paper at ICLR 2027

STEERACT: Temporally Gated Single-Feature Steering of Frozen Action Models

Abstract

Improving pretrained robot action models from a few rollouts requires identifying both what to steer and when to intervene and the sign of the correction. We find that successful and failed executions exhibit distinct temporal patterns in sparse latent space, and that steering one feature per task can improve closed-loop success without updating model weights. Based on this finding, we introduce STEERACT, a reference-guided steering method using a pretrained sparse autoencoder. Offline, STEERACT summarizes sparse-feature activation histories from successful and failed rollouts via temporal descriptors, constructs outcome-conditioned reference banks, and selects the feature with the strongest outcome separation. Online, a support gate assesses whether the observed activation history provides sufficient evidence for reference comparison, while a failure gate checks whether its descriptor is closer to failed than successful references. When both gates are open, a bounded search selects a signed correction predicted to move the current descriptor closer to the success bank. The selected feature direction and reference banks remain fixed; only the intervention coefficient adapts during execution. Using 24 outcome-labeled reference rollouts per task, improves task success by approximately 10% relative to unsteered policies in evaluations spanning the vision–language–action model , the world–action model Fast-WAM, RoboTwin 2.0, DexJoCo, and real-world manipulation. Code and Demo can be found at: https://anonymous-submission-20.github.io/steeract-sub.github.io/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.