acceptodds
Under review as a conference paper at ICLR 2027

AndroidFlux: Evaluating and Reward Modeling Failure Recovery Capabilities in Mobile-Use Multimodal Agents

Abstract

Mobile-use multimodal agents are typically evaluated by task success from a clean initial state, which does not isolate their ability to recover from intermediate errors. This capability also matters when a stronger agent takes over an unfinished task and inherits another agent's mistakes. Hence, We introduce AndroidFlux, a failure-recovery benchmark built on AndroidWorld. Recorded trajectories are restored to controlled checkpoints before, at, and after the first recoverable error, allowing different agents to continue from the same device state and history. A complementary reward-modeling benchmark evaluates recovery-action preferences using downstream rollout outcomes. Moreover, we release a reward-modeling training dataset of 1.8k first-error contexts with candidate recovery actions and an offline reinforcement learning (RL) dataset of 8.5k contexts to enhance both error recovery and overall task success. We also introduce Latent Peer Voting (LPV): each expert scores actions proposed by its peers using conditional likelihoods, and these scores are normalized within each candidate set and combined with performance-based weights to produce preference supervision. Our evaluations show that recovery performance only partially correlates with clean-start success, and that explicit reasoning does not consistently improve recovery. Also, we show that LPV aggregates the internal action preference representation in expert policies and its collective preference can be distilled into a single reward model.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.