acceptodds
Under review as a conference paper at ICLR 2027

REVIVE: Counterfactual On-Policy Diagnosis And Recovery for Interactive Ultrasound Segmentation Agent

Abstract

Recent MLLM-based segmentation agents automate multi-round interactive segmentation by generating corrective prompts for SAM-based models. However, they primarily emphasize action imitation while lacking explicit diagnosis of the current error state before action selection. Moreover, trajectory-level reinforcement learning optimizes sampled actions through sparse reward signals but does not explicitly compare alternative corrections at the same on-policy state, limiting recovery from suboptimal decisions and allowing errors to accumulate across interaction rounds. In clinical settings, human annotators follow a "diagnose-act-verify" self-correction loop: they need to diagnose the current mask errors before acting and then verify their effectiveness from the updated mask. We propose REVIVE, a self-improving segmentation agent that explicitly diagnoses the residual-error state before action, which further translates residual error evidence into action-specific representations. Then, the Counterfactual On-Policy Action Ranking (COAR) is proposed to execute alternative corrective actions at on-policy states and align the action likelihoods with the resulting segmentation outcomes, providing dense state-wise supervision for corrective action-type selection, while its integration with trajectory-level GRPO improves multi-round interaction behavior. Experiments demonstrate that our REVIVE consistently improves segmentation quality and interaction reliability, while reducing error-action misalignment and improving recovery from self-induced errors.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.