VECPO: View-Equivariant Constrained Policy Optimization for Robust Multi-View Remote Sensing Agricultural Reasoning
Abstract
Agricultural multimodal large language models must reason over heterogeneous observations spanning ground-level, unmanned aerial vehicle, and satellite imagery. Although supervised fine-tuning on multi-view data improves domain knowledge, subsequent reinforcement learning faces two core challenges: (i)transformation inconsistency, where predictions become inconsistent under semantics-preserving spatial transformations; and (ii)hidden group regression, where aggregate optimization conceals regressions on underrepresented view–task groups. To address both challenges, we propose View-Equivariant Constrained Policy Optimization (VECPO), a critic-free framework comprising Dual-Granularity Constrained Advantage (DGCA) and the Warm-Start Group Constraint Controller (WGCC). VECPO constructs an identity branch and eligible transformation branches with jointly transformed visual inputs and targets, canonicalizes predictions into a shared reference frame, and assigns correctness-gated equivariance rewards. DGCA couples a source-level relative advantage across all branches with branch-level relative advantages, while WGCC monitors view–task performance against supervised warm-start anchors and supplies group weights that constrain the dual-granularity advantage when a group deteriorates. Experiments on the 13-task AgroMind benchmark, including comparisons with PPO and multiple critic-free policy optimization baselines, show that VECPO achieves strong and competitive performance across diverse agricultural perception and reasoning tasks. Code is available in the supplementary material.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.