acceptodds
Under review as a conference paper at ICLR 2027

VECPO: View-Equivariant Constrained Policy Optimization for Robust Multi-View Remote Sensing Agricultural Reasoning

Abstract

Agricultural multimodal large language models must reason over heterogeneous observations spanning ground-level, unmanned aerial vehicle, and satellite imagery. Although supervised fine-tuning on multi-view data improves domain knowledge, subsequent reinforcement learning faces two core challenges: (i)transformation inconsistency, where predictions become inconsistent under semantics-preserving spatial transformations; and (ii)hidden group regression, where aggregate optimization conceals regressions on underrepresented view–task groups. To address both challenges, we propose View-Equivariant Constrained Policy Optimization (VECPO), a critic-free framework comprising Dual-Granularity Constrained Advantage (DGCA) and the Warm-Start Group Constraint Controller (WGCC). VECPO constructs an identity branch and eligible transformation branches with jointly transformed visual inputs and targets, canonicalizes predictions into a shared reference frame, and assigns correctness-gated equivariance rewards. DGCA couples a source-level relative advantage across all branches with branch-level relative advantages, while WGCC monitors view–task performance against supervised warm-start anchors and supplies group weights that constrain the dual-granularity advantage when a group deteriorates. Experiments on the 13-task AgroMind benchmark, including comparisons with PPO and multiple critic-free policy optimization baselines, show that VECPO achieves strong and competitive performance across diverse agricultural perception and reasoning tasks. Code is available in the supplementary material.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.