acceptodds
Under review as a conference paper at ICLR 2027

Reinforced Attention Learning

Abstract

Post-training via Reinforcement Learning (RL) has catalyzed significant reasoning gains in Large Language Models (LLMs) by scaling test-time computation. While recent efforts attempt to extend this paradigm to Multimodal Large Language Models (MLLMs) by incentivizing verbose rationales, we demonstrate that such gains are marginal for perception; indeed, excessive verbalization can inadvertently degrade performance. We propose Reinforced Attention Learning (RAL), a policy-gradient framework that treats internal attention distributions as an optimization target. Unlike conventional RL that focuses on high-reward token sequences, RAL explicitly explores and rewards optimal information allocation patterns during the generation process. By shifting optimization from "what to generate" to "where to attend", RAL enhances the model's ability to isolate salient signals within complex multimodal contexts. Extensive experiments across diverse image and video benchmarks show that RAL consistently outperforms standard GRPO and other baselines. We further extend this approach to On-Policy Attention Distillation, proving that transferring latent attention behaviors to student models yields superior cross-modal alignment compared to standard knowledge distillation. Our results establish attention policies as a principled, generic alternative for multimodal post-training. We will release our code and data upon publication.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.