acceptodds
Under review as a conference paper at ICLR 2027

PASA: Post-Merge Perception-Reasoning Asymmetry as a Self-Alignment Signal for MLLMs

Abstract

Model merging offers a direct way to transfer specialized reasoning capabilities to multimodal large language models (MLLMs), but successful transfer must also preserve their ability to ground answers in visual evidence. We find that merging a reasoning-specialized language model into an MLLM improves reasoning performance while degrading fine-grained visual perception, a phenomenon we term post-merge perception–reasoning asymmetry. We investigate whether the reasoning capability retained after merging can itself provide supervision for recovering visual perception. Comparisons on the same self-alignment queries show that reasoning-prompted responses are more accurate than perception-prompted responses, suggesting a source of supervision within the model. Building on this observation, we introduce PASA, a two-stage framework for post-merge self-alignment. PASA first learns to generate structured visual-evidence traces through supervised fine-tuning on synthetic data. It then applies asymmetric GRPO: a reasoning-prompted role supplies detached answer anchors, while perception-prompted trajectories are optimized for answer agreement and structural validity. Both roles share parameters, allowing anchors to be regenerated from the updated model without ground-truth answer labels during reinforcement learning. Across model scales and merging methods, PASA recovers much of the post-merge perception loss while further improving reasoning. Evaluations of generated visual evidence show improved grounding, and stage-wise self-training comparisons support the benefit of asymmetric supervision. These results demonstrate how capabilities retained after merging can serve as an internal learning signal for post-merge self-alignment.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.