IGOA: Online Alignment for Unified-Canvas Reasoning via Information Geometry
Abstract
Multimodal Large Language Models (MLLMs) have excelled at processing multiple images as sequential inputs. However, their performance drops when textual instructions and one or more visual instances are jointly rasterized within one canvas, which we define as the Picture-in-Picture (PiP) Gap. We introduce a PiP evaluation framework with a flagship challenge and a paired diagnostic: PiP-All places the query and indexed visual instances in one image, whereas PiP-Visual preserves image order and source pixels at the benchmark-rendering stage, omits the query from the canvas, and supplies it only as language tokens. On three representative models, 96.1%–99.2% of the total gap remains under PiP-Visual, showing that query rasterization alone cannot explain the degradation. We measure this gap end-to-end through each model's released visual processor and inference pipeline. PiP further groups questions by their dominant requirement: subregion referencing or cross-region compositional reasoning. Treating this as a learnable gap, we propose Information-Geometric Online Alignment (IGOA) for multi-context RLVR. IGOA coordinates Original-Interleaved and Unified-Canvas contexts using a scalable diagonal active-Fisher proxy and coordinate-wise log-volume expansion over accumulated Fisher memory. These scores drive a mirror-descent update that prioritizes non-redundant context contributions. Experiments across 13 models show a consistent PiP bottleneck, while IGOA improves PiP performance across four backbone groups and also improves performance on five standard multimodal benchmarks. *Code is available at https://anonymous.4open.science/r/IGOA_PiP-EA76*.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.