acceptodds
Under review as a conference paper at ICLR 2027

SCOPE: Selective Cross-modal Orchestration of Visual Perception Experts

Abstract

Multi-encoder vision-language models (VLMs) benefit from complementary visual representations, but activating every encoder increases inference cost and visual context length. We propose SCOPE, a Mixture-of-Encoders framework that keeps one shared encoder and selects one auxiliary encoder for each image–prompt pair. A lightweight router uses cross-attention between the text prompt and shared visual features, and is trained with dual entropy regularization and auxiliary losses to encourage confident instance-level decisions and diverse routing across samples. On the document-centric and OCR-heavy benchmarks studied here, SCOPE improves over the reported static single-auxiliary baselines and naive all-encoder concatenation on average. Relative to activating all auxiliary encoders, the tabulated vision-encoder and decoder FLOPs estimates decrease by 24–49%; text-conditioned routing also requires a separate text-encoder pass. These results support instance-adaptive selection under a one-auxiliary-stream deployment constraint, where the selected encoder determines the auxiliary token count.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.