acceptodds
Under review as a conference paper at ICLR 2027

MAPLE: Multilingual Alignment of Predictive Latent Embeddings

Abstract

Enhancing Visual-Language Models (VLMs) with external tools improves perception but incurs substantial overhead. Common solutions rely on specialised supervision; however, these are predominantly English-centric and fail when scaling to multilingual spaces. We propose MAPLE (Multilingual Alignment of Predictive Latent Embeddings), a representation-learning framework based on predictive alignment that learns from expert tool-use trajectories in the latent space, eliminating the need for explicit tool invocation at inference time. Specifically, MAPLE learns multilingual predictive embeddings, bypassing reconstruction-based methods that generate latent tokens and suffer from a training-inference mismatch, as well as limited support for multi-step tool use. By preserving the standard vision-language generation pipeline, our framework is model-agnostic and straightforward to train. Crucially, it naturally supports trajectories with multiple tool calls across languages. We evaluate MAPLE on a comprehensive suite of perception benchmarks and introduce their corresponding multilingual versions, demonstrating that it matches or outperforms standard tuning and reconstruction-based latent-reasoning baselines across multiple languages. Finally, we show that reconstruction methods learn static embeddings, demonstrating that predictive alignment is a more principled and globally scalable paradigm.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.