acceptodds
Under review as a conference paper at ICLR 2027

AffiMoE: Fine-Grained Text-Visual Alignment via Token-Wise Expert Routing for Language-Grounded Affordance Reasoning

Abstract

Affordance grounding aims to localize object regions that enable a specified action, requiring models to connect high-level task intent with fine-grained functional structures. Recent LLM- and VLM-based approaches improve open-world affordance reasoning through language supervision, while RAGNet further extends this paradigm with human-like instructions and reasoning-aware segmentation. However, these methods still compress rich multimodal semantics into compact affordance representations, leaving the correspondence between language-level actions and fine-grained functional regions largely implicit. To address this limitation, we propose AffiMoE, a language-grounded affordance reasoning framework that explicitly strengthens fine-grained text-visual alignment through token-wise expert routing. Specifically, AffiMoE uses cross-attention to align action semantics with dense visual features, producing localized representations that better capture action-relevant object parts. Since precise affordance grounding requires both global object context and local functional evidence, we further treat globally aligned multimodal features and locally aligned visual features as complementary experts, and introduce a token-wise mixture-of-experts router to adaptively integrate them at each spatial location. In this way, AffiMoE preserves coherent high-level reasoning while improving the localization of action-relevant functional regions. To rigorously evaluate this capability, we introduce a new benchmark, and extensive experiments show that AffiMoE achieves state-of-the-art performance across multiple benchmarks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.