TableMoE: Reinforcement Learning to Route Connector Experts for Multimodal Table Reasoning
Abstract
Tables are not just pictures of numbers; they are structured objects where rows, headers, units, and grouping all matter. We ask whether a model can learn to combine different table views through answer-level feedback, instead of committing to one rigid representation. TableMoE keeps a pretrained visual connector and adds experts aligned to HTML, JSON, and table-redrawing code. At each visual token, the router selects two experts and mixes their outputs. A reward-guided training stage then adapts the router, experts, and language-model adapters together. Across six benchmarks, this adaptation improves average accuracy by 4.14 percentage points over dense softmax SFT, with a similar gain from the self-critical variant. Analysis of the learned mixture reveals a dominant General expert with spatially varying companions, while a fixed General+JSON pair retains most of the learned routing accuracy on the development subset. These findings support reward-based adaptation of the aligned mixture, but provide more limited evidence that spatially adaptive expert selection is necessary for accurate answers.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.