Pairformer Is an Effective Relational Embodied Learner
Abstract
Transformers thrive on large-scale data, yet robot experience remains costly and sparsely covers physical interactions. We investigate a relational inductive bias for embodied learning: maintaining and composing vector-valued representations of token pairs. We study Pairformer from AlphaFold 3, whose persistent pair states and triangle updates guide token attention while retaining a Transformer-compatible token interface. Diagnostic experiments show that Pairformer retains graph relations for subsequent queries and predicts contact-driven motion more accurately, with attention more closely aligned to collisions in a visual probe. We then test this computation in three embodied components: (i) observation encoders for diffusion policies, (ii) latent dynamics models for model-predictive control, and (iii) pretrained action predictors within . On seven simulated manipulation tasks with 200 demonstrations per task, the encoder and dynamics adaptations improve mean success by 6.9 and 2.85 percentage points over their Transformer baselines. The action-predictor adaptation improves atomic-seen success by 4.23 points, with smaller gains on composite tasks. Together, these results support explicit relational computation as an effective architectural complement to data scaling and pretraining for embodied learning and control.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.