acceptodds
Under review as a conference paper at ICLR 2027

Physis-Embed: Unified VLA Embeddings for Robot In-Context Learning and Evolution

Abstract

Retrieving context for a robot requires matching both its task and its stage of execution. We introduce Physis-Embed, to our knowledge the first vision-language-action (VLA) embedding model to jointly encode all three modalities in both current-state queries and recorded future continuations. Predictive training aligns these endpoints for physical-experience retrieval; at inference, the model retrieves historical continuations as context. We construct multi-source training data and introduce Physis-Embed-Bench, on which our model achieves 47.33% clip-level and 81.77% episode-level R@1, surpassing action-free Gemini Embedding 2 by 22.09 and 33.47 percentage points. In our retrieval studies, removing native action degrades retrieval more than removing vision. This retrieval space supports two applications. Embed2ICL uses retrieved experience for training-free and training-based policy conditioning, increasing success from 47.0% to 56.3% in DROID-Sim and from 86.4% to 92.8% on the LIBERO difficult subset. Both interfaces also improve task progress in real-robot evaluations. Embed2Evolve uses execution feedback to grow task-local memory and refine its use around a frozen GPT-6 Astra policy. Across five ManiSkill tasks, it achieves 68% success, compared with 24% for GPT-6 Astra Direct and 33% for a harness independently evolved with Omni-Embed-Nemotron-3B under the same protocol. Relative to the latter, it uses approximately 79.5% fewer policy tokens per test rollout.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.