Learning from Physical Experience:Context-Space Adaptation forRobotic Manipulation
Abstract
Recent multimodal foundation models exhibit strong perception, reasoning, and in-context learning capabilities, but translating these capabilities into physical action remains challenging. Robotic manipulation demands four critical capabilities: spatial reasoning, dexterous interaction, tool use, and adaptation to execution failures. Through real-world experiments across diverse manipulation tasks, we find that ChatGPT-6-Astra, a state-of-the-art general-purpose reasoner, can serve as a physically actionable, context-responsive manipulation policy, demonstrating effective fine manipulation and spatial reasoning. Building on these findings, we introduce a recursive self-improvement harness for learning from physical experience. A context compiler integrates current observations, system-level guidance, task instructions, and retrieved experience to generate subtask sequences and executable motion plans. Plans, execution results, and task outcomes are accumulated as experience, with optional human feedback providing additional guidance. Reflection on this evidence produces context updates that inform subsequent planning and execution, forming a recursive loop of interaction, experience accumulation, and context adaptation without task-specific parameter updates. Our experiments show that learning from experience enables manipulation adaptation in context space. On challenging tool-use and potato-chip picking tasks where direct reasoning alone yields zero success, successive experience-driven updates support progressive skill consolidation and improved execution. These findings connect general-purpose reasoning with physical skill acquisition and demonstrate how accumulated interaction experience can improve robotic manipulation while keeping model parameters fixed.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.