SPARK: Self-Correcting Embodied Memory via Physical Re-Observation
Abstract
Embodied retrieval-augmented generation (RAG) systems build semantic knowledge bases (KBs) from robot observations and answer language queries over them. Most retain and extend the first episode’s KB, even though perception errors can introduce incorrect labels, merged instances, and missing regions from the outset. We present SPARK, a corrective multi-episode framework for embodied RAG. After the first episode, a semantic critic flags implausible and under-covered regions without ground truth, a targeted planner generates revisit viewpoints, and new observations are incorporated through a selective merge policy that can add, remove, or replace information. We evaluate SPARK on four outdoor AirSim maps under two viewpoint degradations using a perception-built KB, a visibility-matched development-time oracle, and a fixed EQA set scored before and after correction. One corrective episode improves downstream question answering on every map under both degradations while storing 19% fewer objects than append-only updating, with gains up to 2.6× on a single map and condition. Extending correction through Episode 4 shows that the resulting KB can exceed the EQA performance of a single undegraded collection pass. Ablations show contributions from the critic, targeted revisits, and selective merge. SPARK shows that physical re-observation can correct persistent robotic memory by identifying weak regions, collecting new evidence, and selectively updating the KB.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.