acceptodds
Under review as a conference paper at ICLR 2027

SAGE: Semantic Action Grounding for VLM-Guided Offline Long-Horizon Control

Abstract

Offline trajectories contain reusable behaviors for long-horizon control but lack semantic skill annotations, while off-the-shelf vision-language models (VLMs) lack grounding in environment-specific control. Existing methods often assign language to pre-discovered latent skills, which may misalign high-level planning concepts with executable behaviors. In contrast, we propose Semantic Action Grounding and Execution (SAGE), a semantic-first framework that discovers semantic skills before learning their low-level realization, thereby converting reward-free offline trajectories into a planner-readable and controller-executable interface without fine-tuning the foundation model. SAGE extracts candidate skills from sliding-window trajectory segments, consolidates them through semantic merging, relevance filtering, and data-support pruning, and grounds the resulting skill library with a skill-conditioned policy trained by transition-level imitation. A memory-enhanced VLM planner then tracks task progress and selects skills in a closed loop. Across long-horizon tasks, SAGE matches or outperforms the representative offline learning and hierarchical VLM-enhanced baselines on tasks, with relative gains of up to **53%**.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.