Can Vision-Language Models Learn to Execute Classical Search Algorithms?
Abstract
The automated planning community has probed the planning ability of large language models by fine-tuning them on planning-problem corpora, but the training target is almost always the plan (i.e., the action sequence a planner returns), rather than the search that found it. Yet classical planning offers many search algorithms, and how well a model learns them from their search traces is still an open question. As embodied agents increasingly plan from what they see, the question extends to vision-language models (VLMs). We fine-tune a VLM on the search traces of four classical algorithms across 15 PDDL domains and study learning at two levels. For breadth-first search (BFS) and best-first width search (BFWS), whose expansion order follows from bookkeeping rules, the model executes the algorithm, writing every step of the search, including where each new node goes in the open list. For greedy and weighted- best-first search, whose order follows the additive heuristic, the model makes the search-control decision, choosing which node to expand next from images of the open nodes, with no heuristic values shown. Execution is learnable but costly: from a small training set the model solves no tasks under either algorithm, and it reaches success of 1.0 for BFS and 0.93 for BFWS only with over an order of magnitude more data. Search control is learnable from images alone: on held-out tasks, the trained VLM solves 0.28 to 0.62 of tasks within twice the reference budget, against at most 0.04 for random valid choices. Training on width-based traces even transfers beyond planning, raising logical-reasoning accuracy on FOLIO from 0.600 to 0.660. Together, these results establish search traces as a practical supervision target for planning research, alongside plans.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.