acceptodds
Under review as a conference paper at ICLR 2027

Trace2Prog: Compiling and Calibrating Agent Experience into Reusable Visual Programs

Abstract

In recurring visual question answering (QA), different questions may share the same solution procedure despite changes in their visual inputs. However, multimodal agents often reconstruct this procedure from scratch for every question. Reusing compiled procedures can reduce repeated processing, but a program that worked before may fail under new visual conditions. The key challenge is deciding when to use a compiled program and when to fall back to the agent. To this end, we introduce Trace2Prog, a compile–calibrate–deploy framework for recurring visual QA. It compiles past executions into candidate programs, validates them under the source visual condition, and uses a small labeled target set to determine whether they remain suitable for reuse. For each new question, Trace2Prog uses a validated program when suitable and otherwise falls back to the multimodal agent. On 480 controlled requests across four task types and three renderer profiles, Trace2Prog achieves 93.1% accuracy, compared with 48.3% for the baseline tool agent, 86.7% for per-query program synthesis, and 87.1% for success-only compilation. On 3,215 TIR-Bench and PuzzleVQA requests, Trace2Prog improves overall accuracy by 0.25–2.18 percentage points without introducing observed additional errors. Across 2,160 program–request pairs, target calibration reduces harmful replacements from 87 to 24.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.