FlowRSI: Benchmarking Recursive Self-Improvement in Flowchart Parsing
Abstract
Improving after feedback is not the same as learning how to improve: an agent may fix serialization without recovering more visual structure, overwrite a useful strategy, or fail to transfer it to new appearances. We introduce FlowRSI, an offline benchmark of frozen-weight, agent-directed strategy revision in flowchart parsing. Agents inspect labeled examples, receive feedback from two successive mock exams, and commit prompt, memory, or skill states; we independently replay all four states on the same hidden 1,348-image test. Exact graph supervision, strict and content-aware scores, statewise replays, artifact interventions, and clean, corrupted, and unseen-writer strata distinguish feedback attribution, revision stability, and visual transfer. Across nine open systems, final strict relation-set F1 improves for four and declines for five. Qwen3.5-27B gains 55.47 points, yet its content F1 gains only 6.88 while valid output rises from 6.08% to 77.30%: much of the measured gain is contract repair. Qwen3.5-9B instead loses 25.78 points in its last revision, and removing its final artifacts restores 27.48 points in a matched replay. Even the improved 27B state scores 65.87 on held-out scan corruption but only 26.06 on unseen handwriting. The central limitation is thus not merely whether agents can change themselves, but whether they can diagnose the right failure, retain prior competence, and transfer a correction. FlowRSI turns these mechanisms into observable outcomes and suggests that future multimodal self-improvement should couple error attribution with revision safeguards and shift-aware validation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.