FigStruct: Object Hierarchies in the Loop for Agentic Scientific Figure Redesign
Abstract
Improving the visual quality of scientific figures often requires time-consuming manual editing. Researchers must adjust layouts, labels, and connections while preserving the underlying scientific meaning. Image generation can improve presentation, but further revisions require an editable representation of the result. We introduce FigStruct, an agentic framework that couples visual redesign with object-hierarchy-guided SVG correction. A local open-source controller analyzes the input, retrieves design references, and plans a revised composition to guide an image generator. The generated design is converted to SVG and organized into an object hierarchy. Enclosing groups provide diagnostic context, while object bindings identify the SVG elements selected for revision. These bindings are updated after structural changes to support subsequent corrections. Correction rounds run locally, followed by one paid model call for final global refinement. On 150 scientific figures, our best redesign configuration achieves the highest mean scores across all five model-judged visual-quality dimensions among the compared methods. With the generator fixed, agentic guidance improves all five scores for GPT Image 2.0 and four for Qwen Image 2.0 Pro over direct prompting. Reconstructed SVGs attain 99.39% semantic text preservation against the redesigned image and 96.34% strict native-text editability. On 50 complex figures, manual post-editing to a common design target averages 9 minutes 15 seconds for FigStruct outputs, compared with 45 minutes 19 seconds for AutoFigure-Edit outputs. Including retries, mean paid API cost for SVG reconstruction and refinement is $0.009429 per output. This cost is 96.78% lower than that of AutoFigure-Edit and 99.61% lower than that of CraftEditor.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.