VASP: Visual-anchored Agent for Scientific Poster Generation
Abstract
Scientific poster generation condenses a paper into a single page, where visual evidence must clearly communicate its central claims. Existing approaches exhibit two mismatches: a layout mismatch, in which text arrangements or templates constrain visual evidence regardless of its scientific importance, and an execution mismatch, in which models estimate geometry that programs can compute directly. Consequently, key evidence can be compressed or incompletely extracted, while refinement focused on space utilization leaves its association with the claims it supports unchecked. We address both mismatches with VASP, a visual-first framework that combines semantic reasoning with precise execution. Vision-Centric Anchoring makes key visual evidence the basis of layout construction, with its scientific role guiding spatial allocation and the arrangement of related claims. Semantic-Geometric Decoupling assigns scientific interpretation and design decisions to LLMs, while programs compute figure boundaries and check spatial constraints derived from explicitly declared claim–evidence relations. This division supports intact evidence extraction and iterative refinement guided by the intended scientific organization. We further introduce two criterion-level benchmarks that assess visual presentation and author-informed key-information coverage. On the 100-paper Paper2Poster evaluation set, VASP achieves the highest visual-presentation score (94.3 vs. 88.1 for the strongest existing approach) and Paper2Poster VLM-judge score (3.71 vs. 3.52). It also achieves 94.3 mAP in figure extraction, compared with 84.1 for ResearchStudio's agent loop, with complete crops for all 100 sampled figures, and preserves more author-informed key information than existing fixed-canvas approaches.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.