acceptodds
Under review as a conference paper at ICLR 2027

PRISM: Scope-Aware Multi-Agent Prompt Induction for Promptless Question Grounded Visual Synthesis

Abstract

Conventional Text-to-Image (T2I) synthesis and optimization assume that a visual prompt is already available. In reading comprehension (RC) illustration synthesis, however, the input consists of a passage and its questions, with no visual prompt specified. The system must first determine what evidence each illustration should depict and then induce prompts for synthesis. We formulate this prompt induction challenge as Question Grounded Visual Synthesis (QGVS). Its dual-scope constraints require question-scope evidence selection and answer secrecy for each illustration, alongside passage-scope visual consistency across the set. We introduce PRISM, a multi-agent system for RC illustration synthesis. Its scope-aware workflow jointly plans the full illustration set: question-scoped agents select evidence and construct renderer prompts, while passage-wide style planning and cross-question entity routing coordinate visual consistency. The direct mode uses frozen pretrained models. Across four RC dataset sources, direct PRISM improves visual quality over the strongest comparison (GeoMean 66.0 versus 60.6), with especially large consistency gains while preserving distinct question content. This advantage persists across the multiple tested VLM planners and T2I renderer sizes. Further experiments show that the synthesized illustrations yield larger question-answering gains for simulated developing and beginning readers. Finally, GRPO post training yields a PRISM variant with a different performance profile.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.