Prism: Class-Consistent Image Generation from Tabular Data
Abstract
Many real-world datasets are tabular. While recent open-domain foundation models condition image generation primarily on text, a common strategy for tabular data is to serialize rows into natural language and condition a text-to-image diffusion model. This clubs linguistically similar classes together even when they are visually distinct, giving incorrectly conditioned images. We propose PRISM, a tri-fold model, which maps tables directly to images without using intermediate mapping to text. We train a tabular conditioning encoder that converts each attribute into a dedicated cross-attention token for the diffusion model. We then fine-tune the model toward the target class, away from confusable alternatives using a GAN-style discriminative loss on generated images. At inference, we steer towards the correct class via classifier-free guidance maximizing target class alignment without any textual input. We evaluate PRISM on three diverse applications: medical imaging, fine-grained bird classification, and plant disease recognition. Evaluations reveal improvements of up to 116% in class consistency over the best language-based baseline. On the medical dataset, a board-certified dermatologist evaluated the generated images and judged most to be visually indistinguishable from real samples. https://anonymous.4open.science/r/PRISM-8FC2
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.