acceptodds
Under review as a conference paper at ICLR 2027

Can Generative Image Models Reason Through Each Other’s Outputs?

Abstract

Can separately trained generative image models use each other’s outputs for further reasoning? We study this question with models trained from scratch to perform addition modulo 23 by taking images of two operands as input and generating an image of the answer. We split training, development, and test sets between the complete set of 276 possible unordered modulo-23 additions, then train five models with different seeds and select the three that converge. Those three image-to-image transformers trained from scratch generalize to correctly solve 89.1% of held-out addition problems on average. We then ask whether an answer generated by one model can be used by a separately trained model to solve another held-out addition problem. Among cases where both models can solve their respective arithmetic steps independently, passing the first model’s generated answer to the second attains the correct result in 210 of 215 cases. In the remaining five, the generated number is still read correctly but changes the second model’s answer. This suggests that chains of generative image models, by passing images between steps, can reason and solve multi-step arithmetic addition problems.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.