Green Quality Distance: Characterizing Energy and Quality Trade-Off in Open Text- to-Image and Text-to-Video Models
Abstract
Advances in generative technologies have led to the development of models capable of creating high-quality images and videos from text. In turn, the environmental impact of deployment and usage is a growing concern in research and industry. These models are commonly evaluated using qualitative metrics related to perceptual quality and prompt adherence. Recent research focused on measuring the consumed energy or the emitted during training and inference. The notion of Intelligence per watt (Saad-Falcon et al., 2026) enables a rating of Large Language Models based on their quality (e.g. intelligence) relative to the required power (e.g. watts). This metric can however not be applied readily to image and video models. We propose a novel formulation allowing to assess both the visual quality and environmental impact of generative image and video models for a given inference configuration: the Green Quality Distance. Our metric emphasizes the environmental impact of models by gathering data on energy consumption, abiotic resource usage, global warming potential, and water use. By benchmarking configurations (e.g. diffusion steps, GPU power cap) of open-source state-of-the-art models, we show that the right configuration reaches a given quality bar at a fraction of the energy cost, exploiting the full potential of existing models while reducing resource waste.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.