TRACE: Towards Interpretable Carbon Emission Estimation via Reasoning-to-Prediction from Satellite Imagery
Abstract
Accurate monitoring of urban carbon emissions is essential for climate mitigation, yet satellite-based deep models remain difficult to interpret: they predict an emission value without revealing which scene properties support it. Vision-Language Models (VLMs) have the potential to bridge this semantic gap; however, applying them directly to carbon-emission estimation remains challenging because they are prone to visual hallucinations and lack mechanisms tailored to continuous regression. We propose TRACE (Textual Reasoning-Augmented Carbon Estimation), a textual reasoning-augmented framework compatible with different visual encoders in which an Analyst VLM generates image-grounded environmental analyses and a downstream Predictor maps visual and textual features to an emission estimate. In the absence of ground-truth analytical annotations, teacher-generated pseudo-analyses are used to initialize the Predictor, followed by alternating Analyst–Predictor optimization through Periodical Evolution. Experiments using data from five cities show improvements over image-only baselines across four visual encoders. TRACE achieves relative R² gains of 3.6%–6.8%, while blind human evaluation shows improved factual grounding and carbon relevance over unfine-tuned VLM analyses.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.