GHGbench: Benchmarking Greenhouse Gas Emissions Prediction Across Companies and Buildings
Abstract
Annual greenhouse-gas (GHG) disclosures support climate accounting, yet coverage remains incomplete and reporting formats vary across companies, buildings, and jurisdictions. Prediction can support gap analysis when reports are missing or delayed. Measured accuracy depends on which entities, inputs, regions, and years appear in training and test data. Existing studies use different sources and evaluation designs, making model gains hard to compare and performance beyond the training distribution hard to assess. We introduce GHGbench, an open benchmark with separate company and building tracks and a shared evaluation framework. The company track contains 32,830 company-year records from 12,087 companies. The building track harmonizes 491,591 building-year records from 12 public programs across 26 metropolitan areas. Fixed tests cover random annual company records, unseen buildings, held-out regions and cities, later years, forecasting, and multimodal models that combine tabular inputs with satellite-image embeddings. Across both tracks, input availability and distribution shift have larger effects than small differences among leading models. Company-track R² rises from about 0.3 with basic metadata to about 0.9 with company attributes. For buildings, leading models reach about 0.4 R² on unseen buildings when all cities are represented during training; city-average R² falls to about 0.1 when each test city is held out entirely. Multimodal inputs improve average new-city performance, with gains varying across cities. GHGbench releases fixed splits, target provenance, per-record predictions, and reconstruction and validation tools.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.