acceptodds
Under review as a conference paper at ICLR 2027

SEAICE-BENCH: A Benchmark for Sea Ice Prediction and Analysis Capability of LLMs

Abstract

Sea ice analysis requires models to connect remote-sensing imagery, long-horizon geophysical records, and specialized scientific knowledge, yet existing large language model (LLM) benchmarks do not evaluate this combination of capabilities in a sea-ice-specific setting. We introduce SEAICE-BENCH, a benchmark for assessing LLMs and multimodal LLMs on three evidence sources: native Sentinel-2 imagery, historical numerical records, and domain literature. SEAICE-BENCH comprises 136,500 multispectral satellite images across full-scene and crop-based splits, monthly sea ice extent records from November 1978 to January 2026, weekly sea ice motion vectors from November 1978 to December 2024, and a curated corpus of more than 10,000 sea-ice research articles and textbooks that yields 53,785 question-answer pairs spanning six task families and four answer formats. We evaluate ten frontier models, seven proprietary and three open-weight, in a unified zero-shot protocol on sea ice concentration estimation, next-step sea ice extent and motion vector reasoning, and literature-grounded question answering. The results reveal substantial headroom for current systems: models remain weak at processing multiple large image files to estimate sea ice concentration in specific regions and at predicting multidimensional motion vectors. These findings position SEAICE-BENCH as a realistic testbed for measuring progress toward expert-facing LLM systems for sea-ice analysis, forecasting support, and scientific literature synthesis.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.