Expert-Reviewed Benchmark and Domain-Specialized LLM for Subsurface Well-Log Interpretation
Abstract
Well-log interpretation requires reasoning over long, multivariate, depth-indexed measurements. The oil and gas industry relies on it to identify hydrocarbon-bearing zones, estimate reserves, and plan well completions. Although large language models (LLMs) could accelerate this process, how well they handle it is largely unknown. We introduce WellLogBench, the first expert-reviewed benchmark for this task, built with over 1,000 hours of expert review: 1,075 validated QA pairs across six petrophysical categories and three reasoning levels, from single-formation interpretation to cross-formation and multi-well synthesis. To make raw LAS logs usable by LLMs, we develop Petrophys, a Python petrophysics library that turns million-token depth-indexed logs into compact, lithology-aware structured contexts with reproducible calculations; Dynamic Feature Selection (DFS) then cuts the context supplied to the model by 73.5%. We also present WellLog-LLM, a domain-specialized model trained with direct preference optimization (DPO) and group relative policy optimization (GRPO), then refined through inference-time Group Relative Context Optimization (GRCO). Across 15 evaluated LLMs and VLMs, GPT-5.6 Luna (high) leads the public and private splits at 70.4 and 69.3 Overall score. Our WellLog-LLM reaches 66.0 and 66.4 Overall, outperforming the 1T-parameter open-weight Kimi-K2.6 (65.0 and 65.2) despite having approximately 37× fewer parameters. Under the same DFS context, it raises the Overall score of its Qwen3.5-27B base model from 56.3 to 66.0. We will open-source the benchmark, code, and model weights, providing a reproducible foundation for grounded petrophysical reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.