CraftBench: Benchmarking LLM-Generated Programs for Feature-Preserving Scientific Compression
Abstract
The massive volumes of data produced by scientific computing have driven widespread adoption of error-bounded lossy compression, which can achieve high compression ratios while bounding reconstruction error at every data point. However, many downstream analyses depend on derived quantities and features, such as gradients, distributions, spectra, and topological structures, that may be substantially distorted even when the pointwise error bound is satisfied. Existing feature-preserving methods are typically designed for a specific quantity or feature, or a narrow class of them, requiring substantial redesign when the target changes. Recent advances in LLM-based code generation raise the possibility of automating the adaptation of existing compressors to preserve diverse scientific features. We introduce CraftBench, a benchmark for evaluating LLM-generated adaptations of existing compressors using executable feature definitions and task-specific distance functions. CraftBench comprises 83 feature-preservation tasks spanning numerical, spatial, geometric, spectral, composite, and domain-specific analyses. Each task is evaluated under three adaptation modes: error bounding (EB), configuration selection (CS), and error correction (EC). An adaptation is considered successful only if it satisfies four criteria: (1) it reduces the distance of the target derived quantity or feature by at least 50% relative to the base reconstruction; (2) it respects the prescribed pointwise error bound; (3) it is not dominated in feature distance and compressed size by measured uniform-error-bound references, constructed by recompressing the data with progressively tighter global error bounds; and (4) its compressed output remains smaller than a fixed lossless reference. Among 2,475 LLM-generated adaptations, 282 (11.4%) satisfy all four criteria. An additional 509 satisfy the feature-distance reduction, pointwise-error, and lossless-storage requirements but are dominated by the uniform-error-bound references. These results reveal a substantial gap between generating adaptations that improve preservation of the target quantity or feature and producing adaptations that outperform the simple baseline of uniformly tightening the global error bound.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.