acceptodds
Under review as a conference paper at ICLR 2027

SystemOneBench: A Cross-Domain Benchmark for Calibrated Machine-Native Intelligence

Abstract

Machine intelligence is increasingly embedded in software that requires structured decisions rather than open-ended text generation. Existing evaluations mostly measure generative quality, reasoning accuracy or task completion, and give limited evidence about the calibration, efficiency and reliability of models used as machine-consumed decision functions. System One interfaces, now offered as hosted services, serve this use by returning a probability distribution over a supplied answer space; this guarantees that every answer is well formed, not that it is informative or calibrated. We introduce SystemOneBench, a cross-domain benchmark protocol for calibrated structured decision models, instantiated on seven data families spanning financial market microstructure, blockchain transaction dynamics, road traffic, Earth-system observations, maritime operations and autonomous-agent traces. Each task maps a real-world state and a dynamically defined answer space to typed probabilistic decisions whose ground truth is derived from subsequently observed outcomes rather than model-generated labels. The protocol specifies paired, dependence-aware analyses of decision accuracy, proper scoring rules, calibration error, risk–coverage trade-offs, temporal and entity shift and option-set generalization, and a procedure for measuring latency, throughput and energy. Its harness presents free generation, constrained decoding and one-pass option scoring of one Qwen3-0.6B checkpoint, supervised references including learned decision heads, and a commercial System One model with identical states, answer spaces and outcomes. First results show that validity and information separate. Free and constrained generation return a valid distribution on every attempt, yet reproduce the prompt's uniform format example on all 15,491 outputs per readout; one-pass scoring never does. The hosted System One model is schema-valid on every agent-trace request, but its test log loss (0.7986 nats raw, 0.7747 after temperature scaling) exceeds that of a uniform prediction (0.6931); on maritime and weather tasks it assigns zero probability to eleven observed calibration events per family, leaving temperature scaling without a finite-loss solution. Machine-native decisions must be evaluated as probabilities on outcomes, not as outputs that fit a schema.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.