OpenSystemOne: Efficient Parallel Inference for Dynamic Structured Decision Models
Abstract
Machine-facing AI workloads often require selecting, scoring, or verifying a bounded set of alternatives rather than generating natural language. Yet such workloads are commonly served by autoregressive decoders, and whether their computational structure matches the task is rarely measured at the level of the decision. We present OpenSystemOne, an open benchmarking and systems framework for parallel structured decision inference. The framework implements and compares free and constrained autoregressive generation, option likelihood scoring with and without shared-prefix reuse, encoder-based classifiers, query-conditioned decision heads, and learned dynamic-option models under a common System One interface. We measure accuracy and calibration on four public decision benchmarks and probe reordered, reduced, gold-absent and out-of-scope candidate sets in intent routing. On one Blackwell GPU, we measure latency, throughput, device memory and GPU energy on fixed request cohorts and vary context length, option cardinality and batch size one factor at a time. Eliminating answer generation does not by itself lower cost: a one-token constrained decoder already decides after one prefill, and on fixed batch-32 cohorts it completes more decisions per second than our encoder query head (342 versus 224 on BANKING77). The largest measured change was host-side: moving a per-request tokenization check out of timed decoder calls raises throughput 13-fold for requests prepared in advance, but counting that setup leaves no demonstrated first-use gain. In FP32, the precision in which reused and fresh paths agree within tolerance, prefix reuse raises multi-question throughput at higher memory cost. Quality and calibration vary by predictor and workload, and changing candidate sets expose order sensitivity and constant abstention. Characterizing when eliminating token-by-token answer generation pays off end to end, we find that for the tested implementations cost resides in the complete execution path rather than in generated tokens, while quality on the served candidate sets decides which fast path is usable.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.