SCM-Bench: An Auditable Benchmark for Context Efficiency in Multi-Turn Dialogue
Abstract
Context efficiency in multi-turn dialogue concerns the answer produced from a particular history, not only the amount of history retained. SCM-Bench defines a versioned target-turn task and an auditable framework for studying history retention, selection, compression, and model routing. Deployment decisions use the current request and strictly preceding legal history; evaluation uses frozen requirements, anonymous answers, and complete legal history. The framework distinguishes reference-history budgets from provider-native generation usage and shared-matrix construction from policy deployment. Its historical profile contains 3,893 deployment-eligible inputs from 261 sessions, but candidate selection is not a completed formal evaluation. Eight context interfaces and independent versus selected-context routing support controlled comparisons without implying learned joint optimization. The proposed analysis preserves failures, missing outcomes, and constraint applicability. Available development re-evaluations do not supply a qualified quality–efficiency release.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.