Conv-to-Bench: Exploring Automated Benchmark Construction from Conversational Traces
Abstract
The continuous advancement of large language models requires robust, scalable benchmarking. Traditional benchmark creation relies on expensive expert curation, while prior automated methods focus primarily on single-turn interactions and rely heavily on proprietary models. We introduce Conv-to-Bench, a multi-stage framework that automatically transforms authentic human–LLM dialogues into self-contained instructions paired with fine-grained, verifiable requirement checklists using open-weight models for construction and judging. Moving beyond single-turn queries, our pipeline captures multi-turn user refinements to formulate realistic evaluation criteria. Across an aligned 19-model roster evaluated against Arena AI human preferences, Conv-to-Bench achieves comparable rank correlation with human judgments and Brier scores relative to Arena-Hard-Auto, a widely used automated benchmark (– vs.; Brier – vs.), without proprietary API reliance. Ablation experiments demonstrate that atomic checklist verification significantly outperforms pairwise comparisons across Spearman correlation (), separability ( pp), agreement ( pp), and Brier score (). Incorporating multi-turn refinements yields mixed results: Spearman correlation point estimates improve across all LLM judges, but significant gains in correlation and Brier score occur for only one individual judge. Conv-to-Bench builds a 500-instruction benchmark for USD 19.05 and evaluates candidate models for USD 1.62 each on hosted infrastructure. Together, these results show that open-weight models can build benchmarks that recover human preference rankings from conversational data, enabling their application in lower-budget settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.