Reliability-Aware Bayesian Self-Consistency for Evidence Aggregation and Risk-Controlled Abstention
Abstract
Self-consistency improves language-model reasoning by aggregating multiple sampled solutions, but standard aggregation treats observed support as a proxy for confidence without explicitly accounting for whether different evidence sources are trustworthy for a particular question. We introduce Reliability-Aware Bayesian Self-Consistency (RA-BSC), a probabilistic framework for aggregating evidence from multiple generators and verifiers while representing question-dependent uncertainty in source trust. RA-BSC combines a categorical candidate model with structured source-reliability functions and propagates parameter uncertainty through approximate Bayesian inference. We further separate posterior uncertainty from decision-time guarantees by coupling the resulting predictive and selective scores with a held-out Learn-Then-Test procedure for risk-controlled abstention. Controlled comparisons with fixed-reliability, adaptive deterministic, and conventional self-consistency baselines show that structured reliability modeling is particularly useful when evidence sources disagree, while Bayesian marginalization primarily contributes uncertainty-aware selective information rather than universally improving point prediction. Across reasoning benchmarks, our results highlight that answer aggregation, confidence calibration, and certifiable abstention are distinct objectives, and that explicitly representing uncertainty about evidence quality provides a principled way to connect them.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.