CARE-Voice: Constraint-Aware Reliable Evaluation via Selective Decisions for Voice Design
Abstract
An open-ended voice-design brief may specify attributes to change, properties to preserve, permitted exceptions, and conflicting goals, so no single scalar score can carry the evaluation. We introduce CARE-Voice (Constraint-Aware Reliable Evaluation via Selective Decisions for Voice Design), a constraint-aware framework that compiles a brief into a Voice Constraint Graph (VCG), records typed evidence and provenance, audits tool health, and delegates the final Pass/Fail/Abstain decision to a frozen selective scorer. The evaluator may acquire evidence adaptively, but it cannot rewrite the decision contract after observing a candidate. The evidence we assemble supports a reliability claim, not a human-level perceptual claim. CARE-Voice is evaluated on complementary evidence tiers: 96 controlled counterfactual conditions with exact labels, 40 parallel semantic pairs, 90 single-attribute and 52 multi-constraint outputs from a deployed service, and existing blinded-listening studies; no new human ratings are collected, and the underpowered multi-constraint study serves only as exploratory failure analysis. Three findings summarize the evidence. First, global RMS, a common shorthand for vocal energy, agrees with the human majority on only 25% of semantic pairs, whereas typed evidence with selective refusal reaches 100% selective agreement at 42.5% coverage. Second, against strict human majorities on single-attribute pairs, CARE-Voice retains comparable agreement to a strong direct judge (84.0% selective at 75.8% coverage versus 81.8% forced), with overlapping uncertainty and therefore no superiority claim. Third, abstentions and position-inconsistency checks flag half of the pairs on which listeners fail to reach a majority, exposing ambiguity that forced labels conceal. Constraint structure prevents unsupported penalties, evidence authority blocks weak proxies from deciding alone, and selective refusal turns judge instability, identity drift, and tool failures into recorded, auditable events.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.