FOMCBench: A Human-Annotated Benchmark for Federal Reserve Q&A Responsiveness
Abstract
Cautious language can answer a question, while a fluent, on-topic response can leave it unanswered. We introduce FOMCBench, a human-annotated benchmark of answer responsiveness: 2,154 English question–answer pairs from 89 Federal Reserve press conferences (2011–2025). Three finance professionals independently annotated every pair, blind to model predictions. The release provides 6,462 judgments, fixed splits, and human-majority references for binary responsiveness and three-class evasion degree. Three-class inter-annotator agreement is 0.791 (Fleiss’ κ). Human training labels disagree with a three-model majority on 26.3% of pairs for three-class degree and 21.5% for binary responsiveness. We compare eight open-weight and three API models under pair-only and full-meeting inputs, alongside Eva-4B’s native pair-input baseline. All use common references, scoring and meeting-bootstrap uncertainty. On the primary test, meeting-minus-pair binary Macro-F1 differences range from −8.0 to +24.9 percentage points; three-class comparisons can reverse direction. The contrast does not isolate context or model openness. Versioned votes, cached predictions and executable scoring support reproducible comparisons, with explicit limits from shared meetings, subjective judgments and source-text fidelity.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.