acceptodds
Under review as a conference paper at ICLR 2027

LifelineQA: An Evidence-Grounded Benchmark for Situational-Awareness Question Answering in Emergency Operations

Abstract

Disasters disrupt the lifeline systems that communities depend on, including safety and security, food, water, shelter, health services, energy, communications, transportation, and hazardous-material management. During response operations, decision makers need timely and reliable answers about what is happening, which lifelines are affected, where conditions are changing, and how evidence from evolving operational reports should be interpreted. Yet existing question-answering (QA) benchmarks rarely capture this setting, where information is time-sensitive, distributed across reports, and tied to domain-specific lifeline frameworks. We introduce LifelineQA, an expert-curated, evidence-grounded benchmark for lifeline-centered QA in emergency operations. LifelineQA includes three complementary question types: situational questions about incident status, threats, and response resources; lifeline-specific questions about risks, interruptions, and impacts affecting Federal Emergency Management Agency (FEMA) Community Lifelines; and complex questions requiring synthesis across reports, evidence units, or dates, including spatiotemporal reasoning. Each answerable question is linked to verifiable source evidence. We evaluate ten large language models as QA systems across five evidence conditions to assess correctness, evidence support, abstention, and temporal robustness. Results show that models answer accurately when given relevant evidence, but retrieving that evidence and citing sufficient support remain distinct barriers. LifelineQA establishes a new evaluation setting for reliable, evidence-grounded QA in high-stakes emergency operations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.