acceptodds
Under review as a conference paper at ICLR 2027

: Acting on Evidence Streams, Evolving through Disaster Event Streams

Abstract

Disaster-response agents revise their harnesses to improve later decisions by adapting programs and configurations governing evidence and tool use. These updates use feedback from comparing forecasts with later observations. However, forecasts may share inputs or be checked against the same observations. Treating their feedback as independent can overstate support for retaining updates. Existing benchmarks do not trace how these dependencies affect updates and later performance. We introduce Streaming², a benchmark annotating forecast lineage and measuring how updates affect later tasks. Its 1,000 events contain 10,000 action-selection and spatial-analysis tasks, with evidence arriving within events and harnesses evolving across them. Validation experiments show that a baseline agent makes fewer harmful updates when given these annotations. However, some updates still improve performance on past tasks but not on future events. These findings motivate FLARE, which compares content along forecast lineage to generate complete updates and predict their future gains. It learns from measured performance differences between updated and unchanged harnesses on the same later training events. At test time, learned parameters remain frozen while predicted gains guide retention. On Streaming²-Bench, FLARE raises the overall task score by 21.0% relative to the fixed harness and 4.4% relative to MemEvolve, the strongest self-evolving baseline.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.