acceptodds
Under review as a conference paper at ICLR 2027

Measuring Counterfactual Credit in a Replayable Wildfire-Suppression Benchmark

Abstract

Predicting a team's return and assigning credit to individual actions are distinct tasks. FLARE, a stochastic wildfire-suppression benchmark, uses coupled-draw replay to measure immediate placement effects for training and long-horizon action contrasts for grading critics. In an eight-seed MAPPO comparison, simulator credit reduces interquartile mean burned area from 26.6% with shared reward to 20.9%. A fixed non-replay local reward scores 24.7%; simulator credit wins every paired seed against both. Informed learned credit scores 27.0% and meets a prospectively fixed -point equivalence criterion against shared reward. Four informed critics paired with responsive main-comparison actors ( saturated audited decisions) attain factual-return , yet raw credit RMSE is times zero's. Matching critic credit to the finite replay reference's RMS still leaves squared error at times zero's; ten earlier-rule critics show . Held-out MSE fitting instead attenuates current predictions by over 97%, approaching zero-credit error. Current critics also give heterogeneous agents nearly identical action rankings (median Spearman ), with largest-magnitude credits on below-1%-probability actions in rows. These measured structures complement the finite-reference calibration diagnostic. Coupled replay separates feedback performance, numerical scale and action-level structure while quantifying replay precision.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.