DECODE: Decomposable Environments for COntrolled Sequential-Decision Evaluation
Abstract
Benchmarks for sequential decision systems usually report a final score, which shows that a solver fell short of the optimum but not why. DECODE is a generator of partially observed sequential decision environments whose equations are known. This allows the gap between a solver and the exact optimum to be split into components, each measured against a reference computed from those equations in closed form, analytically, or by Monte Carlo with certified precision. The components are the planning horizon, knowledge of the true action response, exact inference over the hidden state, memory, re-observation before each decision, and the inflation of the solver's own value estimate. The generator is instantiated as a consumer-credit collections portfolio and released as five environments that share the same randomly drawn dynamics matrices and differ in a memory parameter, which sets how much an action's effect depends on the client's history, and a noise parameter. In the deterministic environment, with both parameters at zero, where scores are percentages of the exact optimum, planning six months ahead instead of one raises the same solver's score by 15 to 17 percentage points, while switching among the tested solver classes at a fixed horizon changes it by at most 4. The best solver trained on the logs ends 15 points below the optimum. Giving it the true action response together with exact inference recovers 9 of these points, and within that disclosure exact inference is worth nothing measurable. The missing information is linearly recoverable from the logged observables, so the binding constraint is the lack of exploration in the logged data, which none of the tested estimator classes removes. On the same environment, a planner acting on expected future states overstates the value of its own plan even with exact beliefs, so self-reported values cannot replace external evaluation. Re-observing before each decision is worth almost nothing in the deterministic environment and gains value as process noise grows. When the memory parameter is low, all tested recurrent configurations perform the same, so such environments cannot rank memory architectures. A statistic that needs no ground truth relates the released environments to public Fannie Mae loan-performance data. The generator, the five environments, and all artifacts behind the reported values are released with the paper.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.