acceptodds
Under review as a conference paper at ICLR 2027

Learning from Recovery: Localized Episodic Trajectory Memory for Closed-Loop Multiclass Decisions under Binary Feedback

Abstract

We study multiclass decision making when an agent acts before receiving feedback and an error reveals only that the emitted action was wrong, not the hidden correct class. Our primary question is whether a persistent memory bank can turn such binary-feedback experience into better future decisions and more effective bounded retries. We introduce Localized Episodic Trajectory Memory (LETM), which stores context-conditioned winner-over-rejected relations from successful recovery trajectories such as a1- -> a2- -> a3+. The recurrent representation core is frozen; the trajectory-learning path uses no BP/BPTT, autograd-based update, or optimizer. Two evidence layers support the claim. First, mechanism controls show that disabling or scrambling persistent memory sharply degrades both first-decision and five-attempt performance, while random no-repeat retries recover only part of the closed-loop benefit. Second, in a frozen 10-seed closeout on UCI-HAR, N-MNIST, DVS-Gesture, and SHD, adding localized trajectory memory to the Full memory controller improves mean First@1 by +2.96, +2.16, +3.18, and +4.32 percentage points. The First gains are paired-significant on UCI-HAR, N-MNIST, and SHD; DVS-Gesture instead shows significant gains in five-attempt recovery and conditional-recovery AUC. Mean attempts decrease on all four tasks, and Local exceeds Full at all 20 predeclared dataset-by-exposure checkpoints. A Global-Pairwise control and representation diagnostics further show that reusable memory depends on contextual localization and retrieval geometry, not merely on storing class-pair counts.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.