acceptodds
Under review as a conference paper at ICLR 2027

CoRoute: Certified Inverse Learning from Incomplete Trajectories

Abstract

Inverse reinforcement learning from incomplete trajectories alternates between estimating missing decisions and fitting a policy. Repeating the policy fit can be costly, but reusing a stored policy may become inaccurate as the estimated trajectories change. We introduce CoRoute, a learning procedure that checks when such reuse is justified. For finite deterministic MDPs with an absorbing goal, our main result bounds the loss from reusing a policy inside a two-terminal subgraph when both its completed edge counts and incoming flow change. Using the cached flow, it checks whether reuse meets a chosen error budget without computing the new optimal policy. It allows cycles and inexact caches, and controls both directions of divergence between complete path distributions. When only incoming flow changes, a sharp bound requires just the cached mean path exposure, even if paths can be arbitrarily long. Two supporting certificates check approximate completion and the accuracy of each policy-fitting step. Numerical studies approach the sharp bounds and reduce the cost of accepting shared learning proposals. Refining the reuse bound with one potential step certifies reuse of 275 of 768 caches, up from 234. The implemented outward-rounded acceptance tests certify improvement for the stored record kernels.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.