acceptodds
Under review as a conference paper at ICLR 2027

Offline Imitation Error as a Model-Selection Proxy Under Distribution Shift: Evidence from Real-Camera Active PTZ Control

Abstract

Practitioners often screen imitation-learning policies with logged-state imitation error - per-step action error against an expert on logged data - before expensive closed-loop evaluation. We study when this proxy is valid on a real pan-tilt-zoom (PTZ) camera for indoor single-target tracking. (RQ1) Does in-distribution error preserve the ordering of held-out imitation error? Not across data-collection regimes: across three optimisation runs the held-out BC-DAgger separation persists, while in-distribution error favours behaviour-cloning (BC) policies in 178/189 pooled cross-family comparisons; the DAgger reruns share one fixed collection chain, and orderings within the BC family do not survive retraining. (RQ2) Controlled retraining shows that access to the DAgger-collected support is sufficient for most of the held-out gain, and the warm-start fine-tuning trajectory is not required to recover it: plain BC from scratch on the collected states recovers it under matched-epoch and approximately matched-update budgets, without DAgger's in-distribution degradation; we do not show that learner-induced states are uniquely responsible. (RQ3) For our geometry-defined expert, a pretrained RGB branch raises held-out error 10-17x even though the sufficient geometric inputs are given explicitly; in separate controls, constant and shuffled images stay much closer to geometry-only, consistent with the cost depending on state-aligned visual content rather than on the presence or size of the image branch, and a protocol-locked external check agrees in direction but is imprecise. Finally, a pre-specified live test of the same ten checkpoints separates selection from ranking: under free zoom the in-distribution choice has the highest mean telemetry-derived centred-tracking score and more in-frame coverage than the held-out choice in 5/5 blocks, while held-out error orders centred tracking across all ten far better (Spearman +0.71 vs. +0.01). Offline imitation error is thus not one model-selection signal: its validity depends on where it is evaluated, how the candidates acquired their training support, and whether the goal is to select one policy or rank many. All physical-system results come from one PTZ rig.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.