Diagnosing Faults in Reinforcement Learning Simulators with Canonical Polynomial Invariants
Abstract
A large literature incorporates physical structure into learned dynamics on the premise that models respecting the underlying physics should predict better. We test this premise using exact polynomial invariants recovered from trajectories and represented canonically as reduced Gröbner bases over ℚ. We find that exact polynomial invariants provide little additional benefit for residual-based downstream uses. On Gymnasium's Acrobot, consistency regularisation reduces algebraic residuals without improving long-horizon rollout fidelity, while one-step validation error is more strongly associated with long-horizon accuracy; similarly, shaping potentials recovered from systems with mass errors of up to 100% achieve similar learning performance to the physically correct potential. The main benefit observed here instead appears in diagnosis, where canonical representations permit exact ideal-equality and membership decisions, and where a reference generating set supplies a physical vocabulary that generic two-sample tests do not have: handed the same per-generator residuals, a Kolmogorov–Smirnov baseline localises almost as well, so the advantage comes from the generating set and not from the algebra that reads it. We develop a two-level diagnostic comprising screening, which identifies the violated physical constraint, and attribution, which recovers the faulty invariant and identifies the responsible physical parameter. Normal-form deflation removes algebraically trivial multiples, while quotient-space recovery avoids tolerance-based nullspace selection. Across fifteen injected faults, screening localises all fifteen when no observation noise is added, with no observed false alarms. Attribution identifies the responsible parameter in all seven single-parameter faults. Paired difference tests detect all faults but provide no constraint-level localisation, showing that the advantage is localisation and not detection. Finally, applying the diagnostic to 350 release pairs across eleven RL environments finds no evidence of changed simulator dynamics, but exposes a stale forward-kinematics read in Reacher's observation function, an integrator-drift ceiling on Acrobot, and a chaotic-amplification limit on comparisons in InvertedDoublePendulum. In conclusion, exactness is more useful for diagnosing violations of physical constraints than for improving prediction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.