Move, don't think: sample-efficient collision avoidance with lossy proximity sensors
Abstract
Robots must learn to avoid collisions with as few demonstrations as possible to overcome the challenge of collecting large datasets of potentially destructive collision examples. By using egocentric proximity sensor observations as inductive biases, we train a humanoid robot to learn collision avoidance in fewer steps using lossy, low-dimensional sensors with wide overlapping coverage rather than with high-resolution, spatially accurate sensing of the robot's surrounding space. We also find that, perhaps unintuitively, anticipating when and where contact happens on the body of a robot harms its learning sample efficiency. We investigate eight candidate representations for sensor observations as an inductive bias for learning collision avoidance using reinforcement learning, which include latent compression, future contact anticipation, lossy corruption, and privileged state estimation. The candidate representations were tested by training over 200 humanoid robots to play a simplified version of dodgeball end-to-end under an intentionally constrained sample and compute budget.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.