Rethinking Perceptual Representations for 6-DoF Grasping with Stereo Spike Streams
Abstract
Most robotic grasping pipelines first convert sensory observations into explicit 3D representations such as point clouds, which is a computational step not found in biological intelligence. While effective in many standard settings, this reconstruct-then-act paradigm can become a bottleneck under asynchronous sensing, where high dynamic range and strict latency constraints make intermediate reconstruction costly and potential losses for downstream control. This paper explores a fundamentally different, neuro-inspired paradigm for 6-DoF grasp detection. We introduce SpikeGrasp, a framework that mimics the feedforward visuomotor pathway, processing raw, asynchronous events from stereo spike cameras, similarly to retinas, to directly infer grasp poses. Our model fuses these stereo spike streams and uses a recurrent spiking neural network, analogous to high-level visual processing, to iteratively refine grasp hypotheses without ever reconstructing the point cloud. To validate the approach, we construct a benchmark spanning clutter, textureless objects, severe illumination changes, and occlusion, and compare direct inference against reconstruct then grasp baselines. Across synthetic and real high-dynamic-range scenes, SpikeGrasp achieves stronger accuracy, data efficiency, and better robustness under challenging sensing conditions. Our results suggest that explicit point-cloud reconstruction is not universally necessary for embodied manipulation, and that direct learning in the asynchronous observation space is a promising alternative for low-latency robot perception and control, particularly in complex real environments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.