R²D-Grasp: Reward-Refined Diffusion for Physically Reliable Dexterous Grasping
Abstract
Existing dexterous grasp diffusion models can learn plausible and diverse grasp distributions from pose supervision, but they struggle to distinguish subtle pose variations that lead to different physical contact outcomes. We propose Physics-Selective DDPO, which first pretrains a grasp DDPM and then generates local candidates by branching from shared intermediate denoising states. Simulation outcomes are used to construct group-relative advantages and optimize physically sensitive denoising decisions, while the remaining stages preserve the pretrained generative behavior to improve grasp success without sacrificing grasp priors or diversity. Experiments on MultiDex and RealDex demonstrate improved grasp success and cross-object generalization.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.