acceptodds
Under review as a conference paper at ICLR 2027

DexHoldem: Playing Texas Hold'em with Dexterous Embodied System

Abstract

Evaluating embodied systems with real dexterous hardware requires more than isolated motor-skill tests: an agent must perceive a changing scene (e.g. a tabletop), choose a context-appropriate action, execute it with a dexterous hand, and leave the scene usable for later decisions. We introduce **DexHoldem**, a real-world benchmark evaluating Texas Hold'em related dexterous manipulations with a ShadowHand. DexHoldem provides 1,470 teleoperated demonstrations across 14 Texas Hold'em manipulation primitives, a standardized physical policy benchmark, and an agentic perception benchmark that tests whether agents can recover the structured game state needed for embodied decision making. On primitive execution, obtains the highest task completion rate (%), while and tie on scene-preserving success rate (%). On agentic perception, Opus 5.5 narrowly leads on both strict problem-level accuracy (%) and average field-wise accuracy (%); the gap between the two exposes the distance between isolated visual sub-capabilities and complete routing-relevant state recovery. Finally, we instantiate the full embodied-agent loop with one agent–policy pairing over 33 closed-loop hand-level rollouts, in which only % of hands complete; retries restore the failed primitive in 12 of 34 dispatches and resolve prolonged execution stalls in three of the four completed hands, which would otherwise have required manual termination. Only one hand completes with neither a retry nor a human-help request. DexHoldem therefore evaluates dexterous tabletop execution, agentic perception, and embodied decision routing in a shared physical setting. Anonymous Website at https://dexholdempage.github.io/DexHoldem.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.