acceptodds
Under review as a conference paper at ICLR 2027

GenPA: Image Generators as Efficient Robot Policies

Abstract

Pretrained visual generation models have shown great promise in robotic manipulation. However, a representative line of work still treats future visual prediction as the central link between visual generative pretraining and action learning, retaining it as a training objective even when future scenes are no longer generated at inference time. This paradigm organizes policy learning around the proxy task of predicting future visual states. We argue that the instruction-following, spatial understanding, and visual imagination capabilities acquired through image generation pretraining can support control without future visual prediction: image generation models themselves can be directly adapted into powerful action learners. Meanwhile, their native image-editing interface can accommodate diverse dense perception targets, providing additional visual supervision for action learning. Building on this perspective, we propose GenPA, which directly learns actions by fine-tuning an image generation backbone together with a lightweight 19M-parameter regression adapter, while incorporating native perception co-training to strengthen control-relevant visual representations. At deployment, GenPA requires only a single backbone forward pass and action regression, without visual generation or iterative action denoising. Without additional large-scale robot-data pretraining, GenPA demonstrates strong control performance and robustness to perturbations across simulation benchmarks (99.10% success on LIBERO and 83.79% on LIBERO-Plus) and real-world manipulation tasks. Under compiled inference, GenPA takes 20.9 ms per action chunk on a single H100 GPU, achieving 1.7–3.4× speedups over strong VLA and WAM baselines. Our latency analysis further highlights architectural advantages of visual generation backbones for efficient robot control. Ablations across visual tasks and their combinations identify current-depth supervision as a simple, effective training recipe. Together, these results show that image generators can be directly adapted into strong, efficient robot policies without requiring future visual prediction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.