PortraitGen: Anchoring GRPO with Real-Image Trajectories for Photorealistic Portrait Generation
Abstract
Reinforcement Learning like Group Relative Policy Optimization (GRPO) has significantly advanced text-to-image post-training. However, existing methods often favor superficial aesthetics, such as over-saturated colors, leaving critical flaws like AI artifacts and biological implausibilities unresolved. We attribute these limitations to two primary factors: (1) GRPO relies solely on self-sampled groups without real-image trajectory references, constraining optimization to the model's existing generative distribution; (2) existing reward signals are insufficiently sensitive to fine-grained portrait artifacts, such as anatomical inconsistencies and unnatural skin appearance. To address this, we propose PortraitGen, a novel framework tailored for photorealistic portrait generation. First, we introduce a real exemplar into each GRPO group and employ BELM inversion to construct its latent trajectory, enabling the exemplar to participate in both relative comparison and policy optimization. Second, to explicitly steer the model toward photorealism, we introduce a complementary dual-reward mechanism: OmniReward for general quality and AI-Portrait for human-centric fidelity. Furthermore, we curate PortraitBench, a comprehensive portrait-centric benchmark. Extensive experiments demonstrate that PortraitGen significantly outperforms existing baselines, effectively suppressing AI artifacts and improving portrait realism.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.