acceptodds
Under review as a conference paper at ICLR 2027

HuRF: Human-Centric Representation Learning for Generalizable Digital Human Video Forensics

Abstract

Native digital humans, whose appearance, motion, and identity are synthesized or reconstructed end-to-end, pose growing challenges to media authenticity. Existing AI-generated video and face-centric deepfake detectors often rely on generator-specific visual artifacts, temporal irregularities, or editing traces, which do not transfer reliably across heterogeneous digital human generation mechanisms. We therefore seek transferable forensic evidence from the human subject itself, which provides a shared representational basis across different generation mechanisms. Based on this insight, we propose HuRF, a human-centric representation learning framework for generalizable digital human video forensics. HuRF leverages intermediate representations from a pose-oriented human foundation model, retaining dense RGB appearance alongside fine-grained awareness of human structure. It further distills region-level semantic relations from a human parsing teacher to suppress background interference, and adaptively calibrates temporal evidence across generation mechanisms, reducing reliance on unreliable motion-naturalness cues. In this way, the detector focuses on transferable evidence from the human subject rather than generator- or background-specific shortcuts. To systematically evaluate this problem, we construct HuDigivid, containing over 35K videos from 19 known generators, together with nearly 1K collected in-the-wild digital human videos with unknown generators. Extensive experiments reveal distinct failure modes of existing detector families and show that HuRF achieves the strongest overall performance under cross-generator, cross-paradigm, and in-the-wild evaluations while remaining robust to common video degradations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.