C-HAT-Bench: Benchmarking Chinese AI-Text Detection Beyond Fully Generated Text
Abstract
Large Language Models (LLMs) increasingly participate in writing by modifying or extending human drafts, causing machine involvement to vary in both form and extent. Yet most Machine-Generated Text (MGT) detectors are evaluated only on fully human-written versus fully AI-generated text. Because human–AI collaboration can weaken or redistribute cues associated with machine generation, strong performance under this binary setting may overstate detector reliability. This mismatch remains underexplored in Chinese: detection cues are shaped by tokenization and language-specific text distributions, yet controlled resources spanning production settings, domains, and generators remain limited. To fill this gap, we present a **C**hinese **H**uman-**A**I Collaborative **T**ext Detection **Bench**mark (**C-HAT-Bench**), a unified benchmark that links 5,000 human-written source texts from five domains to more than 240,000 variants produced using six generative models under *Prefix-Conditioned Continuation* as a reference setting and three collaborative production modes. We evaluate 21 detectors through four protocols spanning zero-shot and pretrained supervised document-level detection, boundary localization, and cross-condition generalization. Relative to Prefix-Conditioned Continuation, mean AUROC across document-level detectors is 12.0% lower on the collaborative production modes, with the largest detector-specific relative decrease reaching 44.4%. Transfer across collaborative production modes is also asymmetric, indicating that performance under one production setting does not reliably predict performance under another.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.