DACO: DISCRIMINABILITY-AWARE CONTENT OPTIMIZATION FOR AI-GENERATED CONTENT
Abstract
Large language model (LLM) evaluators increasingly rank and score AI-generated or AI-rewritten text. When powerful generative models rewrite content at scale, different texts can become more similar in wording, organization, and presentation. We call this effect content homogenization. We ask whether homogenization changes the internal features used by an evaluator even when its accuracy remains similar, and whether features that remain important after homogenization can guide content optimization. We study these questions through representation analysis using Principal Component Analysis (PCA) components of pairwise hidden-state differences to identify features that predict evaluator preferences. We evaluate three open-weight evaluators (Gemma 4 12B, Ministral 3 14B, and Phi-4 14B) across human-preference judgments, assistant-response ranking, and hiring. We further test generalization with Llama 4 Scout (109B total, 17B active), Llama-3.1-405B-Instruct, Gemini 3.6 Flash, with an additional blind human evaluation. Homogenization changes evaluator accuracy by only about one percentage point while substantially reorganizing which PCA components predict evaluator decisions. This shows that stable accuracy can hide changes in the features associated with evaluator decisions. Held-out functional validation further shows that components that remain important after homogenization can provide useful guidance for content optimization. Based on this finding, we introduce Discriminability-Aware Content Optimization (DACO), a representation-guided content optimization method that uses validated guidance for a single guided rewrite while preserving source content. DACO has zero protected-literal failures while preserving source content. These results show that representation analysis can reveal changes in evaluator features that accuracy alone misses, while also providing a practical basis for content optimization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.