When Models Prefer What Humans Strongly Reject: Revision in Open-Ended Domains
Abstract
AI systems increasingly use frontier language models both to improve an answer and to judge whether it improved. This couples the optimizer to an evaluator that may share its blind spots. We introduce SlopBench, a stress test for this failure mode in open-ended conceptual domains, where no binary verifier can determine whether a revision improved the answer. SlopBench follows 50 answers written by Opus 4.8 through sixteen rounds of recursive revision and evaluates them with five Claude-family and three GPT-family judges. After sixteen rounds, Claude judges choose the deeply revised endpoint over the initial essay in 98–100% of comparisons, while GPT judges do so in 69–86%. By contrast, humans choose the deeply revised endpoint in only 8.0% of judgments–an overwhelming divergence from frontier model preferences. SlopBench shows that whole-essay critique–revision can create a powerful improvement signal for frontier model judges even as blinded humans overwhelmingly judge the resulting writing far worse and less useful. This reversal exposes a consequential blind spot in frontier model evaluation and a concrete target for calibration.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.