Two Asymmetries of LLM Self-Review: Family-Dependent Recall Gaps and Same-Model False-Positive Bias
Abstract
We study self-review as an empirical problem in LLM-as-a-judge settings, using a programmatic injection log as ground truth across a 4×4 writer–verifier design (400 technical reports, 286 flawed and 114 clean, 1,600 reviews) and a 9-verifier pilot (50 instances, 4,345 reviews). We report two asymmetries. First, recall gaps are model-dependent, not universal: in a logistic regression with cluster-robust standard errors clustered by instance, self-review lowers detection odds overall (log-odds -0.60, p=0.00011), but this effect is model-dependent rather than a uniform property: the GPT-family verifiers consistently show a self-review blind spot (odds ratio 0.55), whereas DeepSeek (OR 1.72) and Google (OR 1.70) verifiers are stricter on their own outputs—yet within Google the direction splits (gemini-2.5 verifiers stricter on self, gemini-3-pro-preview blind), so even vendor family is not a reliable unit. Second, error type is asymmetric: at comparable recall (self 73.9% vs. cross 72.6%, sign test p=0.70), self-review produces substantially more false positives on clean reports (FPR 16.0% vs. 0.0%, two-sided exact McNemar p=0.0078), and a same-model fresh-session condition (FPR 18.0%) is consistent with model self-similarity rather than authorship as the source of this bias. A controlled re-parsing study rules out verdict-extraction artifacts. Finally, a counterbalanced probe holds report content fixed and varies only the authorship label (self/other/none); it finds no reliable effect of the label on recall or false-positive rate (paired McNemar p=0.115 and p=0.273, respectively), indicating that the apparent self-review blind spot is a content-ownership effect rather than a label effect. These results support an empirical, non-causal conclusion: self-review reliability must be validated per model, and verifier selection should consider both missed defects and false alarms rather than treating model separation as a proxy for quality.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.