ForenEvolve: Hypothesis-Driven Program Evolution for AI-Generated Image Detection
Abstract
Developing an AI-generated image detector involves coupled choices: which forensic cues to extract, how to preserve them under image processing, and how to adapt pretrained features without overfitting to known generators. Because these choices interact, a revision's value becomes clear only after testing across image sources and processing conditions, so automation requires feedback that shows where a revision helps or fails and a record that carries this evidence forward. We introduce ForenEvolve, an agentic framework for hypothesis-driven evolution of detector programs. Its search space spans three nested capability tiers: frozen forensic feature views and readouts, training-image transformations, and low-rank backbone adaptation with programmable losses. Before each experiment, a language-model agent records a hypothesis and predicted outcome; the revision is then measured across image sources and processed copies of the same images, with controlled revisions compared against their parents. Persistent experimental lineages link each outcome to its prediction, so later revisions build on accumulated evidence. The discovered detector reaches 94.34% mean balanced accuracy across 12 public benchmarks, exceeding the strongest published detector by 9.12 percentage points; it ranks first on nine, including all four excluded from search, where it reduces mean balanced error by 55% relative to the best published result. With the same backbone and training data, it exceeds standard adaptation by 4.14 points overall and outperforms three transplanted state-of-the-art recipes on every benchmark.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.