acceptodds
Under review as a conference paper at ICLR 2027

PostTrainBench-Pro: How Well Can AI Agents Post-Train Domain-Level Capabilities under Controlled Compute?

Abstract

AI agents are increasingly used to post-train models, yet existing benchmarks optimize one benchmark per run, measure compute in wall-clock time, and detect contamination only after training. We introduce PostTrainBench-Pro, which evaluates whether agents can improve a model's capabilities in a domain under controlled compute. PostTrainBench-Pro covers three domains, Math, Code, and Tool-use, each evaluated on a benchmark suite, plus a joint track in which a single checkpoint is evaluated on all three suites. Compute is measured in held GPU-hours, which include only the time during which GPUs are allocated to a job, and the training infrastructure rejects data that overlap evaluation items and accepts only checkpoints produced by its own training jobs. We find that optimizing a single benchmark often lowers performance on the other benchmarks in the same domain. Across seven agent backends and three base models, the strongest agents match a human baseline within single domains, but on the joint track they are more than 9 points behind it, and 68% of domain runs attempted to train on evaluation items. Current agents can match human researchers on one domain, but not yet when one model must cover several.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.