EnvAudit: Auditing the Training Utility of Synthetic Agent Environments under a Fixed Budget
Abstract
Can the target utility of synthetic training be bounded before running the training itself? We give an explicit end-to-end certificate for a finite learner and a class of state-dependent target tasks. The audited quantity is the improvement of the model returned under a resource budget, including base-model fallback on noncompletion. A static enclosure of the future policy combines with local target-transition constraints to give closed-form gain bounds, with attainable endpoints in a fixed-parameter subfamily. Preserving the same unknown environment across policies then yields completion-probability thresholds for choosing between training sources. For two fixed synthetic sources and a prescribed finite learner, the certificates establish opposite rankings at two budgets for every task horizon. At horizon 16, both rankings are determined even though independently bounded utilities overlap. The same task class also exposes a genuine information limit: two legal three-step targets satisfy the same supplied constraints yet give opposite rankings at one fixed budget for the same finite training endpoints. Together, these results specify when structural information certifies a useful decision and demonstrate why it cannot always do so. No future optimizer iterates or target samples are used. The guarantees concern the declared learner, progress-based success rule, and cost contract; they do not establish general neural fine-tuning prediction or empirical savings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.