CreditAgentBench: Measuring Agentic Credit-Model Repair under Production Constraints
Abstract
Repairing an existing credit model differs from developing a new one. A model may rank borrowers well but miscalibrate default probabilities, or improve on a development sample while weakening under temporal shift. A repair is useful only if the other operational requirements remain intact. Credit and financial-risk data make this tension concrete because ranking, calibrated probabilities, prediction error, and temporal stability matter together. We introduce CreditAgentBench, an executable benchmark of aspirational constrained model repair. Each task fixes an incumbent, an improvement goal, and protected boundaries that must each be met; a participant explores on visible data and commits one final model. Evaluation reports progress toward the goal and the margin of the least-satisfied constraint. The task bank has 57 Natural and Controlled cells from six sources. In the matched 27-task Controlled comparison within the 1,026-assignment reported visible-development set, HPO attains higher average margins, while the dynamic agent crosses both boundaries slightly more often. That joint-validity advantage does not persist when the committed models are replayed on sealed data. Five further model–provider backends change both the agent–HPO comparison and the fraction of runs that yield a scoreable model. These results show why repair progress, joint validity, and final-model availability must be reported separately.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.