ProoFit: Confidential Auditing of Model Overfitting
Abstract
As machine learning models are increasingly deployed in high-stakes settings, auditing their quality becomes essential. A critical property of interest is whether the model is overfitted. Yet direct audits are often impossible: model owners cannot reveal model weights without risking exposure of proprietary IP. While zero-knowledge proofs, a cryptographic primitive, offer a compelling theoretical solution, naive approaches require proving the correctness of entire training trajectories–an operation that is prohibitively expensive for modern models. In this paper, we introduce ProoFit, a practical system for auditing overfitting without violating model confidentiality. Rather than proving properties of the training process, ProoFit employs novel tests that depend only on trained model weights, leveraging theoretical connections between generalization and model complexity. Our experiments demonstrate that ProoFit significantly outperforms baseline cryptographic approaches, achieving a to improvement in runtime.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.