acceptodds
Under review as a conference paper at ICLR 2027

A Framework for Copyright and Reconstruction Protection via Max Information

Abstract

We develop a formal framework for understanding and limiting the impact of sensitive data on data analysis and model training procedures. Specifically, we introduce the notion of piecewise information limited (PIL) algorithms. We focus in particular on the power of this framework as a copyright protection standard, for which it has several practical advantages over existing copyright protection frameworks. First, our framework does not require that only one (or some prespecified number of) datapoints depend on a copyrighted work, as needed in previous frameworks. Second, our framework allows us to quantify at a granular level to what degree the prompt increases the risk of a generative model regurgitating a copyrighted work. Third, our definition is strictly more general than the commonly leveraged technique of differential privacy and allows for a class of noiseless mechanisms which are not differentially private. Finally, our framework is auditable via memorization testing, thus allowing regulators to verify whether training procedures adhere to the protection standard. In this vein, we propose a new algorithm for memorization testing and empirically demonstrate its capability. In addition to copyright, our framework also provides formal protection guarantees for the two closely related domains of data reconstruction and data memorization. As such, we believe our framework may serve more generally as a way to understand responsible data usage.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.