acceptodds
Under review as a conference paper at ICLR 2027

Bag Design for Model Evaluation from Aggregate Labels

Abstract

Aggregate labels provide responses for groups of instances, or bags, when individual outcomes are unavailable. Existing bag-curation methods target gradient recovery or prediction quality, while evaluating a fixed model requires understanding how grouping affects uncertainty about its ranking performance. Identical group means can support different values of the area under the receiver operating characteristic curve (AUC). We formulate bag design before feedback as minimizing the range of AUC values compatible with exact population group means. When each instance has its own response probability, we derive an exact width formula that sums contributions from score ranks within bags. For distinct scores and equal-capacity bags, adjacent-rank grouping minimizes worst-case width at a fixed overall response probability, provided its product with bag size is an integer. When instances with identical observed features share a response probability, their occurrences across bags couple the constraints. Using specified reference probabilities, we compute the coupled AUC endpoints by linear programming and search for width-reducing swaps that preserve bag sizes. Keeping each repeated state within one bag recovers the sum obtained by treating instance probabilities separately. On twelve-instance blocks, width-minimizing assignments that spread repeated states across bags reduce the median coupled-to-separable width ratio to 0.104, demonstrating that evaluation-width minimization and training-oriented curation can select different bags. In the separate-probability setting, adjacent-rank grouping yields narrower intervals than balanced random assignment across click-log, census and network-traffic workloads. On the click log at a fixed feedback budget, its interval width is 2.4 × 10⁻⁴ versus a balanced-random median of 0.992. These results support designing aggregate-label groups around the performance measure to be evaluated.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.