acceptodds
Under review as a conference paper at ICLR 2027

SUADS: Sampling-Uncertainty-Aware Data Selection for Large Language Models

Abstract

Gradient-based data selection has shown promise for data-efficient fine-tuning of large language models (LLMs). However, we find that its gains diminish or disappear on challenging coding tasks, sometimes falling below random selection. Our analysis identifies finite-reference sampling uncertainty as a key source of this unreliability: estimating the target gradient direction from a limited reference set introduces noise that scales with the incoherence of the references. Under a stylized model, we prove that this uncertainty attenuates the expected gain over random selection by a multiplicative factor determined by reference reliability, independent of the selection fraction. This degradation occurs even when population influence scores exactly capture one-step training utility, revealing a bottleneck beyond instance-level score accuracy. Guided by this analysis, we propose SUADS, a clustering-based method that combines within-cluster gradient aggregation with stratified selection to exploit more reliable local reference directions. Across Qwen3.5-4B, Llama-3.1-8B, and Gemma-2-9B on MBPP, HumanEval, and LiveCodeBench, SUADS is the only selection method whose average exceeds random selection in every setting and has the highest average pass@1 of any method under matched selection budgets. On Gemma-2-9B it matches full-data fine-tuning on average with 5% of the data.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.