acceptodds
Under review as a conference paper at ICLR 2027

TokenFlow: Inference-Free Token Selection with Diversity Preservation for LLM Fine-Tuning

Abstract

Supervised fine-tuning (SFT) is critical for aligning large language models (LLMs), yet its efficiency is hindered by the uneven utility of tokens in instruction datasets. Existing token-level selection methods rely on reference models or auxiliary inference, introducing significant computational overhead. In this paper, we propose TokenFlow, a plug-and-play, inference-free framework for token-efficient SFT. TokenFlow exploits intrinsic signals from the standard SFT forward pass and consists of two synergistic modules that determine which tokens to optimize and how to optimize them: MINT (Maximizing INstruction-aware Token utility) and Tempered Cross-Entropy (TCE). MINT reformulates token selection as an instruction-aware subset utility maximization problem. Unlike existing methods that independently rank tokens, MINT considers token interactions and subset-level utility to select complementary token subsets under a fixed budget. TCE preserves response diversity by incorporating semantically valid alternatives from the model's predictive distribution while maintaining correctness supervision. Extensive experiments across diverse LLM families demonstrate the effectiveness of TokenFlow. Training on only 60% of the tokens, TokenFlow outperforms full-dataset fine-tuning by 2.78 points on the OpenLLM Leaderboard and consistently surpasses state-of-the-art selection baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.