acceptodds
Under review as a conference paper at ICLR 2027

DVSD: Dual-View Self-Distillation for Long-Context Evidence Compression and Answering

Abstract

Long-context reasoning tasks such as question answering and agent memory suffer from poor effective utilization of lengthy inputs, despite continuous expansions to model context windows. A key capability toward solving this problem is enabling models to autonomously compress long contexts into task-relevant compact evidence and leverage the compressed information for downstream tasks. We propose Dual-View Self-Distillation (DVSD), a training framework that jointly optimizes evidence compression and answer generation. A shared self-teacher provides supervision through two complementary views: an answer-privileged view uses reference answers to guide evidence construction without requiring evidence annotations, while an evidence-conditioned view forms answer targets from the question and student-generated evidence, without access to the original context or reference answer. Both views deliver dense token-level guidance along the student's on-policy trajectories, jointly training evidence construction and answer generation. Across three backbone models, DVSD achieves the highest overall scores among all baselines on LongBench v2, RULER, and InfiniteBench. It also generalizes beyond training lengths, improves agent-memory retrieval and test-time learning, and largely preserves the general capabilities.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.