acceptodds
Under review as a conference paper at ICLR 2027

Evaluating Relational Context Compression at Realized Token Budgets

Abstract

Context compression must preserve answerable information within a token budget, yet shorter text alone cannot isolate the effect of its representation. We study Telegraph English, which rewrites retrieved passages into relational statements, using a pre-registered encoder–consumer matrix and an open compiler trained by distillation. Six of seven encoders violated the registered output schema on most or all inputs. Only 33 of 9,600 budgeted outputs fell within the registered band: at or below the requested budget and no more than two percent short. These findings motivate evaluating answer F1 at realized token ratios. The matrix was collected, but all registered confirmatory comparisons were unavailable because their comparator or calibration conditions were missing or ineligible. The remaining answer-F1 comparisons with full prose are descriptive and score construction failures as empty answers. The second compiler failed its pre-registered acceptance test against a budget-matched pruning baseline. Its out-of-distribution component was unscorable because answer keys were absent, and citation-reach judging was not reached. Context-compression comparisons require measured token budgets, explicit construction failures, and available comparator evidence.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.