acceptodds
Under review as a conference paper at ICLR 2027

SynKV: Rethinking Attention Structure in LLMs for Long-Context Code Generation

Abstract

As LLM-based code generation scales to repository-level inference, context lengths routinely exceed 100K tokens, sharply intensifying GPU memory pressure. Existing KV-cache compression methods remain limited in this setting: token-level selection struggles under aggressive KV budgets, while fixed-size grouping can break coherent code structures and substantially degrade generation performance. Program syntax provides a structural perspective, organizing code into locally coherent blocks. This raises a natural question: **does code-LLM attention mirror this block organization for KV retention?** Query-key interactions reveal syntax-aligned blocks whose recurring properties provide actionable selection signals. Motivated by this finding, we develop SynKV, a syntax-aware block-level framework for KV retention in long-context code generation that alleviates inference memory pressure while preserving generation quality. SynKV restructures flat token sequences into syntax-aligned blocks, then combines hierarchical dependency retrieval with lightweight summary-based relevance estimation to retain critical KV entries. Across four diverse code-generation benchmarks and six complementary metrics, SynKV advances the frontier in both memory scalability and generation quality. Notably, SynKV exhibits pronounced advantages under extreme long-context inference: (i) at the 1M-token scale, SynKV cuts the device-memory footprint from TB scale to tens of GB with only 0.005% KV retention, while existing systems run out of memory; and (ii) SynKV consistently sustains generation quality under aggressive compression, outperforming the best-performing method by over 40%. Beyond SynKV, our findings open a new direction in structured-attention design for scaling long-context LLM inference, balancing memory efficiency and generation quality.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.