acceptodds
Under review as a conference paper at ICLR 2027

Level of Detail Attention: Allocating Resolution Instead of Sparsity

Abstract

Most sparse attention methods attend to a subset of keys, ignoring the rest and thereby amplifying any mistakes while entirely removing the ignored set from softmax normalization. We introduce Level of Detail (LoD) Attention, a training-free approximation to dense softmax attention that represents less important keys and values at reduced resolution instead of discarding them. For language inference, LoD requires no changes to pretrained model weights and achieves prefill work and work per decode step. Experiments on Qwen3.8 and K2 Horizon show that LoD largely preserves aggregate quality. On Qwen3.8, two-tier LoD accelerates attention computation by over 6x during both prefill and decode in our 128K prompt batch size 8 benchmarks. Its similarity-based grouping scheme also permits the reduction of the LLM key-value cache to a compact INT4 representation, saving VRAM with little additional aggregate quality loss. To show the versatility of our method across modalities we also demonstrate LoD for MiniMax-H3 video generation. We will release our code under the Apache 2 license.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.