acceptodds
Under review as a conference paper at ICLR 2027

LITE: Layer-Wise Intermediate Token-Dropping for Efficient Structured-Grid PDE Emulation

Abstract

Transformer-based neural surrogates for partial differential equations achieve strong predictive accuracy on structured grids, but their training cost grows rapidly with resolution, and in three dimensions it becomes the binding constraint on what can be trained at all. We introduce LITE (Layer-wise Intermediate Token-dropping), which makes sparse–dense token processing work in hierarchical PDE transformers. Rather than routing tokens around the entire network, LITE applies a drop–process–reconstruct cycle independently at each encoder and decoder stage while keeping the latent stage dense. Before each stage an -norm importance score retains one token per spatial block, concentrating computation on physically active regions. The stage's transformer blocks operate on that subset alone. A depthwise spread convolution then propagates the result to the dropped positions, and a learned linear fusion recombines it with the dense features carried past the stage untouched. Because full resolution is restored between stages, multi-scale feature flow and skip connections remain intact and no full-token fine-tuning phase is needed, so peak memory never returns to that of the dense model. In 2D, LITE halves wall-clock training time while reducing normalized RMSE by on a 16-equation benchmark. Whole-network dropping at the same keep rate degrades sharply over a ten-step rollout, ending up worse than the dense model it aims to accelerate. In 3D at , dropping cuts per-sample latency and peak memory by – without increasing terminal rollout error, so a 167M-parameter model trains with lower latency and less memory than a 42M-parameter dense backbone. On full volumes, three times the linear extent used in prior work on this data, the dense baselines exhaust memory, while LITE trains and improves rollout accuracy. At these resolutions token dropping is thus a precondition for training rather than an optimization.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.