PREFILL DOES NOT NEED THE FULL FFN
Abstract
Recent studies have shown that computation during prefill can be reduced, with several approaches focusing on skipping entire Transformer layers. In this work, we investigate prefill computation at a finer granularity by reducing feed-forward network (FFN) capacity. We ask two questions: **(Q1)** To what extent can FFN capacity be reduced during prefill while preserving model quality? **(Q2)** How should the retained FFN capacity be distributed across layers? To answer these questions, we introduce a __Nested FFN__ architecture, in which prefill activates only a fixed subset of intermediate FFN channels while decode uses the full FFN, with both phases sharing the same parameters. Our experiments show that substantial FFN capacity (37.5% in multiple settings) can be removed during prefill, including in early Transformer layers.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.