acceptodds
Under review as a conference paper at ICLR 2027

J-Space-Guided Structural Search for Language Model Compression

Abstract

Structured pruning reduces language-model cost but risks removing computations needed across tasks. Capacity allocation alone leaves unresolved which equal-sized channel groups preserve useful semantic content. We propose candidate-dynamic J-space structural search to address this problem. The framework combines semantic mask characterization, joint capacity–identity optimization, and surrogate-assisted evaluation. First, a frozen Jacobian lens decomposes dense and pruned activations independently, exposing changes in sparse semantic support rather than imposing the dense model's support on every candidate. Second, a mixed-discrete representation combines component capacities and concrete masks under an exact parameter budget, while second-order compensation adjusts retained weights. Third, an online surrogate selects promising structures for full-model evaluation and updates a Pareto archive from measured behavior and semantic drift. Together, these components connect local semantic measurements to global structure selection without treating surrogate predictions as observed performance. On Qwen3.5-4B with 11.84% of parameters removed, two selected structures improve question-weighted accuracy by 1.10 and 1.19 percentage points over an equal-parameter non-J structural reference across eight benchmarks, with slightly lower WikiText-2 perplexity. The structures approach a conventional allocation reference, while results on CommonsenseQA, SciQ, and COPA reveal task-dependent trade-offs. These findings support semantic information as a useful structural-search signal, while leaving its advantage under fully matched search compute unresolved.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.