acceptodds
Under review as a conference paper at ICLR 2027

Layerwise Function Anchoring: Preserving Sub-Module Functions on Sampled Hidden States for Continual Domain Adaptation

Abstract

Fine-tuning a language model on a narrow domain erodes the broad capabilities it arrived with, and can lead to catastrophic forgetting. The classical remedy, rehearsal, mitigates forgetting by replaying pretraining text at every stage. We present Layerwise Function Anchoring (LFA), which preserves capability with a statistic instead. LFA estimates, once, the distribution of hidden states entering each sub-module of the frozen original when text is passed through: a compact, sequence-free, task-agnostic statistic (under a tenth of the model's size). During adaptation LFA penalizes each sub-module's output drift independently on point samples , bypassing resource-intensive self-attention. No exemplars are stored, no text is generated during adaptation, and the statistic can be updated exactly as domains accumulate. Two knobs do the work: prevents forgetting by penalizing layerwise functional drift, and , an L2-SP penalty on all weights, covers the directions does not see. Anchoring sub-modules independently risks compounding error through depth; every whole-network measurement below (perplexity, judged answers, rule-scored generation) bounds that risk, though none isolates it, and the cleanest bound is tight: on Qwen3-0.6B adapted to a philosophy corpus under LoRA, unanchored adaptation inflates general-text perplexity by over 400% and LFA holds it to +1.1%. LFA's time overhead falls with model size, +14.8% per pass at 0.6B, +6.3% at 1.7B and +5.5% at 3B, at lower peak memory than Learning without Forgetting (LwF) at 0.6B and 1.7B. Judged by a frontier LLM, LFA matches tuned LwF at each method's best dose in paired per-question tests and is not separable from tuned replay, both strengthened with the same L2-SP penalty. Merging a learned domain's statistics into supports continual learning with no retained text: after a second domain LFA leads replay and LwF on the new domain's judged quality at all three ranks, and trails replay on old-domain perplexity. After a third domain its old-domain judged quality ends tied with replay, and LFA is ahead of both rehearsal and distillation on rule-scored generation (GSM8K, IFEval), where every perplexity on given text puts replay first. LFA's anchoring statistic need be produced only once, allowing model providers to support continual learning without divulging proprietary training data. When not available, the statistic can be formed using self-generated sequence samples, so LFA can run data-free, with nothing beyond the new domain's own text (one model, one domain). At Qwen3-1.7B (one seed) the continual-learning ordering is the same and every method loses less against the base model. All experiments use LoRA; seed coverage is partial.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.