Layer-Wise Causal Evidence of Script-Dependent Adaptation in Low-Resource Translation
Abstract
Parameter-efficient adaptation of translation models is judged by corpus-level metrics, while the question of where adaptation occurs is answered with a familiar set of layer-wise instruments: adapter norms, representational similarity, linear probes and layer ablations. We fine-tune IndicTrans2-1B with rank-64 LoRA under rsLoRA scaling, one adapter per translation direction shared by four language–script pairs, across eight pipelines covering English, Manipuri in two scripts, Assamese and Bodo. Adaptation is uneven and unordered by data volume: EnglishBodo, the smallest corpus, gains 8.69 points chrF++ and EnglishMeetei Mayek 4.10 points, whereas Bengali-script Manipuri is unchanged and Assamese, half of the training mixture, loses 3.34 points. Re-translating the training corpus for synthetic data changes chrF++ by less than one point in all forty comparisons and does not compound over rounds. The instruments prove less informative than they appear. Cross-lingual CKA mainly reflects whether two pipelines saw the same sentences, reaching 0.932 in the decoder for the one pair with identical inputs and at most 0.09 for the eleven pairs without; two probes reduce to one subword-fertility statistic; a part-of-speech probe reaches 0.94 accuracy at 0.03 selectivity; a script probe is perfect at every layer because its labels encode pipeline identity; and adapter norms never significantly predict ablation effects, whose load-bearing count ranges from 2 to 32 layers on differences of at most 1.05 chrF++. We release our code and resources.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.