Does Fine-Tuning Still Beat In-Context Learning Under LoRA?
Abstract
Does the advantage of fine-tuning (FT) over in-context learning (ICL) on synthetic formal languages persist under parameter-efficient adaptation? We compare full FT, LoRA and 4-bit QLoRA with single-pass ICL on Qwen2.5 base models from 0.5B to 3B parameters, using matched sets of 32 training strings and a published discriminative ROC-AUC evaluation. Under the published protocol, which selects each run's best epoch using test scores, all tested fine-tuning methods achieve higher mean AUC than ICL at every evaluated model size. However, at 3B, this ranking reverses when the same runs are evaluated after 50 epochs. To test whether the advantage survives validation-based checkpoint selection, we repeat the 3B comparison across four languages, choosing epochs on fresh validation strings and evaluating on separate held-out test strings. Full FT, LoRA and QLoRA achieve mean test AUCs of 0.942, 0.921 and 0.919, respectively, versus 0.873 for ICL; only the full-FT advantage remains statistically significant after correction for multiple comparisons. Selecting epochs on the test strings instead raises mean AUC only slightly, while all three fine-tuning methods again fall below ICL on average at epoch 50. Under the published protocol, adapter-versus-full-FT comparisons also depend on the adapter learning rate. Thus, in this 32-example setting at 3B, validation-based stopping preserves full FT's advantage over single-pass ICL, with weaker evidence for the adapters, whereas training for 50 epochs reverses the comparison.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.