acceptodds
Under review as a conference paper at ICLR 2027

Examining the Temporal Dynamics of Human Alignment by Probing Model Learning Windows

Abstract

Critical periods in biological development are temporal windows when experiences leave lasting effects on a system. Timing of critical periods differs based on the function being learned, with different sensory, cognitive, and linguistic systems following distinct schedules. In deep neural networks, critical periods have been identified primarily through task accuracy, collapsing this diversity into a single curve. We ask whether critical periods differ when measured through representational alignment with the human visual cortex. We train ResNet, ConvNeXt, and ViT under supervised ImageNet classification, ViT under contrastive language-image pretraining (CLIP), and pretrained CLIP-ViT under a fine-tuning task, injecting transient perturbations to inputs and targets at different training stages. We measure the effects at three levels: immediate change in internal representations (via similarity to human behaviorally-derived embeddings and fMRI activations), the trajectory representations follow as normal training resumes, and generalizability of learned embeddings via linear probes on out-of-distribution tasks. Early input-side perturbations cause the most disruption to representational structure, and models that recover task performance also recover human alignment. In supervised classifiers, human alignment and task performance recover from all perturbations. In CLIP, early and mid perturbations permanently damage task accuracy, yet after mid-training perturbations, brain alignment recovers in every case and behavioral alignment in two of three, showing the two objectives can dissociate. Late input-side perturbations have a different effect on ImageNet classifiers: representations of within-category objects are pulled closer together, distinguishing between-category boundaries and transiently improving human alignment. Critical periods in neural networks depend on what is measured, not just when. If late perturbations push representations toward human-like organization with little harm to performance, targeted interventions could steer representational alignment.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.