acceptodds
Under review as a conference paper at ICLR 2027

Continual Visual and Verbal Learning Through a Child's Egocentric Input

Abstract

Children learn the meanings of words from a continuous, temporally structured stream of egocentric experience. Recent work shows that neural networks can also learn word-referent mappings from a child's egocentric video recordings, but they cycle through the shuffled data for hundreds of epochs, contrasting with how children actually encounter their environment. We introduce BabyCL, a continual multimodal framework that combines streaming visual representation learning with an image–text contrastive objective to learn in a single chronological pass from naturalistic child-headcam video–language streams. BabyCL combines a multi-stage temporal segmentation of the stream with a dual replay buffer that independently manages visual and multimodal histories, and it jointly optimizes three objectives on a shared backbone. Against a streaming baseline that shares BabyCL's replay, schedule, and labeled-pair exposure but omits its visual objectives, BabyCL's frozen visual features yield consistently higher linear-probe accuracy, both in domain and on a separate child-headcam corpus. BabyCL also improves four-alternative forced-choice word–referent accuracy over this baseline and narrows the gap to offline training. Together, these results show that meaningful word-referent mappings can emerge under training conditions much closer to a child's actual experience.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.