Welcome to the BabyVLM Workshop: Toward Developmentally Plausible Multimodal Systems, to be held at NeurIPS 2026 in Atlanta, Georgia, USA.
Workshop Summary
Great effort has been invested into optimizing language model (LM) pretraining at massive scales in terms of pretraining dataset size; leading models today are trained on datasets of well over a trillion words. This is significantly more than the 100 million words that humans hear during typical development. How can we close the gap between human and machine language learners? Recent work suggests that engagement with other modalities may be the key to unlocking sample-efficient language learning.
Integrating more modalities raises critical new research questions and challenges relative to the text-only settings observed so far:
- How can we effectively integrate other modalities into LM training? Can this improve sample-efficiency relative to text-only training?
- By what standards should multimodal language learners be evaluated? What new benchmarks are needed to accurately assess learning of language and its grounding?
- Can we foster new kinds of research questions in computational cognitive science—for example, learnability in the presence or absence of visual signals?
The BabyVLM Workshop is designed to be the central forum for addressing these questions. In partnership with members of the BabyLM Challenge team, we invite researchers to contribute to the development of a community centered on questions of learnability, sample-efficient multimodal pretraining, cognition, and human-inspired evaluation of multimodal systems. The workshop will concentrate on key questions where the interaction of language and other modalities is key for learning, or for evaluating the capabilities of sample-efficient language learners.
Workshop Themes
Developmentally plausible multimodal language models. This theme explores the challenge of pretraining multimodal LMs using only as much data as a human has access to when first learning language. We will accept original work and kick off a competition, the BabyVLM Challenge, designed to encourage developments in sample-efficient vision language modeling.
Developmentally aligned evaluation. Most multimodal evaluation datasets focus on artificial tasks, or are highly difficult relative to the learning signals available to a human learner. To address this gap, we will encourage the submission of evaluations and benchmarks inspired by human language and vision capabilities. We aim to solicit evaluations that not only measure psychometric fit to human metrics, but also enable a more fine-grained view of multimodal language learning during earlier stages of pretraining.
Longitudinal egocentric learning. Cognitive scientists and machine learning researchers alike have benefited from the development of longitudinal, egocentric datasets. Studying and improving them may enable high-impact research.
Submission Format
- Page Limit: Submissions will be solicited as full papers of at most 8 pages. Paper submissions should discuss research or deployed work.
- Archival Option: This is a non-archival workshop. Calls for non-archival submissions will be made in July 2026.
- Reviewing: Submissions will be reviewed via a double-blind review process on OpenReview. Each paper will receive at least three reviews. See the full Call for Papers for details.
Important Dates (all times AoE):
- Call for Papers: July 2026
- Submission Deadline: August 22, 2026
- Notification of Acceptance: September 29, 2026
- Workshop Date: NeurIPS 2026
Keynote Speakers
See the Speakers & Panelists page for bios and the confirmed panel.
Schedule
All talks include a Q&A. The schedule is tentative and subject to change. Full details are on the Schedule page.
Opening Remarks Organizers
Invited Talk 1 Michael C. Frank (Stanford, confirmed)
Invited Talk 2 Uri Hasson (Princeton, confirmed)
Oral Highlights 1
Coffee Break
Poster Session
Lunch
Invited Talk 3 Freda Shi (U Waterloo, confirmed)
Invited Talk 4 Kristen Grauman (UT Austin, confirmed)
BabyVLM: Kickoff & Tutorial Organizers
Coffee Break
Invited Talk 5 Cordelia Schmid (Inria, France)
Oral Highlights 2
Panel Discussion Closing the Data Efficiency Gap
Paper Award & Closing Remarks Organizers
Workshop Venue
The workshop will be co-located with NeurIPS 2026 in Atlanta, Georgia, USA.
How can we get multimodal models to learn as efficiently as human infants? Learn with us at the BabyVLM workshop!
Contact us at bgong@bu.edu.
