Welcome to the BabyVLM Workshop: Toward Developmentally Plausible Multimodal Systems, to be held at NeurIPS 2026 in Atlanta, Georgia, USA.

Workshop Summary

Great effort has been invested into optimizing language model (LM) pretraining at massive scales in terms of pretraining dataset size; leading models today are trained on datasets of well over a trillion words. This is significantly more than the 100 million words that humans hear during typical development. How can we close the gap between human and machine language learners? Recent work suggests that engagement with other modalities may be the key to unlocking sample-efficient language learning.

Integrating more modalities raises critical new research questions and challenges relative to the text-only settings observed so far:

The BabyVLM Workshop is designed to be the central forum for addressing these questions. In partnership with members of the BabyLM Challenge team, we invite researchers to contribute to the development of a community centered on questions of learnability, sample-efficient multimodal pretraining, cognition, and human-inspired evaluation of multimodal systems. The workshop will concentrate on key questions where the interaction of language and other modalities is key for learning, or for evaluating the capabilities of sample-efficient language learners.

Workshop Themes

  1. Developmentally plausible multimodal language models. This theme explores the challenge of pretraining multimodal LMs using only as much data as a human has access to when first learning language. We will accept original work and kick off a competition, the BabyVLM Challenge, designed to encourage developments in sample-efficient vision language modeling.

  2. Developmentally aligned evaluation. Most multimodal evaluation datasets focus on artificial tasks, or are highly difficult relative to the learning signals available to a human learner. To address this gap, we will encourage the submission of evaluations and benchmarks inspired by human language and vision capabilities. We aim to solicit evaluations that not only measure psychometric fit to human metrics, but also enable a more fine-grained view of multimodal language learning during earlier stages of pretraining.

  3. Longitudinal egocentric learning. Cognitive scientists and machine learning researchers alike have benefited from the development of longitudinal, egocentric datasets. Studying and improving them may enable high-impact research.

Submission Format

Submit on OpenReview

Important Dates (all times AoE):

Keynote Speakers

Michael C. Frank Stanford University Confirmed
Uri Hasson Princeton University Confirmed
Freda Shi University of Waterloo Confirmed
Kristen Grauman UT Austin Confirmed
Cordelia Schmid Inria, France Unconfirmed

See the Speakers & Panelists page for bios and the confirmed panel.

Schedule

All talks include a Q&A. The schedule is tentative and subject to change. Full details are on the Schedule page.

Opening Remarks Organizers

Invited Talk 1 Michael C. Frank (Stanford, confirmed)

Invited Talk 2 Uri Hasson (Princeton, confirmed)

Oral Highlights 1

Coffee Break

Poster Session

Lunch

Invited Talk 3 Freda Shi (U Waterloo, confirmed)

Invited Talk 4 Kristen Grauman (UT Austin, confirmed)

BabyVLM: Kickoff & Tutorial Organizers

Coffee Break

Invited Talk 5 Cordelia Schmid (Inria, France)

Oral Highlights 2

Panel Discussion Closing the Data Efficiency Gap

Paper Award & Closing Remarks Organizers

Workshop Venue

The workshop will be co-located with NeurIPS 2026 in Atlanta, Georgia, USA.

Atlanta, Georgia, USA Exact venue and room to be announced.

How can we get multimodal models to learn as efficiently as human infants? Learn with us at the BabyVLM workshop!

Contact us at bgong@bu.edu.