The BabyVLM Workshop
Toward Developmentally Plausible Multimodal Systems


How can we get multimodal models to learn as efficiently as human infants?

NeurIPS 2026 · Atlanta, Georgia, USA
Vision Language Development

About the Workshop

Leading language models are trained on well over a trillion words—vastly more than the ~100 million words a child hears during typical development. The BabyVLM Workshop asks how we can close this gap, building on the idea that engagement with other modalities—especially vision—may be key to sample-efficient language learning.

We bring together machine learning, computer vision, cognitive science, and developmental psychology around learnability, sample-efficient multimodal pretraining, and human-inspired evaluation. In partnership with the BabyLM Challenge team, we invite contributions—from full papers to challenge entries—on:

  • Effectively integrating other modalities into LM training to improve sample-efficiency
  • New benchmarks that assess the learning of language and its grounding
  • Computational cognitive science—e.g., learnability with or without visual signals

Read the Call for Papers

Call for Papers

Important Dates

July 2026
Call for Papers opens
Aug 22, 2026
Submission deadline
Sep 29, 2026
Notification of acceptance
NeurIPS 2026
Workshop day, Atlanta

Submission Format

  • Non-archival submissions. Calls for non-archival submissions will be made in July.
  • Page limit. Submissions will be solicited as full papers of at most 8 pages.
  • Scope. Paper submissions should discuss research or deployed work.

Reviewing

The initial review of submissions will be handled by the program committee via a double-blind review process on OpenReview. We will continue to invite additional reviewers as needed to ensure that each paper receives at least three reviews, with a maximum of three papers assigned to any given reviewer. To manage conflicts of interest, no organizer will be involved in assessing a submission from someone within the same organization, and no organizers nor any students with a conflict of interest with the organizers will submit papers to the workshop.

Workshop Themes

  • Developmentally plausible multimodal language models. This theme explores the challenge of pretraining multimodal LMs using only as much data as a human has access to when first learning language. We will accept original work and kick off a competition, the BabyVLM Challenge, designed to encourage developments in sample-efficient vision language modeling.
  • Developmentally aligned evaluation. Most multimodal evaluation datasets focus on artificial tasks, or are highly difficult relative to the learning signals available to a human learner. To address this gap, we will encourage the submission of evaluations and benchmarks inspired by human language and vision capabilities. We aim to solicit evaluations that not only measure psychometric fit to human metrics, but also enable a more fine-grained view of multimodal language learning during earlier stages of pretraining.
  • Longitudinal egocentric learning. Cognitive scientists and machine learning researchers alike have benefited from the development of longitudinal, egocentric datasets. Studying and improving them may enable high-impact research.

Recommended Reading

Training Data

Baselines

Developmental Grounding

  • NIH Baby Toolbox. A collection of validated measures for assessing cognitive, sensory, motor, and social-emotional development in children aged 0–3.

Interesting Findings

Organizers

Paula Buttery

Paula Buttery

University of Cambridge
Leshem Choshen

Leshem Choshen

Weizmann Institute of Science
Boqing Gong

Boqing Gong

Boston University / Google DeepMind
Aaron Mueller

Aaron Mueller

Boston University
Suchir Salhan

Suchir Salhan

University of Cambridge
Shengao Wang

Shengao Wang

Boston University
Wenqi Wang

Wenqi Wang

Boston University
Max Whitton

Max Whitton

Boston University
Read organizer bios

Paula Buttery (pjb48@cam.ac.uk) is the Deputy Head of the Department of Computer Science & Technology at the University of Cambridge and a Full Professor of Language and Machine Learning. She is co-Director of the Cambridge Language Sciences Interdisciplinary Research Centre and Principal Investigator of the Cambridge Institute for Automated Language Teaching and Assessment (ALTA). Paula, together with her research group, focuses on designing and pretraining performant Small (and Baby) Language Models, developing interpretable and efficient text-based and multimodal language model architectures. Her group has won four outstanding paper awards for their research on BabyLM.

Leshem Choshen (leshem.choshen@weizmann.ac.il) is a Zuckerman Faculty Senior Scientist (eq. Assistant Professor) at the Weizmann Institute of Science. He co-founded the BabyLM challenge and has led various open science and community research efforts, including TextArena (supporting millions of users at its peak) and EveryEvalEver (including a shared task and wide community outreach). His work on future humane AI includes broad evaluation efforts and human-focused research.

Boqing Gong (bgong@bu.edu) is a computer science faculty member at Boston University and a part-time research scientist at Google DeepMind. His research on machine learning and computer vision focuses on visual recognition, video, and AI models' generalization and efficiency. He received an NSF CAREER award for his project on BabyVLM.

Aaron Mueller (amueller@bu.edu) is an Assistant Professor of CS at Boston University. His research centers on language modeling, developmentally inspired evaluation methods, and interpretability. He co-founded and regularly organizes the BabyLM Challenge. He is actively involved in the NLP and ML communities through research, as well as event planning and conference chairing (e.g., publicity chair for EMNLP 2025, program chair for the New England Mechanistic Interpretability Workshop).

Suchir Salhan (sas245@cam.ac.uk) is a PhD Candidate at the University of Cambridge, working on Small Language Models and Cognitively-Inspired AI. He is an organiser of the BabyLM Shared Task and Workshop at EMNLP 2026, and has won an Outstanding Paper Award for his work on interpretable approaches for multimodal language modeling. His research focuses on developmental interpretability, language model learning dynamics, and tokenization.

Shengao Wang (wsashawn@bu.edu) is a Ph.D. student at Boston University co-advised by Prof. Boqing Gong and Prof. Venkatesh Saligrama. His research interests lie among vision language model pretraining, developmental learning, and robotics. He was the team lead of the BabyVLM-v2 project. He also participated in the second BabyLM challenge and contributed to the evaluation codebase.

Wenqi Wang (wqwang@bu.edu) is a second-year Ph.D. student at Boston University, advised by Prof. Boqing Gong. His research focuses on vision language model evaluation and world models. He was a co-author on the BabyVLM-v2 team.

Max Whitton (maxwh@bu.edu) is pursuing his Ph.D. at Boston University, advised by Prof. Boqing Gong. He does research at the intersection of computer vision and developmental learning, and was a co-author on the BabyVLM-v2 team.

View the program committee

Our program committee draws on eight institutions across four countries (the US, the UK, the Netherlands, and Germany), with expertise in multimodal machine learning, language modeling, evaluation, and training dynamics: Canaan Breiss (U Chicago), Bastian Bunzeck (Bielefeld U), Andrew Caines (U Cambridge), Arjun Chandra (Boston U), Bianca-Mihaela Ganescu (U Cambridge), Fuhang Kuang (Boston U), Lingxiao Li (Boston U), Mahir Patel (Boston U), Yulu Qin (Boston U), Yuan Qing (Boston U), Maan Qraitem (Boston U), Arijit Ray (Boston U / Ai2), Naomi Saphra (Harvard, Kempner Institute), Ece Takmaz (Utrecht U), Yuwen Tan (Boston U), Zilu Tang (Boston U), Nazia Tasnim (Boston U), Piotr Teterwak (Boston U), Dheeraj Varghese (U Amsterdam), Manushree Vasu (Boston U), and Michael Wakeham (Boston U).

Invited Speakers

Michael C. Frank

Michael C. Frank

Stanford University (confirmed)
Uri Hasson

Uri Hasson

Princeton University (confirmed)
Freda Shi

Freda Shi

University of Waterloo (confirmed)
Kristen Grauman

Kristen Grauman

UT Austin (confirmed)
Cordelia Schmid

Cordelia Schmid

Inria, France (unconfirmed)

Confirmed Panelists

Panel discussion: Closing the Data Efficiency Gap

Michael C. Frank

Michael C. Frank

Stanford University
Kristen Grauman

Kristen Grauman

UT Austin
Freda Shi

Freda Shi

University of Waterloo
Uri Hasson

Uri Hasson

Princeton University

Tentative Schedule

All talks include a Q&A. Schedule is tentative and subject to change.

15 sessions 8:50a–5:45p
Session8:50a · 10 min

Opening Remarks

Organizers
Invited9:00a · 35 min

Invited Talk 1

Michael C. Frank · Stanford
Invited9:35a · 35 min

Invited Talk 2

Uri Hasson · Princeton
Oral highlights10:10a · 30 min

Oral Highlights 1

Contributed papers
Break10:40a · 20 min

Coffee Break

Posters11:00a · 1h

Poster Session

All accepted papers
Break12:00p · 1h 30m

Lunch

Invited1:30p · 35 min

Invited Talk 3

Freda Shi · U Waterloo
Invited2:05p · 35 min

Invited Talk 4

Kristen Grauman · UT Austin
Session2:40p · 35 min

BabyVLM: Kickoff & Tutorial

Organizers
Break3:15p · 15 min

Coffee Break

Invited3:30p · 35 min

Invited Talk 5

Cordelia Schmid · Inria
Oral highlights4:05p · 30 min

Oral Highlights 2

Contributed papers
Panel4:35p · 1h

Panel: Closing the Data Efficiency Gap

Panelists
Session5:35p · 10 min

Paper Award & Closing Remarks

Organizers

Workshop Venue

Atlanta, Georgia, USA

The BabyVLM Workshop is co-located with NeurIPS 2026 in Atlanta—the capital of Georgia and the largest city in the US Southeast, home to Georgia Tech, Emory, and a thriving AI and research community.

NeurIPS is expected to take place at the Georgia World Congress Center in downtown Atlanta. The airport (ATL) is one of the world's best-connected hubs, and the venue sits near Centennial Olympic Park, the Georgia Aquarium, and ample hotels.

Exact room and session details will be announced closer to the workshop.

The BabyVLM Challenge

The BabyVLM Challenge is a separate but related shared task, run in partnership with organizers of the BabyLM Challenge and focused on developmentally plausible, sample-efficient vision language modeling. It comes with new datasets and evaluations designed specifically for small vision language models.

We will announce and promote the challenge at the workshop—including a tutorial to introduce attendees to the task—and then aim to have results presented at CVPR the following year. No prior BabyLM experience is required.

How It Will Run

1

Announced at NeurIPS 2026

The task is launched and promoted at the workshop, with a tutorial introducing the data and evaluations.

2

Build a sample-efficient VLM

Pretrain on developmentally plausible data—roughly what a child sees and hears—and evaluate.

3

Results at CVPR

Findings and winning entries are aimed for presentation at CVPR the following year.

Details, datasets, and timeline will be released around the workshop—check back for updates.

Explore the benchmark data

Challenge Steering Committee (confirmed)
  • Michael C. Frank (Psychology, Stanford)
  • Chen Yu (Psychology, UT Austin)
  • Kristen Grauman (CS, UT Austin)
  • Kate Saenko (CS, Boston University)

Data Explorer

Every benchmark task is a vision–language conversation — a prompt carrying one or more <image> tokens, paired with the ground-truth answer. The three sources differ in whose eyes the footage comes from. Choose a source, then a task, to browse sample items.

Showing a curated subset for fast browsing. For the complete corpus and methodology, see the BabyVLM-v2 project page.

Sources & Their Role
Task
Samples

Contact

Venue

NeurIPS 2026 · Atlanta, Georgia, USA

Email Us

  • Paula Buttery — pjb48@cam.ac.uk
  • Leshem Choshen — leshem.choshen@weizmann.ac.il
  • Boqing Gong — bgong@bu.edu
  • Aaron Mueller — amueller@bu.edu
  • Suchir Salhan — sas245@cam.ac.uk
  • Shengao Wang — wsashawn@bu.edu
  • Wenqi Wang — wqwang@bu.edu
  • Max Whitton — maxwh@bu.edu

Acknowledgements

Special thanks to Michael Hua, Jason Liu, and Karthik Srikumar for their help building this website.