The BabyVLM Workshop
Toward Developmentally
Plausible Multimodal Systems
How can we get multimodal models to learn as efficiently as human infants?
NeurIPS 2026 · Atlanta, Georgia, USAAbout the Workshop
Leading language models are trained on well over a trillion words—vastly more than the ~100 million words a child hears during typical development. The BabyVLM Workshop asks how we can close this gap, building on the idea that engagement with other modalities—especially vision—may be key to sample-efficient language learning.
We bring together machine learning, computer vision, cognitive science, and developmental psychology around learnability, sample-efficient multimodal pretraining, and human-inspired evaluation. In partnership with the BabyLM Challenge team, we invite contributions—from full papers to challenge entries—on:
- Effectively integrating other modalities into LM training to improve sample-efficiency
- New benchmarks that assess the learning of language and its grounding
- Computational cognitive science—e.g., learnability with or without visual signals
Call for Papers
Important Dates
Submission Format
- Non-archival submissions. Calls for non-archival submissions will be made in July.
- Page limit. Submissions will be solicited as full papers of at most 8 pages.
- Scope. Paper submissions should discuss research or deployed work.
Reviewing
The initial review of submissions will be handled by the program committee via a double-blind review process on OpenReview. We will continue to invite additional reviewers as needed to ensure that each paper receives at least three reviews, with a maximum of three papers assigned to any given reviewer. To manage conflicts of interest, no organizer will be involved in assessing a submission from someone within the same organization, and no organizers nor any students with a conflict of interest with the organizers will submit papers to the workshop.
Workshop Themes
- Developmentally plausible multimodal language models. This theme explores the challenge of pretraining multimodal LMs using only as much data as a human has access to when first learning language. We will accept original work and kick off a competition, the BabyVLM Challenge, designed to encourage developments in sample-efficient vision language modeling.
- Developmentally aligned evaluation. Most multimodal evaluation datasets focus on artificial tasks, or are highly difficult relative to the learning signals available to a human learner. To address this gap, we will encourage the submission of evaluations and benchmarks inspired by human language and vision capabilities. We aim to solicit evaluations that not only measure psychometric fit to human metrics, but also enable a more fine-grained view of multimodal language learning during earlier stages of pretraining.
- Longitudinal egocentric learning. Cognitive scientists and machine learning researchers alike have benefited from the development of longitudinal, egocentric datasets. Studying and improving them may enable high-impact research.
Recommended Reading
Training Data
- SAYCam: A Large, Longitudinal Audiovisual Dataset Recorded From the Infant's Perspective. Sullivan et al. Open Mind, 2022.
- The BabyView Dataset: High-Resolution Egocentric Videos of Infants' and Young Children's Everyday Experiences. Long et al. arXiv, 2024.
Baselines
- BabyVLM: Data-Efficient Pretraining of VLMs Inspired by Infant Learning. Wang et al. ICCV, 2025.
- BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation Models. Wang et al. CVPR, 2026.
- Visual Instruction Tuning. Liu, Li, Wu, and Lee. NeurIPS, 2023.
- Looking to Learn: Token-wise Dynamic Gating for Low-Resource Vision-Language Modelling. Ganescu, Salhan, Caines, and Buttery. BabyLM Workshop, 2025.
Developmental Grounding
- NIH Baby Toolbox. A collection of validated measures for assessing cognitive, sensory, motor, and social-emotional development in children aged 0–3.
Interesting Findings
- Findings of the Third BabyLM Challenge: Accelerating Language Modeling Research with Cognitively Plausible Data. Charpentier et al. BabyLM Workshop, 2025.
Organizers
Paula Buttery
University of Cambridge
Leshem Choshen
Weizmann Institute of Science
Boqing Gong
Boston University / Google DeepMind
Aaron Mueller
Boston University
Suchir Salhan
University of Cambridge
Shengao Wang
Boston University
Wenqi Wang
Boston University
Max Whitton
Boston UniversityRead organizer bios
Paula Buttery (pjb48@cam.ac.uk) is the Deputy Head of the Department of Computer Science & Technology at the University of Cambridge and a Full Professor of Language and Machine Learning. She is co-Director of the Cambridge Language Sciences Interdisciplinary Research Centre and Principal Investigator of the Cambridge Institute for Automated Language Teaching and Assessment (ALTA). Paula, together with her research group, focuses on designing and pretraining performant Small (and Baby) Language Models, developing interpretable and efficient text-based and multimodal language model architectures. Her group has won four outstanding paper awards for their research on BabyLM.
Leshem Choshen (leshem.choshen@weizmann.ac.il) is a Zuckerman Faculty Senior Scientist (eq. Assistant Professor) at the Weizmann Institute of Science. He co-founded the BabyLM challenge and has led various open science and community research efforts, including TextArena (supporting millions of users at its peak) and EveryEvalEver (including a shared task and wide community outreach). His work on future humane AI includes broad evaluation efforts and human-focused research.
Boqing Gong (bgong@bu.edu) is a computer science faculty member at Boston University and a part-time research scientist at Google DeepMind. His research on machine learning and computer vision focuses on visual recognition, video, and AI models' generalization and efficiency. He received an NSF CAREER award for his project on BabyVLM.
Aaron Mueller (amueller@bu.edu) is an Assistant Professor of CS at Boston University. His research centers on language modeling, developmentally inspired evaluation methods, and interpretability. He co-founded and regularly organizes the BabyLM Challenge. He is actively involved in the NLP and ML communities through research, as well as event planning and conference chairing (e.g., publicity chair for EMNLP 2025, program chair for the New England Mechanistic Interpretability Workshop).
Suchir Salhan (sas245@cam.ac.uk) is a PhD Candidate at the University of Cambridge, working on Small Language Models and Cognitively-Inspired AI. He is an organiser of the BabyLM Shared Task and Workshop at EMNLP 2026, and has won an Outstanding Paper Award for his work on interpretable approaches for multimodal language modeling. His research focuses on developmental interpretability, language model learning dynamics, and tokenization.
Shengao Wang (wsashawn@bu.edu) is a Ph.D. student at Boston University co-advised by Prof. Boqing Gong and Prof. Venkatesh Saligrama. His research interests lie among vision language model pretraining, developmental learning, and robotics. He was the team lead of the BabyVLM-v2 project. He also participated in the second BabyLM challenge and contributed to the evaluation codebase.
Wenqi Wang (wqwang@bu.edu) is a second-year Ph.D. student at Boston University, advised by Prof. Boqing Gong. His research focuses on vision language model evaluation and world models. He was a co-author on the BabyVLM-v2 team.
Max Whitton (maxwh@bu.edu) is pursuing his Ph.D. at Boston University, advised by Prof. Boqing Gong. He does research at the intersection of computer vision and developmental learning, and was a co-author on the BabyVLM-v2 team.
View the program committee
Our program committee draws on eight institutions across four countries (the US, the UK, the Netherlands, and Germany), with expertise in multimodal machine learning, language modeling, evaluation, and training dynamics: Canaan Breiss (U Chicago), Bastian Bunzeck (Bielefeld U), Andrew Caines (U Cambridge), Arjun Chandra (Boston U), Bianca-Mihaela Ganescu (U Cambridge), Fuhang Kuang (Boston U), Lingxiao Li (Boston U), Mahir Patel (Boston U), Yulu Qin (Boston U), Yuan Qing (Boston U), Maan Qraitem (Boston U), Arijit Ray (Boston U / Ai2), Naomi Saphra (Harvard, Kempner Institute), Ece Takmaz (Utrecht U), Yuwen Tan (Boston U), Zilu Tang (Boston U), Nazia Tasnim (Boston U), Piotr Teterwak (Boston U), Dheeraj Varghese (U Amsterdam), Manushree Vasu (Boston U), and Michael Wakeham (Boston U).
Invited Speakers
Michael C. Frank
Stanford University (confirmed)
Uri Hasson
Princeton University (confirmed)
Freda Shi
University of Waterloo (confirmed)
Kristen Grauman
UT Austin (confirmed)
Cordelia Schmid
Inria, France (unconfirmed)Confirmed Panelists
Panel discussion: Closing the Data Efficiency Gap
Michael C. Frank
Stanford University
Kristen Grauman
UT Austin
Freda Shi
University of Waterloo
Uri Hasson
Princeton UniversityTentative Schedule
All talks include a Q&A. Schedule is tentative and subject to change.
Opening Remarks
OrganizersInvited Talk 1
Michael C. Frank · StanfordInvited Talk 2
Uri Hasson · PrincetonOral Highlights 1
Contributed papersCoffee Break
Poster Session
All accepted papersLunch
Invited Talk 3
Freda Shi · U WaterlooInvited Talk 4
Kristen Grauman · UT AustinBabyVLM: Kickoff & Tutorial
OrganizersCoffee Break
Invited Talk 5
Cordelia Schmid · InriaOral Highlights 2
Contributed papersPanel: Closing the Data Efficiency Gap
PanelistsPaper Award & Closing Remarks
OrganizersWorkshop Venue
Atlanta, Georgia, USA
The BabyVLM Workshop is co-located with NeurIPS 2026 in Atlanta—the capital of Georgia and the largest city in the US Southeast, home to Georgia Tech, Emory, and a thriving AI and research community.
NeurIPS is expected to take place at the Georgia World Congress Center in downtown Atlanta. The airport (ATL) is one of the world's best-connected hubs, and the venue sits near Centennial Olympic Park, the Georgia Aquarium, and ample hotels.
Exact room and session details will be announced closer to the workshop.
The BabyVLM Challenge
The BabyVLM Challenge is a separate but related shared task, run in partnership with organizers of the BabyLM Challenge and focused on developmentally plausible, sample-efficient vision language modeling. It comes with new datasets and evaluations designed specifically for small vision language models.
We will announce and promote the challenge at the workshop—including a tutorial to introduce attendees to the task—and then aim to have results presented at CVPR the following year. No prior BabyLM experience is required.
How It Will Run
Announced at NeurIPS 2026
The task is launched and promoted at the workshop, with a tutorial introducing the data and evaluations.
Build a sample-efficient VLM
Pretrain on developmentally plausible data—roughly what a child sees and hears—and evaluate.
Results at CVPR
Findings and winning entries are aimed for presentation at CVPR the following year.
Details, datasets, and timeline will be released around the workshop—check back for updates.
Challenge Steering Committee (confirmed)
- Michael C. Frank (Psychology, Stanford)
- Chen Yu (Psychology, UT Austin)
- Kristen Grauman (CS, UT Austin)
- Kate Saenko (CS, Boston University)
Data Explorer
Every benchmark task is a vision–language conversation — a prompt carrying one or more
<image> tokens, paired with the ground-truth answer. The three sources differ in
whose eyes the footage comes from. Choose a source, then a task, to browse sample items.
Showing a curated subset for fast browsing. For the complete corpus and methodology, see the BabyVLM-v2 project page.
Contact
Venue
NeurIPS 2026 · Atlanta, Georgia, USA
Email Us
- Paula Buttery — pjb48@cam.ac.uk
- Leshem Choshen — leshem.choshen@weizmann.ac.il
- Boqing Gong — bgong@bu.edu
- Aaron Mueller — amueller@bu.edu
- Suchir Salhan — sas245@cam.ac.uk
- Shengao Wang — wsashawn@bu.edu
- Wenqi Wang — wqwang@bu.edu
- Max Whitton — maxwh@bu.edu
Acknowledgements
Special thanks to Michael Hua, Jason Liu, and Karthik Srikumar for their help building this website.