The BabyVLM Challenge
In partnership with organizers of the BabyLM Challenge, we will launch the BabyVLM Challenge at the workshop. This is a shared task focused on developmentally plausible and sample-efficient vision language modeling. We have designed this challenge with new datasets and evaluations specifically designed for small vision language models. We are planning a tutorial at the workshop to introduce attendees to the challenge. It will be advertised to BabyLM Challenge participants.
Previously, multimodal submissions to the BabyLM Challenge have been limited, likely due to computer vision researchers not attending NLP conferences as frequently as machine learning conferences, and due to the data not being specifically designed for multimodal learning nor effective vision modeling. Our intention in establishing this challenge is to build a wider community of researchers interested in the intersection of developmentally plausible machine learning and multimodal modeling, especially (but not exclusively) those training vision language models.
