A multilingual vision-language embedding model that maps images and text to a shared embedding space.