A multimodal embedding model that encodes both text and images into a shared vector space for cross-modal retrieval.