A 1-billion parameter multimodal embedding model from NVIDIA, based on Llama, designed for vision-language retrieval and embedding tasks.