A multimodal embedding model from the Qwen family, capable of processing both text and images for cross-modal retrieval and similarity.