A multimodal large language model from the Qwen family that processes and generates images alongside text.