A 4-bit quantized embedding model based on Qwen3 with 4B parameters, using W4A16 and group size 128 for efficient retrieval.