A 1.7-billion-parameter text-to-speech model from Qwen with 12Hz output and custom voice cloning capabilities.