A text-to-speech model from Microsoft based on the SpeechT5 architecture, enabling high-quality speech synthesis.