A text-to-speech model based on the SpeechT5 architecture, designed for generating natural-sounding speech from text.