A specialized variant of OpenAI's GPT-4o model for transcribing audio with speaker diarization to distinguish multiple speakers.