A unified speech model from k2-fsa for voice activity detection, speaker diarization, and speech recognition.