A Wav2Vec2 model fine-tuned on Common Voice using eSpeak phoneme labels for cross-lingual speech recognition.