A fine-tuned Wav2Vec2 model for cross-lingual speech recognition using eSpeak phoneme labels and Common Voice data.