We evaluate Microsoft Azure Pronunciation Assessment (PA) for Arabic-speaking children using three datasets: two word-level datasets (Arabic Speech Mispronunciation Detection Dataset (ASMDD) and Down syndrome) and one sentence-level dataset (Arabic Pronunciation Assessment for Saudi Arabian Students (APASAS)). Supplying reference text to PA stabilizes Accuracy, Fluency, Completeness, and Pronunciation scores, while no-reference mode inflates high scores since it grades recognition hypotheses rather than prompts. Errors cluster around difficult Arabic sounds (emphatics, pharyngeals, uvular /ق/, interdentals) for both readers and Azure PA. Standardized audio front-end processing (noise conditioning, voice-activity detection, pre-emphasis, loudness normalization) yields small but consistent improvements on word-level corpora. On APASAS, preprocessing reduces Azure’s over-rating and better matches teacher labels. With intended text supplied, Azure PA is suitable for tracking intelligibility and formative feedback; no-reference is best for basic screening.
Evaluating Azure Pronunciation Assessment for Mispronunciation Detection in Children Learning Arabic
35 views
2 Downloads