Emphatic consonants are the hardest part of Quranic speech recognition. ص ض ط ظ differ from their plain counterparts س د ت ذ mostly at the back of the tongue, and in v1 and v2 they failed at three times the rate of plain letters.
Correct recitation depends on these letters, so the model can't lean on a language model to guess them. Every phoneme has to come from the sound.
صِرَٰطَ ٱلَّذِينَ أَنْعَمْتَ عَلَيْهِمْ غَيْرِ ٱلْمَغْضُوبِ عَلَيْهِمْ وَلَا ٱلضَّآلِّينَ
v23.0×emphatic vs plain error
v31 : 1on all three held-out sets
Small enough for a phone
The model has 65.5 million parameters and runs with INT8 quantization at a real-time factor of 0.07 on low-end and mobile CPUs, with no internet connection.
Zipformer Quran 3.1Phoneme-level speech recognition · 1.43% PER on held-out studio reciters View the model