Benchmark
Phoneme error rate (PER) on held-out sets: reciters, phones and a professional the model never heard. Lower is better.
Zipformer Quran 3.1, the current model
All 600 clips of quranic-asr-benchmark v1.1, every one scored, with the corrected references. Greedy decoding in the streaming setting, no language model.
- All 600 clips
- 4.13%
- 25,989 reference phonemes
- Held-out studio reciters
- 2.06%
- 200 / 200 clips
- Real phone recordings
- 5.96%
- 200 / 200 clips
- A professional reciter it never heard
- 5.09%
- 200 / 200 clips
About this benchmark
quranic-asr-benchmark v1.1 is Quran Lab's held-out test: 600 clips in three sets of 200, none of them used in training. Every model is scored on the same clips, so versions can be compared.
Clean studio recitation the model was not trained on.
Recitation on phone microphones, held out from training.
One professional voice absent from every training source.
What it cannot tell you
- 1
It scores the written text, not tajweed judgement
PER compares the output to the canonical phonemes of each ayah. It does not measure whether a reciter's mistake is caught. On a separate set of 159 recordings with human-marked tajweed deviations, 3.1 kept 22.7% of them and wrote the canonical form for the rest.
- 2
The Quran is one closed text
Every ayah in the benchmark also occurs in training, recited by other people. What is held out is the recording and the voice, never the words.
- 3
The audit checks recordings, not voices
No benchmark recording appears in training. That rules out the same file, not an acoustically identical copy under another name.
- 4
Three sets, one professional
The unseen-reciter set is a single voice. It shows how the model meets a new reciter, not how it meets every reciter.
Zipformer Quran 2 and 3
On quranic-asr-benchmark v1.1, as corrected on the model card.
Phoneme error rate. Lower is better. No language model is used: every phoneme comes from the sound, so a mistake in recitation stays visible instead of being corrected away.
Emphatic letters
ص ض ط ظ used to fail at a multiple of the rate of plain letters. In v3 the gap is statistically zero on all three held-out sets.
What it's measured against
Test your own model
The leaderboard scores any model on the same held-out sets. It runs on Hugging Face.