Benchmark

Phoneme error rate (PER) on held-out sets: reciters, phones and a professional the model never heard. Lower is better.

Zipformer Quran 3.1, the current model

All 600 clips of quranic-asr-benchmark v1.1, every one scored, with the corrected references. Greedy decoding in the streaming setting, no language model.

All 600 clips
4.13%
25,989 reference phonemes
Held-out studio reciters
2.06%
200 / 200 clips
Real phone recordings
5.96%
200 / 200 clips
A professional reciter it never heard
5.09%
200 / 200 clips

About this benchmark

quranic-asr-benchmark v1.1 is Quran Lab's held-out test: 600 clips in three sets of 200, none of them used in training. Every model is scored on the same clips, so versions can be compared.

Held-out studio reciters200 EveryAyah clips

Clean studio recitation the model was not trained on.

Real phone recordings200 clips

Recitation on phone microphones, held out from training.

A professional reciter it never heard200 QUL clips, Al-Nufais

One professional voice absent from every training source.

What it cannot tell you

  1. 1

    It scores the written text, not tajweed judgement

    PER compares the output to the canonical phonemes of each ayah. It does not measure whether a reciter's mistake is caught. On a separate set of 159 recordings with human-marked tajweed deviations, 3.1 kept 22.7% of them and wrote the canonical form for the rest.

  2. 2

    The Quran is one closed text

    Every ayah in the benchmark also occurs in training, recited by other people. What is held out is the recording and the voice, never the words.

  3. 3

    The audit checks recordings, not voices

    No benchmark recording appears in training. That rules out the same file, not an acoustically identical copy under another name.

  4. 4

    Three sets, one professional

    The unseen-reciter set is a single voice. It shows how the model meets a new reciter, not how it meets every reciter.

Zipformer Quran 2 and 3

On quranic-asr-benchmark v1.1, as corrected on the model card.

v2, previous release v3, base of the current v3.1
  • Held-out studio reciters 3 EveryAyah reciters, 200 clips
    5.19%
    1.43%
    −72.4%
  • Real phone recordings Phone-mic recitation held out from training, 200 clips
    7.92%
    3.65%
    −53.9%
  • A professional reciter it never heard QUL reciter Al-Nufais, 200 clips
    11.91%
    9.10%
    −23.6%

Phoneme error rate. Lower is better. No language model is used: every phoneme comes from the sound, so a mistake in recitation stays visible instead of being corrected away.

Emphatic letters

ص ض ط ظ used to fail at a multiple of the rate of plain letters. In v3 the gap is statistically zero on all three held-out sets.

v23.0×emphatic errors per plain-letter error
v31 : 1statistically zero gap

What it's measured against

1,100+
Reciters
36+ full readings
6,236
Ayat phonemized
every ayah of the Quran
251
Quranic phonetic symbols
tajweed included
59,000+
Hours in the corpus
not all used in training

Test your own model

The leaderboard scores any model on the same held-out sets. It runs on Hugging Face.

Open the leaderboard