Zipformer Quran 3.1

Phoneme-level speech recognition for Quran recitation. It writes what was recited phoneme by phoneme, tajweed included, streaming on a phone with no internet. There is no language model, so a mistake is written as it was recited, never corrected away.

Task
Streaming phoneme recognition
Base
v3, fine-tuned for madd length
Parameters
65.8M
Language model
None, by design
Runs
On device, offline: ONNX and Apple Neural Engine
iOS package
63 MB, 8-bit weights
Phonetic alphabet
251 symbols, tajweed included
Licence
NPL-1.2
Updated
28 Aug 2026

Hear it work

Al-Fatiha with the word being read and the phonemes the model writes for it.

Transcript
سورة الفاتحة
٣

The Beneficent, the Merciful.

ررَ rra حِ ḥi ۦۦ ii maddم m
0:14 / 0:46

Recite with it

  • Read 3 layouts

    Read along from the mushaf page, ayah by ayah, or as a teleprompter. The text follows your voice.

  • Memorise Hints

    The words stay hidden until you recite them. Stuck? A hint shows the next word, one per tap.

  • Shuffle Scored

    Random ayat from the surahs, juz or pages you choose. Recite each from memory, then see your score.

Results

All 600 clips of quranic-asr-benchmark v1.1, measured by the team on 14 September 2026.

Full benchmark
All 600 clips
4.13%
25,989 reference phonemes
Held-out studio reciters
2.06%
200 / 200 clips
Real phone recordings
5.96%
200 / 200 clips
A professional reciter it never heard
5.09%
200 / 200 clips

Use it

Access is granted on Hugging Face after agreeing to NPL-1.2. Download only the files you need.

Python
from huggingface_hub import snapshot_download

# only v3.1 (INT8 ONNX) and its symbol table, not the whole repository
path = snapshot_download(
    "Quran-Lab/zipformer_p-arabic-v3",
    allow_patterns=["zipformer_p_arabic_v3.1.int8.onnx", "tokens.txt"],
)

Known limits

Measured, not hypothetical. From the model card.

  1. 1

    It hears the rule, not always the slip.

    Trained almost entirely on correct recitation, it tends to write the rule when a reciter deviates. It reported full ikhfa for 81 of 94 clips where people heard a plain noon. Don't use it alone to judge ikhfa or similar contrasts.

  2. 2

    Free-choice madd lengths are unreliable in the output.

    Labels fixed munfasil and ʿarid at 4 harakat while reciters may choose 2, 4 or 6. v3.1 retunes this; grade durations from forced-alignment timing, not from the symbols.

  3. 3

    Children are its weakest region.

    Reciters under 12 score 2 to 3 times worse than adults on qdat_bench.

  4. 4

    Chunk size depends on the audio.

    On recitation, streaming chunks of 16 or 24 frames match full context. On other audio, full context is about 2 PER points better.

Research