Zipformer Quran 3.1
Phoneme-level speech recognition for Quran recitation. It writes what was recited phoneme by phoneme, tajweed included, streaming on a phone with no internet. There is no language model, so a mistake is written as it was recited, never corrected away.
- Task
- Streaming phoneme recognition
- Base
- v3, fine-tuned for madd length
- Parameters
- 65.8M
- Language model
- None, by design
- Runs
- On device, offline: ONNX and Apple Neural Engine
- iOS package
- 63 MB, 8-bit weights
- Phonetic alphabet
- 251 symbols, tajweed included
- Licence
- NPL-1.2
- Updated
- 28 Aug 2026
Hear it work
Al-Fatiha with the word being read and the phonemes the model writes for it.
The Beneficent, the Merciful.
Recite with it
ReadRead along from the mushaf page, ayah by ayah, or as a teleprompter. The text follows your voice.
MemoriseThe words stay hidden until you recite them. Stuck? A hint shows the next word, one per tap.
ShuffleRandom ayat from the surahs, juz or pages you choose. Recite each from memory, then see your score.
Results
All 600 clips of quranic-asr-benchmark v1.1, measured by the team on 14 September 2026.
- All 600 clips
- 4.13%
- 25,989 reference phonemes
- Held-out studio reciters
- 2.06%
- 200 / 200 clips
- Real phone recordings
- 5.96%
- 200 / 200 clips
- A professional reciter it never heard
- 5.09%
- 200 / 200 clips
Use it
Access is granted on Hugging Face after agreeing to NPL-1.2. Download only the files you need.
from huggingface_hub import snapshot_download
# only v3.1 (INT8 ONNX) and its symbol table, not the whole repository
path = snapshot_download(
"Quran-Lab/zipformer_p-arabic-v3",
allow_patterns=["zipformer_p_arabic_v3.1.int8.onnx", "tokens.txt"],
)Known limits
Measured, not hypothetical. From the model card.
- 1
It hears the rule, not always the slip.
Trained almost entirely on correct recitation, it tends to write the rule when a reciter deviates. It reported full ikhfa for 81 of 94 clips where people heard a plain noon. Don't use it alone to judge ikhfa or similar contrasts.
- 2
Free-choice madd lengths are unreliable in the output.
Labels fixed munfasil and ʿarid at 4 harakat while reciters may choose 2, 4 or 6. v3.1 retunes this; grade durations from forced-alignment timing, not from the symbols.
- 3
Children are its weakest region.
Reciters under 12 score 2 to 3 times worse than adults on qdat_bench.
- 4
Chunk size depends on the audio.
On recitation, streaming chunks of 16 or 24 frames match full context. On other audio, full context is about 2 PER points better.
Research
- Sep 2026 Proving the potential of ternary quantization in ASR An internal experiment: Zipformer Quran 3.1 at −1, 0 and +1 reaches 3.34% phoneme error, but it is overfit to the canonical text. Not for feedback yet; ready to build a canonicalizer on. ASR
- Sep 2026 Zipformer Quran 3.1: it now writes what was said 4.13% phoneme error on all 600 benchmark clips, madd held as it was recited, and the Neural Engine. ASR
- Sep 2026 Zipformer Quran 3: closing the emphatic-letter gap How ص ض ط ظ went from three times the error of plain letters to parity on every held-out set. ASR