From what should be said to what was said
Trained only on correct recitation, v3 learned each ayah's text by heart: with a different vowel spliced in, it still wrote the canonical vowel 269 times out of 270. v3.1 was fine-tuned for two epochs on planted mistakes whose labels follow the audio, with 30% controls that change nothing.
| Reciters it never heard | v3 | v3.1 |
|---|---|---|
| Wrong mid-ayah vowel reported | 0.0% | 27.5% |
| Deleted qalqalah reported | 0.3% | 46.1% |
| Controls still written as the text | n/a | 97.3–100% |
In the field, an engine built on v3.1 caught a deliberately wrong damma in Al-Buruj 85:11 end to end.
What changed
v3.1 is a fine-tune of v3 that retunes madd length. The architecture, the streaming interface and the symbol table are the same, so it drops in wherever v3 runs, including on iPhone. Every weight is different, so test it on your own recordings before switching.
The reason is in the labels. v3 learned munfasil and ʿarid at a fixed 4 harakat, while reciters may rightly hold them for 2, 4 or 6.
For grading how long a madd was held, forced-alignment timing from MFA Quran Hafs is still more reliable than counting symbols. See the models
On Apple's Neural Engine
v3.1 ships as two CoreML packages for iOS 18 and later: 8-bit weights at 63 MB, recommended, and a 124 MB half-precision copy. The encoder is split into two functions, because as one program it goes over the Neural Engine's compile budget and quietly falls back to the CPU.
Checked against its own INT8 ONNX export on 35 recitations, v3.1 on the Neural Engine produced the same 554 tokens on all 35.
A corrected number
An earlier version of the model card gave 4.92% phoneme error for a professional reciter the model never heard. That figure does not reproduce on the shipped weights. The corrected value is 9.10%, and every page now uses it.
An audit after release also found 24 clips of the three studio benchmark reciters in the training data: 0.001% of the fine-tune set, and none of the 200 benchmark clips. It is reported as is.
v3.1 on the full benchmark
- All 600 clips
- 4.13%
- 25,989 reference phonemes
- Held-out studio reciters
- 2.06%
- 200 / 200 clips
- Real phone recordings
- 5.96%
- 200 / 200 clips
- A professional reciter it never heard
- 5.09%
- 200 / 200 clips
For v3's own numbers against v2, see the benchmark. See the benchmark
Zipformer Quran 3.1Streaming phoneme recognition · NPL-1.2 View the model