Zipformer Quran 3.1: it now writes what was said

v3 wrote what an ayah should say. v3.1 reports a wrong vowel 27.5% of the time where v3 never did, holds madd as it was recited, and runs on Apple's Neural Engine. On all 600 benchmark clips it scores 4.13%.

From what should be said to what was said

Trained only on correct recitation, v3 learned each ayah's text by heart: with a different vowel spliced in, it still wrote the canonical vowel 269 times out of 270. v3.1 was fine-tuned for two epochs on planted mistakes whose labels follow the audio, with 30% controls that change nothing.

Reciters it never heardv3v3.1
Wrong mid-ayah vowel reported0.0%27.5%
Deleted qalqalah reported0.3%46.1%
Controls still written as the textn/a97.3–100%

In the field, an engine built on v3.1 caught a deliberately wrong damma in Al-Buruj 85:11 end to end.

What changed

v3.1 is a fine-tune of v3 that retunes madd length. The architecture, the streaming interface and the symbol table are the same, so it drops in wherever v3 runs, including on iPhone. Every weight is different, so test it on your own recordings before switching.

The reason is in the labels. v3 learned munfasil and ʿarid at a fixed 4 harakat, while reciters may rightly hold them for 2, 4 or 6.

Training labels4harakat, always
What reciters choose2 · 4 · 6harakat, all correct

For grading how long a madd was held, forced-alignment timing from MFA Quran Hafs is still more reliable than counting symbols. See the models

On Apple's Neural Engine

v3.1 ships as two CoreML packages for iOS 18 and later: 8-bit weights at 63 MB, recommended, and a 124 MB half-precision copy. The encoder is split into two functions, because as one program it goes over the Neural Engine's compile budget and quietly falls back to the CPU.

Checked against its own INT8 ONNX export on 35 recitations, v3.1 on the Neural Engine produced the same 554 tokens on all 35.

A corrected number

An earlier version of the model card gave 4.92% phoneme error for a professional reciter the model never heard. That figure does not reproduce on the shipped weights. The corrected value is 9.10%, and every page now uses it.

An audit after release also found 24 clips of the three studio benchmark reciters in the training data: 0.001% of the fine-tune set, and none of the 200 benchmark clips. It is reported as is.

v3.1 on the full benchmark

All 600 clips
4.13%
25,989 reference phonemes
Held-out studio reciters
2.06%
200 / 200 clips
Real phone recordings
5.96%
200 / 200 clips
A professional reciter it never heard
5.09%
200 / 200 clips

For v3's own numbers against v2, see the benchmark. See the benchmark

Zipformer Quran 3.1Streaming phoneme recognition · NPL-1.2 View the model