A phone with little memory should still be able to follow a recitation, offline. Ternary quantization is the most extreme way to shrink a model that still works: every weight in a matrix becomes −1, 0 or +1, with one scale per row.
We applied it to 278 weight matrices of Zipformer Quran 3.1, 63.3 million of its 65.8 million parameters. Done directly, the error rate more than triples. So we trained the model once more, raising the ternary strength gradually while it learned, over 143,927 recordings (484 hours).
Accuracy comes back, and then some
| Phoneme error, 600 benchmark clips | Original | Ternary, direct | Ternary, trained |
|---|---|---|---|
| All 600 clips | 4.13% | 13.65% | 3.34% |
| Held-out studio reciters | 2.06% | 9.70% | 1.41% |
| Real phone recordings | 5.96% | 23.83% | 5.52% |
| A professional reciter it never heard | 5.09% | 11.24% | 3.92% |
On 16 other reciters (3,200 clips) it also beat the original: 2.22% against 3.96% on clean audio, and 3.04% against 4.48% with room echo added. Packed at two bits per weight, the recognition path fits in 22.4 MiB and reproduced every one of the benchmark's predictions.
Why it is not ready
Those benchmarks score correct recitation. On QDAT, 159 recordings of real readers with their mistakes annotated by people, the result turns around. The ternary model writes what the text should be, not what was said.
Its phoneme error against the human transcripts rose from 10.57% to 13.29%. A model that corrects the reader cannot tell the reader what to correct.
What it is good for
That same habit is useful when the canonical reading is exactly what you want: following along in the mushaf, aligning audio to the text, or cleaning up transcripts. As a canonicalizer, a model this small is ready to build on.
To keep the mistakes, the next run needs labelled recordings of real mistakes, which the training data does not yet have.
Zipformer Quran 3.1The production model this experiment started from View the model