Proving the potential of ternary quantization in ASR

Zipformer Quran 3.1 with almost every weight reduced to −1, 0 or +1. An internal experiment: it proves the idea, and it is not a production model.

A phone with little memory should still be able to follow a recitation, offline. Ternary quantization is the most extreme way to shrink a model that still works: every weight in a matrix becomes −1, 0 or +1, with one scale per row.

We applied it to 278 weight matrices of Zipformer Quran 3.1, 63.3 million of its 65.8 million parameters. Done directly, the error rate more than triples. So we trained the model once more, raising the ternary strength gradually while it learned, over 143,927 recordings (484 hours).

Accuracy comes back, and then some

Phoneme error, 600 benchmark clipsOriginalTernary, directTernary, trained
All 600 clips4.13%13.65%3.34%
Held-out studio reciters2.06%9.70%1.41%
Real phone recordings5.96%23.83%5.52%
A professional reciter it never heard5.09%11.24%3.92%

On 16 other reciters (3,200 clips) it also beat the original: 2.22% against 3.96% on clean audio, and 3.04% against 4.48% with room echo added. Packed at two bits per weight, the recognition path fits in 22.4 MiB and reproduced every one of the benchmark's predictions.

Why it is not ready

Those benchmarks score correct recitation. On QDAT, 159 recordings of real readers with their mistakes annotated by people, the result turns around. The ternary model writes what the text should be, not what was said.

Original22.7%of 194 real tajweed deviations kept
Ternary, trained5.7%kept; 94.3% were written as the canonical reading

Its phoneme error against the human transcripts rose from 10.57% to 13.29%. A model that corrects the reader cannot tell the reader what to correct.

What it is good for

That same habit is useful when the canonical reading is exactly what you want: following along in the mushaf, aligning audio to the text, or cleaning up transcripts. As a canonicalizer, a model this small is ready to build on.

To keep the mistakes, the next run needs labelled recordings of real mistakes, which the training data does not yet have.

Zipformer Quran 3.1The production model this experiment started from View the model