The data behind the models: a 59,000-hour corpus, and what the models took from it

Quran Lab holds a corpus of more than 59,000 hours of recitation. No model trains on all of it: each takes a selection, built for what that model has to learn. This is how Zipformer Quran 3 and 3.1 were fed, with each number taken from the runs.

59,000+ hin the corpus, not all trained on
1,161new reciters, segmented by our aligner
1,850,886clips in the v3 training mix
5,424 hheard per epoch, for 10 epochs
48,732planted mistakes that taught v3.1 to listen
103,409madd lengths measured from the audio
301 hQuranTTS, ear-verified, 16 reciters
6,236ayat phonemized, round trip exact

Voices, not hours

v2 read emphatic letters badly on reciters it had never heard: emphatic ط was wrong 24.4% of the time, plain ت 2.6%. Adding noise didn't help. The cause was voices: the Quran side of training had only 33 reciters.

So the selection from the corpus was spent on voices. Hafs only, deduplicated, at least 40 kbps, files under 500 seconds, and 0.6 to 1 hour per reciter, shortest files first. Each surah was split into ayat by our own forced aligner, cut at the middle of the silence between them, then decoded and kept only if it matched its own ayah.

Reciters attempted 1,164
Reciters that survived 1,161
Clips segmented 533,636
Clips kept after the check 450,330 · 85%

What a recitation really sounds like

Isolated, perfect ayat don't teach a model what happens in a real session. Each of these came from a field report, was traced to a measured cause, and was fixed with data rather than architecture.

ProblemMeasured causeData that fixed itResult
The next ayah lost its opening wordsEvery training clip was one isolated ayah19,575 windows of 2 to 4 consecutive ayat, with 380 ms of the clip's own room tone between them1/12 → 12/12
A repeated ayah was dropped the second timeNo training clip ever repeated an ayah6,000 clips of the same ayah recited twice or three times, time-stretched so it isn't a copy3/12 → 11/12
Held leen couldn't be gradedOnly 8 ayat in the whole Quran end in leen at a stopTheir 314 real clips, heard 25 times per epoch65× more exposure
Fast recitation dropped wordsError jumped past 1.3× speed, exactly where augmentation stoppedSpeed from 0.8 to 1.5×, and tempo-only change that keeps pitchCliff removed
Phone recordings outdoorsThe noise corpus contains no windSynthetic wind, 29× more energy below 500 HzRobust outdoors

The training mix

v3 heard 1,850,886 clips, about 5,424 hours per epoch, for 10 epochs. Each tier has its own weight, set on purpose and printed on every launch.

  • EveryAyah studio recitationIsolated ayat, 15 s cap removed
    419 h×4
  • Broad Arabic speechMASC, MGB2, Common Voice and others
    1,400–1,500 h×1
  • Segmented scrape450,330 clips, 1,161 reciters
    669 h×2
  • Context windows and repeats25,575 clips of 2 to 4 ayat, and the same ayah twice
    129 h×4
  • Real phone recitationweighted by label quality
    92 h×1–4

Source hours per tier. The multiplier is how many times each tier is heard per epoch.

Planted mistakes

A model trained only on correct recitation learns each ayah's text by heart. When a different vowel was spliced into an ayah, v3 still preferred the written vowel 269 times out of 270. It wrote what the ayah should say, not what was said.

v3.1 was fine-tuned on counter-examples whose labels follow the audio. 30% of every tier is a control: the same audio spliced back with the label unchanged, so the splice itself tells the model nothing.

KindExampleHow it was madeRowsHeld out
Which soundFatha recited as dammaSplice in a real different vowel from the same reciter37,266714
Whether it happenedQalqalah burst missingCut the release burst out11,466435
How longMadd held 2 where 4 is writtenMeasure the hold, write the nearest legal length103,40920,796 changed
Wrong vowels reported0.0% → 27.5%v3 → v3.1, reciters it never heard
Dropped qalqalah reported0.3% → 46.1%controls stayed canonical 97 to 100%

Beyond recognition

The same data work feeds the other models and the public datasets.

What hasn't been used yet

Most of the corpus has not been trained on yet. v3 used 44.6% of the 224,233 EveryAyah clips on disk, from 36 reciters, and about 324,000 real phone recitations are still unlabelled. The next recognition model, TajweedFormer, was going to be called v4; it got its own name because its architecture is new.