Voices, not hours
v2 read emphatic letters badly on reciters it had never heard: emphatic ط was wrong 24.4% of the time, plain ت 2.6%. Adding noise didn't help. The cause was voices: the Quran side of training had only 33 reciters.
So the selection from the corpus was spent on voices. Hafs only, deduplicated, at least 40 kbps, files under 500 seconds, and 0.6 to 1 hour per reciter, shortest files first. Each surah was split into ayat by our own forced aligner, cut at the middle of the silence between them, then decoded and kept only if it matched its own ayah.
What a recitation really sounds like
Isolated, perfect ayat don't teach a model what happens in a real session. Each of these came from a field report, was traced to a measured cause, and was fixed with data rather than architecture.
| Problem | Measured cause | Data that fixed it | Result |
|---|---|---|---|
| The next ayah lost its opening words | Every training clip was one isolated ayah | 19,575 windows of 2 to 4 consecutive ayat, with 380 ms of the clip's own room tone between them | 1/12 → 12/12 |
| A repeated ayah was dropped the second time | No training clip ever repeated an ayah | 6,000 clips of the same ayah recited twice or three times, time-stretched so it isn't a copy | 3/12 → 11/12 |
| Held leen couldn't be graded | Only 8 ayat in the whole Quran end in leen at a stop | Their 314 real clips, heard 25 times per epoch | 65× more exposure |
| Fast recitation dropped words | Error jumped past 1.3× speed, exactly where augmentation stopped | Speed from 0.8 to 1.5×, and tempo-only change that keeps pitch | Cliff removed |
| Phone recordings outdoors | The noise corpus contains no wind | Synthetic wind, 29× more energy below 500 Hz | Robust outdoors |
The training mix
v3 heard 1,850,886 clips, about 5,424 hours per epoch, for 10 epochs. Each tier has its own weight, set on purpose and printed on every launch.
- EveryAyah studio recitationIsolated ayat, 15 s cap removed419 h×4
- Broad Arabic speechMASC, MGB2, Common Voice and others1,400–1,500 h×1
- Segmented scrape450,330 clips, 1,161 reciters669 h×2
- Context windows and repeats25,575 clips of 2 to 4 ayat, and the same ayah twice129 h×4
- Real phone recitationweighted by label quality92 h×1–4
Source hours per tier. The multiplier is how many times each tier is heard per epoch.
Planted mistakes
A model trained only on correct recitation learns each ayah's text by heart. When a different vowel was spliced into an ayah, v3 still preferred the written vowel 269 times out of 270. It wrote what the ayah should say, not what was said.
v3.1 was fine-tuned on counter-examples whose labels follow the audio. 30% of every tier is a control: the same audio spliced back with the label unchanged, so the splice itself tells the model nothing.
| Kind | Example | How it was made | Rows | Held out |
|---|---|---|---|---|
| Which sound | Fatha recited as damma | Splice in a real different vowel from the same reciter | 37,266 | 714 |
| Whether it happened | Qalqalah burst missing | Cut the release burst out | 11,466 | 435 |
| How long | Madd held 2 where 4 is written | Measure the hold, write the nearest legal length | 103,409 | 20,796 changed |
Beyond recognition
The same data work feeds the other models and the public datasets.
What hasn't been used yet
Most of the corpus has not been trained on yet. v3 used 44.6% of the 224,233 EveryAyah clips on disk, from 36 reciters, and about 324,000 real phone recitations are still unlabelled. The next recognition model, TajweedFormer, was going to be called v4; it got its own name because its architecture is new.