SAWT v4: 48 kHz restoration of archive recitations

Tape hiss, echo and repeated compression stand between many historical recitations and their listeners. SAWT v4 takes them away at 48 kHz, and checks every window it renders against what the recording said.

Restoration with restraint

Removing noise from a recitation is only useful if the recitation survives. A restorer that sounds fluent but swaps a consonant has changed what was recited, and a quality score will not notice. SAWT was built around that difference: it has to be clean, and it has to stay accountable to the recording.

Recitation also makes the difference measurable. Old recordings rarely have a clean copy to compare against, but every one of them can be checked against the text of the Quran. The text is used only to grade the output; SAWT never sees it.

One part reads, another renders

An anchor reads the damaged recording: a frozen w2v-BERT 2.0 network with 42.3M trainable LoRA parameters, which estimates the speech content, the pitch, the speaker and a first clean latent. Its output is fixed for a given input.

A generator renders the sound under those estimates: a 372M-parameter diffusion transformer trained with flow matching in the 48 kHz latent space of DAC-VAE. It works in 8 s windows, 64 steps each. The design keeps the reading with the anchor; the generator supplies texture.

Every window is then decoded and checked three times before it is kept:

CheckThreshold
Energy: the window has not collapsed0.70
Content: features read back from the output agree with the input0.85
Speaker: the voice matches the window before it0.90

A failed check draws the window again, up to three times. Windows that still fail are listed in a sidecar file, so an archive can refuse them or send them for review.

Listen

The same recording before and after, on one timeline. Switch while it plays to hear the same instant.

SAWT v4 · restoration
Yunus 10:24

A 44.1 kHz archive recording, restored to 48 kHz.

إِنَّمَا مَثَلُ ٱلْحَيَوٰةِ ٱلدُّنْيَا كَمَآءٍ أَنزَلْنَـٰهُ مِنَ ٱلسَّمَآءِ فَٱخْتَلَطَ بِهِۦ نَبَاتُ ٱلْأَرْضِ مِمَّا يَأْكُلُ ٱلنَّاسُ وَٱلْأَنْعَـٰمُ حَتَّىٰٓ إِذَآ أَخَذَتِ ٱلْأَرْضُ زُخْرُفَهَا وَٱزَّيَّنَتْ وَظَنَّ أَهْلُهَآ أَنَّهُمْ قَـٰدِرُونَ عَلَيْهَآ أَتَىٰهَآ أَمْرُنَا لَيْلًا أَوْ نَهَارًا فَجَعَلْنَـٰهَا حَصِيدًا كَأَن لَّمْ تَغْنَ بِٱلْأَمْسِ ۚ كَذَٰلِكَ نُفَصِّلُ ٱلْـَٔايَـٰتِ لِقَوْمٍ يَتَفَكَّرُونَ

0:00 / 1:08

Switch while it plays to hear the same instant, before and after. Headphones help. SAWT v4 paper

On 120 archive recitations

Low-bitrate uploads, tape and halls, scored against the canonical text with Zipformer Quran as the recogniser. V3.1 is the lab's internal predecessor, never released.

MeasureSourceV3.1SAWT v4
Phoneme error vs the text0.06180.08760.0713
With the recitation mode0.0684
Speaker consistency0.9230.941
Floor in pauses−69.6 dBFS

Against V3.1, phoneme error falls by 19% and predicted quality (DistillMOS) rises by 0.123. On heavily damaged synthetic recitation, where a clean copy exists, phoneme error falls from 0.152 at the input to 0.078.

What it still gets wrong

The source recordings keep the lowest phoneme error. Aligning every output against the text shows where the added error comes from: consonant deletions rise from 25 to 62, while vowel and length errors do not grow. Ten recordings carry 77% of it, and none of them was flagged by the checks. By ear, the affected consonants are rendered weakly, sometimes not at all.

The optional recitation mode answers part of this. It draws three candidates per window and keeps the one whose phonemes drift least from the input, without any transcript. On the archive it lowers phoneme error from 0.0713 to 0.0684; it restores about 3 seconds of audio per second instead of 10 to 16. It can rank what the generator offers; it cannot guarantee that one of the three is right.

Beyond recitation

SAWT was trained on about 900 hours of clean speech in 35 languages, so the same model restores any voice.

  • Reverberation. At a 3 s room and 20 dB noise, word error is 0.520 for SAWT, 0.815 for Sidon and 0.938 for V3.1: the lowest of the three in all four conditions where all three were measured. That is still half the words wrong; such output needs review.
  • LibriTTS-R, 200 utterances, against Google's Miipher-1. The recognised words change less (drift 0.042 against 0.068), the voice is closer (0.988 against 0.962), and fewer frames drop out (0.67% against 1.90%). Word error against the reference text is 0.077 against 0.082, not a clear difference, and the unprocessed input reads 0.066.
  • Miipher-2's own 16 demonstration clips. Miipher-2 leads, by 0.53 in DistillMOS. The clips are published side by side at comparison.quranlab.ai.

A quality score is not a fidelity check

Across the reverberation sweep, DistillMOS stays between 3.8 and 4.3 while word error reaches 0.96. The predictor also listens at 16 kHz, so the 8 to 24 kHz band a 48 kHz restorer rebuilds is outside what it can hear. SAWT is reported against quality, words, phonemes, voice and dropouts together, including the comparisons it loses.

Cost and licence

Training ran on eight H100 GPUs for 150,000 steps; rented GPU time for the whole programme cost about US$1,200. On one RTX 4090, batched inference restores 10 to 16 seconds of audio per second, so 1,000 hours take roughly 63 to 100 GPU hours.

The weights and code are released under the SAWT License 2.0, as a public trust for non-commercial use. They may not be sold, restored audio and anything trained on it carry the same terms, and the model may not be used to make a recording say what was not said. Archives should keep the original and label what was restored.

SAWT v448 kHz speech restoration · SAWT License 2.0 Model card