Super Sonique / listening room

Less noise.
More voice.

Super Sonique meets Sidon. Listen to 40 real-noise clips and 10 artificially noised clean clips, side by side with their matching regular Auphonic references.

Non-Studio EDM · Sidon100 step 110,000 · EMASampling Heun 8 · 15 evaluationsAudio 48 kHz · full clipsReferences Regular Auphonic · non-Studio
How to read the scores

MOS score recovery.

Recovery = 100 × (denoised MOS − noisy MOS) / (reference MOS − noisy MOS). Scores are predicted by Microsoft DNSMOS P.835, not human ratings. Negative recovery means degradation; above 100% means a score above the reference. A clean–noisy gap below 0.1 is marked n/a. Summary percentages use cohort mean scores, not the mean of clip percentages.

Audio metrics · overall scores and methodology

NISQA and DNSMOS predict overall speech quality (higher is better). SpkSim measures speaker-embedding cosine similarity to the actual noisy input, using WavLM Base Plus SV—not similarity to the clean reference. The input’s self-similarity is shown as a dash. These are the restoration metrics in Sidon §4.2; the paper also uses MMS word error rate for English and character error rate for multilingual speech.

ASR-reference WER, not verified WER: multilingual Parakeet TDT 0.6B v3 transcribes each regular clean reference and each comparison waveform with automatic language detection. Word error rate measures disagreement against that automatic reference text, which may itself contain errors. This replaces the paper’s MMS recognizer at your request; lower is better. The clean reference’s zero is a self-comparison, not a transcription-accuracy result. Empty reference text is marked n/a.

This panel includes multiple languages. Interpret learned quality scores cautiously where the language or recording conditions differ from the scoring model’s training coverage; NISQA is a prediction, not a language-independent human rating.

Scoring protocol and comparison limits

All tables below score uncompressed, pre-playback float waveforms. “Actual input” is what the denoisers received; on real recordings, it is not the codec reconstruction in the noisy player. The MOS recovery cards and sort order retain their original scoring basis (public MP3 for real clips; float WAV for artificial clips), so their MOS values can differ from the DNSMOS table.

Cohort tables summarize all 40 real or 10 artificial clips and are not changed by source filters. NISQA, DNSMOS and SpkSim are arithmetic means; ASR-reference WER divides total substitutions, deletions and insertions by total reference words, not a mean of clip percentages. These metrics are not a reproduction of Sidon’s published benchmark. The paper does not specify exact NISQA or DNSMOS checkpoint revisions, and our ASR model and automatic reference text differ from its protocol. Reference scores are context, not an upper bound; higher SpkSim does not by itself imply better denoising.

Ground truthMatching regular Auphonic reference
NoisyReal codec baseline / actual synthetic input
Super SoniqueNon-Studio EDM · Sidon100 · EMA step 110,000
Sidon v0.1Original Sidon, not DialogueSidon

Loading the listening room…

Real recordings

More real-noise clips

The remaining test pairs, ordered by Super Sonique MOS recovery. The noisy listening baseline is a KVAE reconstruction, as in the previous site; both denoisers receive the original noisy waveform. MOS is calculated on the public MP3s.

Controlled corruption

Artificial noise · +5 dB

Ten held-out regular clean speech clips, mixed with recorded noise using the Sidon augmentation pipeline. Noise comes from the training noise bank: this tests held-out speech under a matched noise distribution, not unseen noise. No reverb, bandwidth, codec, clipping augmentation or packet loss is enabled. The mixer’s final saturation still applies; +5 dB is measured before saturation.

The noisy player is the exact saved corrupted input from the previous comparison. Following the evaluation guide, MOS uses pre-playback float WAVs for this section. Expand codec diagnostics to hear noisy and clean KVAE reconstructions. Best Super Sonique MOS recovery first.

ASR-reference WER is diagnostic: the artificial-noise panel has unverified language labels and may include languages outside Parakeet's supported set.