How the score is calculated. Each synthesised clip is graded two ways and the two are averaged:
Every sample comes from
- Accuracy = 1 − CER (intelligibility). An ASR judge — the best model for that language in the nsanku ASR benchmark — transcribes the clip, and the transcript is compared with the sentence that was synthesised. CER (shown as a percentage, lower is better) is the share of characters wrong; Accuracy is its complement.
- SBS (SpeechBERTScore, acoustic similarity). The clip is compared with a real human recording of the same sentence
(
microsoft/wavlm-large, layer 6). Higher is better. - Composite = (Accuracy + SBS) / 2, higher is better. Example: CER 6.49% and SBS 0.7126 give Accuracy 0.9351 and Composite (0.9351 + 0.7126) / 2 = 0.8238.
Every sample comes from
ghana-speech-eval; a language with several text sources (domains) is sampled evenly across them.
The domain columns show the Composite computed on that source alone.
Global Leaderboard
Per-Language
All Results
Best model per language, ranked by Composite score. The domain columns show the Composite on each text source.
✓ Scored
✗ No valid output
| Model | Reason |
|---|
0 rows
| Language | Model | Track | Mode | Params | Composite | Accuracy | CER (%) | SBS | Clips | Failed |
|---|