Fasih-TTS-V1 · فَصِيح
Modern Standard Arabic (Fusha) text-to-speech with a professional male voice.
Fine-tuned from Coqui XTTS v2 · 467M parameters · 24 kHz · streaming-ready
Fasih (فَصِيح, "eloquent") is a single-speaker Modern Standard Arabic TTS model with a news-anchor-style male voice. It is the voice of Muslim (مسلم), a religious Q&A assistant. On held-out text its synthesized speech is as intelligible to an ASR judge as the original human recordings, and it does not loop, skip or cut off early.
Listen
| Greeting | Fiqh explanation |
|---|---|
| السَّلَامُ عَلَيْكُمْ... أَنَا مُسْلِم، مُسَاعِدُكَ الصَّوْتِيُّ | الْوُضُوءُ شَرْطٌ لِصِحَّةِ الصَّلَاةِ... |
Try your own text in the live demo. Undiacritized input works too.
At a glance
![]() |
![]() |
![]() |
| Human-level intelligibility. 1.3% CER, below the 1.8% ASR floor measured on the original human recordings. | Stable generation. 0 loops, skips or early cut-offs in 24 stress generations, which are the usual failure modes of XTTS. | Real-time. RTF ≈ 0.60 and ≈ 675 ms to first streamed audio on one RTX 2080 Ti (FP32). |
- Correct iʿrāb from bare text. Training text was fully diacritized, and a built-in CATT diacritizer adds tashkīl to undiacritized input, including case endings.
- Text front-end for production. Numbers are expanded to words, a sacred-term lexicon fixes religious vocabulary, and long passages are chunked automatically. Raw assistant text goes in; speech comes out.
Evaluation
Intelligibility (CER)
Character error rate (CER) between the intended text and a Whisper-large-v3 transcription of the synthesized audio. Both sides are diacritics-stripped and orthography-normalized. The human originals are scored the same way and set the ASR floor.
| Test set | Clips | Mean CER | Worst CER |
|---|---|---|---|
| Varied MSA sentences | 8 | 1.3% | 2.2% |
| Same sentence ×4 (variance) | 4 | 2.0% | 2.0% |
| Long text (auto-chunked) | 2 | 0.8% | 0.9% |
| Hard stress (numbers, lists, terms) | 6 | 2.1% | 8.2% |
| Human originals (ASR floor) | 8 | 1.8% | 4.8% |
SILMA open-source Arabic TTS benchmark
SILMA's benchmark (MSA, 10 fixed sentences), scored by two independent ASR judges, Whisper-large-v3 and NVIDIA NeMo Arabic FastConformer, plus UTMOS as a naturalness proxy.
| Model | WER · Whisper ↓ | WER · NeMo ↓ | UTMOS ↑ |
|---|---|---|---|
| Fasih-TTS-V1 | 6.5 | 2.5 | 3.16 |
| XTTS v2 (base) | 10.3 | 2.5 | 2.99 |
| chatterbox | 12.8 | 5.4 | 3.20 |
| silma_tts | 11.1 | 5.8 | 3.15 |
| omnivoice | 15.3 | 7.3 | 3.62 |
| habibi_specialized | 21.9 | 23.3 | 2.33 |
Fasih has the lowest WER under both judges; with NeMo it ties the base XTTS v2. On naturalness (UTMOS) it ranks third. The model is tuned for pronunciation accuracy, which matters most for religious content. WER measures intelligibility, not naturalness; SILMA's own benchmark is a human listening comparison.
Per-clip results and all audio are in
NightPrince/Fasih-TTS-Benchmark.
Reproduce with scripts/silma_compare.py (Whisper), scripts/nemo_compare.py (NeMo) and
scripts/utmos_compare.py (UTMOS) in the GitHub repo.
How it works
raw Arabic text ─▶ normalize ─▶ numbers → words ─▶ CATT diacritization ─▶ sacred-term lexicon ─▶ ≤160-char chunks
│
24 kHz speech ◀─ HiFi-GAN decoder ◀─ GPT (fine-tuned) ◀─ shipped speaker latents ◀┘
Fasih has 466.9M parameters. Only the XTTS GPT (441.0M) was fine-tuned. The HiFi-GAN decoder and the DVAE stayed frozen, which keeps
the vocoder intact while the GPT learns Fusha prosody and pronunciation. The voice ships as
precomputed conditioning latents (speaker_latents.pt), so no reference audio is needed.
Quick start
pip install "coqui-tts==0.27.5"
import torch
from huggingface_hub import snapshot_download
from TTS.tts.configs.xtts_config import XttsConfig
from TTS.tts.models.xtts import Xtts
path = snapshot_download("NightPrince/Fasih-TTS-V1")
config = XttsConfig()
config.load_json(f"{path}/config.json")
model = Xtts.init_from_config(config)
model.load_checkpoint(config, checkpoint_path=f"{path}/model.pth",
vocab_path=f"{path}/vocab.json", use_deepspeed=False)
model.cuda().eval()
# The Fasih voice ships as precomputed latents. No reference audio needed.
lat = torch.load(f"{path}/speaker_latents.pt", map_location="cuda")
out = model.inference("السَّلَامُ عَلَيْكُمْ وَرَحْمَةُ اللَّهِ", "ar",
lat["gpt_cond_latent"], lat["speaker_embedding"],
temperature=0.65, repetition_penalty=2.0, enable_text_splitting=False)
# out["wav"] → 24 kHz mono waveform
For correct iʿrāb, feed diacritized Fusha, or use the text front-end from the
GitHub repo (TextPipeline.prepare_chunks). It
diacritizes, expands numbers, applies the lexicon and chunks raw text for you.
Intended use
In scope: reading MSA / Fusha explanatory religious and educational content that a qualified person has written or reviewed.
Out of scope / prohibited
- Qur'anic recitation. Recitation requires tajwīd and human reciters. Route āyāt to real recordings.
- Autonomous religious rulings. The model only voices text and does not check whether it is correct.
- Impersonation or misinformation. Do not use this voice to synthesize false statements.
Training
| Base model | coqui/XTTS-v2 |
| Parameters | 466.9M total: GPT 441.0M (fine-tuned) + HiFi-GAN decoder 25.9M (17.8M waveform decoder + 8.0M speaker encoder, frozen) |
| Data | NightPrince/Arabic-professional-original-voice: 1297 clips, ~2.4 h, one male speaker |
| Text | Fully diacritized. Plain transcripts were diacritized with CATT, which was checked to be close to human gold labels |
| Trained parts | GPT only (HiFi-GAN decoder and DVAE frozen) |
| Optimizer | AdamW, LR 5e-6, betas (0.9, 0.96), weight decay 0.01 |
| Batch | 1 × gradient accumulation 24, gradient checkpointing |
| Precision / hardware | FP32 on a single RTX 2080 Ti (Turing has no BF16, and XTTS's GPT is unstable under FP16 autocast) |
| Best validation loss | 2.622 |
Checkpoint format. model.pth holds inference-only weights (1.87 GB, FP32, the same size
as the base XTTS-v2). The original 5.6 GB training checkpoint also contained the AdamW optimizer
state and the frozen DVAE. Neither is used at inference, so both have been removed. All 963 model
tensors and the generated audio are bit-identical to the original. To fine-tune further, point
Coqui's XTTS recipe xtts_checkpoint at this model.pth and take dvae.pth / mel_stats.pth
from coqui/XTTS-v2.
Limitations
- Correct iʿrāb needs diacritized text. The front-end adds diacritics automatically.
- Number gender agreement (
خمسةvsخمس) is not always correct. - The source audio is 128 kbps MP3, which limits fidelity.
- Training used ~2.4 h from a single speaker. The auto-diacritization of 371 training clips is about 95%+ accurate and was not fully human-verified.
In the wild
- Arabic TTS Arena: Fasih-TTS-V1 is live for blind A/B voting against other Arabic TTS models (PR #6, merged).
- Muslim (مسلم): a deployed Arabic voice AI platform for grounded Islamic knowledge. Fasih is its voice.
License
Fine-tuned from Coqui XTTS v2 under the Coqui Public Model License (CPML): non-commercial use with attribution required, and derivatives inherit these terms. The diacritizer is CATT (MIT).
Copyright
Copyright 2026 Yahya Elnawasany (NightPrince). The Fasih-TTS-V1 model, its voice and generated
audio, and the "Fasih / فَصِيح" name and branding are copyright the author. The model is distributed
under the Coqui Public Model License (non-commercial, attribution); the accompanying code is MIT.
Do not use the model or its outputs to impersonate, misrepresent, or generate misleading religious
content. Full terms are in COPYRIGHT and THIRD_PARTY_NOTICES.md in the repository.
Citation
This model is described in Muslim: A Deployed Arabic Voice AI Platform for Grounded Islamic Knowledge (arXiv:2609.31511).
@software{fasih_tts_v1_2026,
author = {Yahya Elnawasany (NightPrince)},
title = {Fasih-TTS-V1: Arabic Fusha Professional-Male Text-to-Speech},
year = {2026},
url = {https://github.com/NightPrinceY/Fasih-TTS-V1},
note = {Fine-tuned from Coqui XTTS v2}
}
Source: https://github.com/NightPrinceY/Fasih-TTS-V1 · Author: Yahya Elnawasany (NightPrince) · https://nightprincey.github.io/Portfolio-App/
- Downloads last month
- 158
Model tree for NightPrince/Fasih-TTS-V1
Base model
coqui/XTTS-v2Dataset used to train NightPrince/Fasih-TTS-V1
Space using NightPrince/Fasih-TTS-V1 1
Collection including NightPrince/Fasih-TTS-V1
Paper for NightPrince/Fasih-TTS-V1
Article mentioning NightPrince/Fasih-TTS-V1
Evaluation results
- CER (%) — held-out synthesis vs Whisper-large-v3 on Arabic Professional Original Voiceself-reported1.300
- Human-recording ASR floor (%) on Arabic Professional Original Voiceself-reported1.800


