Text-to-Speech
Arabic
coqui
tts
arabic
arabic-tts
fusha
msa
xtts
xtts-v2
speech-synthesis
voice-cloning
Eval Results (legacy)

Fasih-TTS-V1 — The voice Muslim answers in

 Fasih-TTS-V1 · فَصِيح

Modern Standard Arabic (Fusha) text-to-speech with a professional male voice.
Fine-tuned from Coqui XTTS v2 · 467M parameters · 24 kHz · streaming-ready

Live demo arXiv GitHub Arabic TTS Arena License

Fasih (فَصِيح, "eloquent") is a single-speaker Modern Standard Arabic TTS model with a news-anchor-style male voice. It is the voice of Muslim (مسلم), a religious Q&A assistant. On held-out text its synthesized speech is as intelligible to an ASR judge as the original human recordings, and it does not loop, skip or cut off early.


Listen

Greeting Fiqh explanation
السَّلَامُ عَلَيْكُمْ... أَنَا مُسْلِم، مُسَاعِدُكَ الصَّوْتِيُّ الْوُضُوءُ شَرْطٌ لِصِحَّةِ الصَّلَاةِ...

Try your own text in the live demo. Undiacritized input works too.


At a glance

1.3% CER vs 1.8% for the human recordings 0 failures in 24 stress generations 675 ms first audio, RTF 0.60, 24 kHz
Human-level intelligibility. 1.3% CER, below the 1.8% ASR floor measured on the original human recordings. Stable generation. 0 loops, skips or early cut-offs in 24 stress generations, which are the usual failure modes of XTTS. Real-time. RTF ≈ 0.60 and ≈ 675 ms to first streamed audio on one RTX 2080 Ti (FP32).
  • Correct iʿrāb from bare text. Training text was fully diacritized, and a built-in CATT diacritizer adds tashkīl to undiacritized input, including case endings.
  • Text front-end for production. Numbers are expanded to words, a sacred-term lexicon fixes religious vocabulary, and long passages are chunked automatically. Raw assistant text goes in; speech comes out.

Evaluation

Intelligibility (CER)

Character error rate by test set

Character error rate (CER) between the intended text and a Whisper-large-v3 transcription of the synthesized audio. Both sides are diacritics-stripped and orthography-normalized. The human originals are scored the same way and set the ASR floor.

Test set Clips Mean CER Worst CER
Varied MSA sentences 8 1.3% 2.2%
Same sentence ×4 (variance) 4 2.0% 2.0%
Long text (auto-chunked) 2 0.8% 0.9%
Hard stress (numbers, lists, terms) 6 2.1% 8.2%
Human originals (ASR floor) 8 1.8% 4.8%

SILMA open-source Arabic TTS benchmark

SILMA benchmark, MSA, WER with Whisper

SILMA's benchmark (MSA, 10 fixed sentences), scored by two independent ASR judges, Whisper-large-v3 and NVIDIA NeMo Arabic FastConformer, plus UTMOS as a naturalness proxy.

Model WER · Whisper ↓ WER · NeMo ↓ UTMOS ↑
Fasih-TTS-V1 6.5 2.5 3.16
XTTS v2 (base) 10.3 2.5 2.99
chatterbox 12.8 5.4 3.20
silma_tts 11.1 5.8 3.15
omnivoice 15.3 7.3 3.62
habibi_specialized 21.9 23.3 2.33

Fasih has the lowest WER under both judges; with NeMo it ties the base XTTS v2. On naturalness (UTMOS) it ranks third. The model is tuned for pronunciation accuracy, which matters most for religious content. WER measures intelligibility, not naturalness; SILMA's own benchmark is a human listening comparison.

Per-clip results and all audio are in NightPrince/Fasih-TTS-Benchmark. Reproduce with scripts/silma_compare.py (Whisper), scripts/nemo_compare.py (NeMo) and scripts/utmos_compare.py (UTMOS) in the GitHub repo.


How it works

raw Arabic text ─▶ normalize ─▶ numbers → words ─▶ CATT diacritization ─▶ sacred-term lexicon ─▶ ≤160-char chunks
                                                                                                       │
                       24 kHz speech ◀─ HiFi-GAN decoder ◀─ GPT (fine-tuned) ◀─ shipped speaker latents ◀┘

Fasih has 466.9M parameters. Only the XTTS GPT (441.0M) was fine-tuned. The HiFi-GAN decoder and the DVAE stayed frozen, which keeps the vocoder intact while the GPT learns Fusha prosody and pronunciation. The voice ships as precomputed conditioning latents (speaker_latents.pt), so no reference audio is needed.


Quick start

pip install "coqui-tts==0.27.5"
import torch
from huggingface_hub import snapshot_download
from TTS.tts.configs.xtts_config import XttsConfig
from TTS.tts.models.xtts import Xtts

path = snapshot_download("NightPrince/Fasih-TTS-V1")
config = XttsConfig()
config.load_json(f"{path}/config.json")
model = Xtts.init_from_config(config)
model.load_checkpoint(config, checkpoint_path=f"{path}/model.pth",
                      vocab_path=f"{path}/vocab.json", use_deepspeed=False)
model.cuda().eval()

# The Fasih voice ships as precomputed latents. No reference audio needed.
lat = torch.load(f"{path}/speaker_latents.pt", map_location="cuda")
out = model.inference("السَّلَامُ عَلَيْكُمْ وَرَحْمَةُ اللَّهِ", "ar",
                      lat["gpt_cond_latent"], lat["speaker_embedding"],
                      temperature=0.65, repetition_penalty=2.0, enable_text_splitting=False)
# out["wav"] → 24 kHz mono waveform

For correct iʿrāb, feed diacritized Fusha, or use the text front-end from the GitHub repo (TextPipeline.prepare_chunks). It diacritizes, expands numbers, applies the lexicon and chunks raw text for you.


Intended use

What Fasih will not do: it never recites. The Qur'an has its reciters.

In scope: reading MSA / Fusha explanatory religious and educational content that a qualified person has written or reviewed.

Out of scope / prohibited

  • Qur'anic recitation. Recitation requires tajwīd and human reciters. Route āyāt to real recordings.
  • Autonomous religious rulings. The model only voices text and does not check whether it is correct.
  • Impersonation or misinformation. Do not use this voice to synthesize false statements.

Training

Base model coqui/XTTS-v2
Parameters 466.9M total: GPT 441.0M (fine-tuned) + HiFi-GAN decoder 25.9M (17.8M waveform decoder + 8.0M speaker encoder, frozen)
Data NightPrince/Arabic-professional-original-voice: 1297 clips, ~2.4 h, one male speaker
Text Fully diacritized. Plain transcripts were diacritized with CATT, which was checked to be close to human gold labels
Trained parts GPT only (HiFi-GAN decoder and DVAE frozen)
Optimizer AdamW, LR 5e-6, betas (0.9, 0.96), weight decay 0.01
Batch 1 × gradient accumulation 24, gradient checkpointing
Precision / hardware FP32 on a single RTX 2080 Ti (Turing has no BF16, and XTTS's GPT is unstable under FP16 autocast)
Best validation loss 2.622

Checkpoint format. model.pth holds inference-only weights (1.87 GB, FP32, the same size as the base XTTS-v2). The original 5.6 GB training checkpoint also contained the AdamW optimizer state and the frozen DVAE. Neither is used at inference, so both have been removed. All 963 model tensors and the generated audio are bit-identical to the original. To fine-tune further, point Coqui's XTTS recipe xtts_checkpoint at this model.pth and take dvae.pth / mel_stats.pth from coqui/XTTS-v2.

Limitations

  • Correct iʿrāb needs diacritized text. The front-end adds diacritics automatically.
  • Number gender agreement (خمسة vs خمس) is not always correct.
  • The source audio is 128 kbps MP3, which limits fidelity.
  • Training used ~2.4 h from a single speaker. The auto-diacritization of 371 training clips is about 95%+ accurate and was not fully human-verified.

In the wild

  • Arabic TTS Arena: Fasih-TTS-V1 is live for blind A/B voting against other Arabic TTS models (PR #6, merged).
  • Muslim (مسلم): a deployed Arabic voice AI platform for grounded Islamic knowledge. Fasih is its voice.

License

Fine-tuned from Coqui XTTS v2 under the Coqui Public Model License (CPML): non-commercial use with attribution required, and derivatives inherit these terms. The diacritizer is CATT (MIT).

Copyright

Copyright 2026 Yahya Elnawasany (NightPrince). The Fasih-TTS-V1 model, its voice and generated audio, and the "Fasih / فَصِيح" name and branding are copyright the author. The model is distributed under the Coqui Public Model License (non-commercial, attribution); the accompanying code is MIT. Do not use the model or its outputs to impersonate, misrepresent, or generate misleading religious content. Full terms are in COPYRIGHT and THIRD_PARTY_NOTICES.md in the repository.

Citation

This model is described in Muslim: A Deployed Arabic Voice AI Platform for Grounded Islamic Knowledge (arXiv:2609.31511).

@software{fasih_tts_v1_2026,
  author = {Yahya Elnawasany (NightPrince)},
  title  = {Fasih-TTS-V1: Arabic Fusha Professional-Male Text-to-Speech},
  year   = {2026},
  url    = {https://github.com/NightPrinceY/Fasih-TTS-V1},
  note   = {Fine-tuned from Coqui XTTS v2}
}

Source: https://github.com/NightPrinceY/Fasih-TTS-V1 · Author: Yahya Elnawasany (NightPrince) · https://nightprincey.github.io/Portfolio-App/

Downloads last month
158
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 2 Ask for provider support

Model tree for NightPrince/Fasih-TTS-V1

Base model

coqui/XTTS-v2
Finetuned
(80)
this model

Dataset used to train NightPrince/Fasih-TTS-V1

Space using NightPrince/Fasih-TTS-V1 1

Collection including NightPrince/Fasih-TTS-V1

Paper for NightPrince/Fasih-TTS-V1

Article mentioning NightPrince/Fasih-TTS-V1

Evaluation results

  • CER (%) — held-out synthesis vs Whisper-large-v3 on Arabic Professional Original Voice
    self-reported
    1.300
  • Human-recording ASR floor (%) on Arabic Professional Original Voice
    self-reported
    1.800