Ghana AudioDiT

Text-to-speech for 43 Ghanaian languages and Ghanaian English, fine-tuned from LongCat-AudioDiT-1B. It speaks without reference audio: you give it text, it chooses a voice (reproducible with a seed).

Supported by Ghana NLP

Try it in the browser · Code, server and training scripts

Listen

Held-out sentences (never seen in training) unless noted, each with two seeds (two voices).

Language Text Voice 1 Voice 2
Asante Twi Ɛno na ɛyɛ aba ahodoɔ nyinaa mu ketewa deɛ, nanso ɛnyini a, ɛne ɛfan nyinaa mu kɛseɛ,
Akuapem Twi Na afei, me mma, muntie me; monyɛ aso mma nea meka.
Fante Na mesee hom dɛ, Hom edu Amorifo nkokodo nkurow a EWURADZE hɛn Nyankopɔn dze rema hɛn no ho.
Ewe Gbe ma gbe la, woaɖe gbe ɖe edzi abe atsiaƒu ene.
Dagbani Zaŋmi dam din kpɛma n-ti ninvuɣ’ so ŋun yɛn kpi, ka zaŋ wain ti ninvuɣ’ shɛb’ bɛn be nandahima pam ni,
Gonja Isɔɔ ka lar na be kaman nɛ mo sipo malɛ pɛ mbe kenaŋkuŋ to m bɛ mo so m ba lar.
Dangme Ye nyɛmimɛ, nyɛɛ hyɛ bɔ nɛ i suɔ nyɛ ha! Nyɛɛ hyɛ bɔ nɛ ye hɛ ngɛ jae ngɛ nyɛ he ha! Nyɛ lɛ nyɛ haa mi bua jɔmi;
Hausa Wani yakan ɗaukaka wata rana fiye da sauran ranaku, wani kuwa duk ɗaya ne a wurinsa.
English (Ghanaian) Good evening, and welcome to the news from Accra. Here are the top stories of the day. (written for this page)
Twi + English (code-switching) Ɛnnɛ, Microsoft de Abilities for Jobs Program yi baeɛ, na ɛsɛ sɛ yɛn mmeranteɛ ne mmabaa bɔ mmɔden fa so. (written for this page)

Quick start

pip install git+https://github.com/GhanaOpenAI/ghana-audiodit
from ghana_audiodit import GhanaTTS

tts = GhanaTTS.from_pretrained("ghanaopenai/ghana-audiodit")   # ~6 GB download; ~4 GB VRAM

out = tts.synthesize("Akwaaba! Wo ho te sɛn ɛnnɛ?", language="Asante_Twi_twi")
out.save("twi.wav")
print(out.seed)                     # the voice used; pass seed=... to get it again

tts.synthesize("Woezɔ! Aleke nèfɔ ŋdi sia?", language="Ewe_ewe", seed=42).save("ewe.wav")

No voice cloning. The model was trained mostly without voice prompts (85 % of examples), and prompted generation, whether from your own recordings or from fixed reference speakers, was not reliable enough to ship. It is used without reference audio: the voice comes from the seed.

Write text in the normal spelling of the language. The package converts it to the africa-g2p universal spelling the model was trained on, keeping English words as written, so code-switched text works too:

tts.synthesize("Ɛnnɛ, Microsoft de Abilities for Jobs Program yi baeɛ.", language="Asante_Twi_twi")

language accepts a key ("Asante_Twi_twi"), a name ("Ewe") or an ISO code ("dag"). Long text is split into sentence chunks of up to ~18 s of speech, all generated with the same seed; the voice usually stays close, but can drift between chunks of a long passage. Options: seed, steps (16; more is slower and slightly cleaner), cfg_strength (4.0), speed (1.0).

Model size and hardware

Parameters 1.42 B: diffusion transformer 982 M, UMT5-base text encoder 282 M, Wav-VAE 156 M
Download 5.7 GB (fp32 weights)
GPU memory ~4 GB (default: transformer in bfloat16, as in training): 3.2 GB weights, 3.9 GB peak for an 18 s passage. dtype="float32": ~6 GB. A 6 GB+ GPU is enough (T4, L4, A10G, RTX 3060 and up; on T4, which lacks bfloat16, use dtype="float32")
Speed H200, bfloat16, 16 steps: 18 s of speech in about 4 s; a short sentence in about 1–2 s
Output 24 kHz mono

In our check, bfloat16 and float32 were equally intelligible (omniASR CER within 0.3 points). CPU inference works but is slow.

Serving

This is a diffusion (flow-matching) model: it refines the whole utterance in 16 parallel denoising steps rather than generating token by token, so LLM servers such as vLLM do not apply. The GitHub repository has a ready API server (FastAPI; the one behind the demo), also packaged as a Docker image that runs on any machine with an NVIDIA GPU (your own server, any cloud VM, RunPod, Modal, Kubernetes):

docker run --gpus all -p 8000:8000 ghcr.io/ghanaopenai/ghana-audiodit:latest   # model included, runs offline

pip install "ghana-audiodit[server] @ git+https://github.com/GhanaOpenAI/ghana-audiodit"
uvicorn ghana_audiodit.server:app --host 0.0.0.0 --port 8000                    # or without Docker
curl -X POST https://<your-endpoint>/synthesize -F "text=Akwaaba! Wo ho te sɛn?" \
     -F language=Asante_Twi_twi -F seed=42 -o out.wav

Without the package

The weights load with plain transformers (trust_remote_code=True), but you then have to reproduce the text preparation yourself: universal spelling via africa-g2p (with pyspellchecker installed), the model's text normalisation, and a duration in latent frames (rates.json gives frames per character per language). The package does all of this; see ghana_audiodit/tts.py.

from transformers import AutoModel, AutoTokenizer
model = AutoModel.from_pretrained("ghanaopenai/ghana-audiodit", trust_remote_code=True).cuda().eval()
tokenizer = AutoTokenizer.from_pretrained(model.config.text_encoder_model)

Languages

Language language= ISO 639-3
Akuapem Twi Akuapem_Twi_twi twi
Anyin Anyin_any any
Asante Twi Asante_Twi_twi twi
Avatime Avatime_avn avn
Bimoba Bimoba_bim bim
Bissa Bissa_bib bib
Buli Buli_bwu bwu
Chumburung Chumburung_ncu ncu
Dagaare Dagaare_dga dga
Dagbani Dagbani_dag dag
Dangme Dangme_ada ada
Deg Deg_mzw mzw
English (Ghanaian) English_eng eng
Ewe Ewe_ewe ewe
Fante Fante_fat fat
Fulfulde (Maasina) Fulfulde_Maasina_ffm ffm
Gikyode Gikyode_acd acd
Gonja Gonja_gjn gjn
Hausa Hausa_hau hau
Kabiye Kabiye_kbp kbp
Kasem Kasem_xsm xsm
Konkomba Konkomba_xon xon
Konni Konni_kma kma
Kusaal Kusaal_kus kus
Lelemi Lelemi_lef lef
Mampruli Mampruli_maw maw
Nawuri Nawuri_naw naw
Ninkare (Gurenɛ) Ninkare_gur gur
Nkonya Nkonya_nko nko
Ntcham (Bassar) Bassar_Ntcham_bud bud
Ntrubo Ntrubo_ntr ntr
Nzema Nzema_nzi nzi
Paasaal Paasaal_sig sig
Sehwi Sehwi_sfw sfw
Sekpele Sekpele_lip lip
Selee Selee_snw snw
Sisaala (Tumulung) Sisaala_Tumulung_sil sil
Siwu Siwu_akp akp
Southern Birifor Birifor_Southern_biv biv
Tampulma Tampulma_tpm tpm
Tem Tem_kdh kdh
Tuwuli Tuwuli_bov bov
Vagla Vagla_vag vag

Evaluation

Character error rate (CER) of omniASR CTC-300M transcribing the model's speech: 12 held-out sentences (Asante Twi, Ewe and Dagbani, 4 each), generated without a voice prompt with 3 seeds each (36 clips, three different voices per sentence), averaged. Lower is better.

The model reads universal spelling, and omniASR writes what it hears in normal spelling, so both the transcription and the reference are converted to universal spelling before comparing; this ignores spelling-only differences such as ɔ vs o, which universal spelling merges anyway. The CER on normal spelling, without conversion, is shown for reference.

CER (universal spelling) CER (normal spelling)
Real recordings (omniASR's own error rate on the human speech) 19.8 % 22.2 %
LongCat-AudioDiT-1B (base) 35.5 % 43.3 %
This model 14.6 % 21.3 %

By language (universal): Asante Twi 10.8 %, Dagbani 16.2 %, Ewe 16.8 %. The model's speech is at least as intelligible to omniASR as the real recordings. The set is small, so treat differences of a couple of points as noise. Flow-matching validation loss fell from 1.208 (base) to 0.933 without a prompt.

Training

  • Data: about 10 hours per language from ghananlpcommunity/ghana-speech (Hausa: 3 hours) — 199k clips, 407 hours, 43 languages. Most languages are read Bible text; English is conversational Ghanaian English.
  • Text: transcripts converted to africa-g2p universal spelling (English left as is).
  • Audio: resampled to 24 kHz and RMS-normalised to −23 dBFS. Many recordings were mastered loud (−13 to −15 dBFS), which made the base model's latents several times their normal size; normalising fixed it.
  • Method: conditional flow matching with LoRA (rank 64, attention and feed-forward) plus full training of the small embedding, AdaLN and output layers (136M trainable parameters). 85 % of examples had no voice prompt, so no-prompt generation is trained directly; text and prompt were dropped together 10 % of the time for classifier-free guidance, built the same way as inference builds its unconditional input.
  • Schedule: AdamW (β 0.9/0.95), learning rate 1e-4 with 500 warm-up steps and cosine decay over 40,000 steps, batch 64, one H200. This checkpoint is step 28,000 (about 9.6 epochs), chosen for the lowest no-prompt CER among checkpoints scored every 4,000 steps. Later checkpoints had slightly lower validation loss but were less intelligible (mainly in Ewe).

Fine-tuning

lora/ holds the adapter and the fully trained layers of this checkpoint, so you can continue training (for a new language, domain or voice) instead of starting from the base model. The training pipeline (latent caching, universal-spelling manifests, the trainer, side CER evaluation) is in GhanaOpenAI/ghana-audiodit under training/; see training/README.md.

Files

Path Contents
model.safetensors, config.json Merged model: base + this fine-tune (fp32)
configuration_audiodit.py, modeling_audiodit.py Model code for trust_remote_code
rates.json Speaking rate per language (latent frames per character)
lora/ LoRA adapter + fully trained layers, for further fine-tuning
samples/ The audio on this page

Limitations

  • Most training text is Bible readings, so the reading style leans formal.
  • Write numbers as words: clips containing digits were excluded from training.
  • No voice cloning: prompted generation was not reliable enough to ship (see above).
  • The voice comes from the seed: the same seed and text give the same voice, but one seed is not guaranteed to be the same voice across different texts, and long passages can drift between chunks.
  • English words inside Ghanaian text are pronounced from their English spelling, which the model saw mostly in the English portion of the data.
  • Do not use this model to imitate a real person without their consent.

License

The weights are released under CC-BY-NC-4.0, following the training data. The base model LongCat-AudioDiT and its code are MIT-licensed (LICENSE-longcat-audiodit).

Acknowledgements

Built by Ghana Open AI, supported by Ghana NLP. Thanks to Meituan for LongCat-AudioDiT, the africa-g2p project for the universal spelling, Meta for Omnilingual ASR, and the zjubinchen fork whose LoRA trainer this training code started from.

Downloads last month
29
Safetensors
Model size
1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ghananlpcommunity/ghana-audiodit

Finetuned
(13)
this model

Dataset used to train ghananlpcommunity/ghana-audiodit