Instructions to use ghananlpcommunity/ghana-audiodit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ghananlpcommunity/ghana-audiodit with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-to-speech", model="ghananlpcommunity/ghana-audiodit", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("ghananlpcommunity/ghana-audiodit", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Ghana AudioDiT
Text-to-speech for 43 Ghanaian languages and Ghanaian English, fine-tuned from LongCat-AudioDiT-1B. It speaks without reference audio: you give it text, it chooses a voice (reproducible with a seed).
Supported by Ghana NLP
Try it in the browser · Code, server and training scripts
Listen
Held-out sentences (never seen in training) unless noted, each with two seeds (two voices).
| Language | Text | Voice 1 | Voice 2 |
|---|---|---|---|
| Asante Twi | Ɛno na ɛyɛ aba ahodoɔ nyinaa mu ketewa deɛ, nanso ɛnyini a, ɛne ɛfan nyinaa mu kɛseɛ, | ||
| Akuapem Twi | Na afei, me mma, muntie me; monyɛ aso mma nea meka. | ||
| Fante | Na mesee hom dɛ, Hom edu Amorifo nkokodo nkurow a EWURADZE hɛn Nyankopɔn dze rema hɛn no ho. | ||
| Ewe | Gbe ma gbe la, woaɖe gbe ɖe edzi abe atsiaƒu ene. | ||
| Dagbani | Zaŋmi dam din kpɛma n-ti ninvuɣ’ so ŋun yɛn kpi, ka zaŋ wain ti ninvuɣ’ shɛb’ bɛn be nandahima pam ni, | ||
| Gonja | Isɔɔ ka lar na be kaman nɛ mo sipo malɛ pɛ mbe kenaŋkuŋ to m bɛ mo so m ba lar. | ||
| Dangme | Ye nyɛmimɛ, nyɛɛ hyɛ bɔ nɛ i suɔ nyɛ ha! Nyɛɛ hyɛ bɔ nɛ ye hɛ ngɛ jae ngɛ nyɛ he ha! Nyɛ lɛ nyɛ haa mi bua jɔmi; | ||
| Hausa | Wani yakan ɗaukaka wata rana fiye da sauran ranaku, wani kuwa duk ɗaya ne a wurinsa. | ||
| English (Ghanaian) | Good evening, and welcome to the news from Accra. Here are the top stories of the day. (written for this page) | ||
| Twi + English (code-switching) | Ɛnnɛ, Microsoft de Abilities for Jobs Program yi baeɛ, na ɛsɛ sɛ yɛn mmeranteɛ ne mmabaa bɔ mmɔden fa so. (written for this page) |
Quick start
pip install git+https://github.com/GhanaOpenAI/ghana-audiodit
from ghana_audiodit import GhanaTTS
tts = GhanaTTS.from_pretrained("ghanaopenai/ghana-audiodit") # ~6 GB download; ~4 GB VRAM
out = tts.synthesize("Akwaaba! Wo ho te sɛn ɛnnɛ?", language="Asante_Twi_twi")
out.save("twi.wav")
print(out.seed) # the voice used; pass seed=... to get it again
tts.synthesize("Woezɔ! Aleke nèfɔ ŋdi sia?", language="Ewe_ewe", seed=42).save("ewe.wav")
No voice cloning. The model was trained mostly without voice prompts (85 % of examples), and prompted generation, whether from your own recordings or from fixed reference speakers, was not reliable enough to ship. It is used without reference audio: the voice comes from the seed.
Write text in the normal spelling of the language. The package converts it to the africa-g2p universal spelling the model was trained on, keeping English words as written, so code-switched text works too:
tts.synthesize("Ɛnnɛ, Microsoft de Abilities for Jobs Program yi baeɛ.", language="Asante_Twi_twi")
language accepts a key ("Asante_Twi_twi"), a name ("Ewe") or an ISO code ("dag").
Long text is split into sentence chunks of up to ~18 s of speech, all generated with the same
seed; the voice usually stays close, but can drift between chunks of a long passage. Options:
seed, steps (16; more is slower and slightly cleaner), cfg_strength (4.0), speed (1.0).
Model size and hardware
| Parameters | 1.42 B: diffusion transformer 982 M, UMT5-base text encoder 282 M, Wav-VAE 156 M |
| Download | 5.7 GB (fp32 weights) |
| GPU memory | ~4 GB (default: transformer in bfloat16, as in training): 3.2 GB weights, 3.9 GB peak for an 18 s passage. dtype="float32": ~6 GB. A 6 GB+ GPU is enough (T4, L4, A10G, RTX 3060 and up; on T4, which lacks bfloat16, use dtype="float32") |
| Speed | H200, bfloat16, 16 steps: 18 s of speech in about 4 s; a short sentence in about 1–2 s |
| Output | 24 kHz mono |
In our check, bfloat16 and float32 were equally intelligible (omniASR CER within 0.3 points). CPU inference works but is slow.
Serving
This is a diffusion (flow-matching) model: it refines the whole utterance in 16 parallel denoising steps rather than generating token by token, so LLM servers such as vLLM do not apply. The GitHub repository has a ready API server (FastAPI; the one behind the demo), also packaged as a Docker image that runs on any machine with an NVIDIA GPU (your own server, any cloud VM, RunPod, Modal, Kubernetes):
docker run --gpus all -p 8000:8000 ghcr.io/ghanaopenai/ghana-audiodit:latest # model included, runs offline
pip install "ghana-audiodit[server] @ git+https://github.com/GhanaOpenAI/ghana-audiodit"
uvicorn ghana_audiodit.server:app --host 0.0.0.0 --port 8000 # or without Docker
curl -X POST https://<your-endpoint>/synthesize -F "text=Akwaaba! Wo ho te sɛn?" \
-F language=Asante_Twi_twi -F seed=42 -o out.wav
Without the package
The weights load with plain transformers (trust_remote_code=True), but you then have to
reproduce the text preparation yourself: universal spelling via africa-g2p (with
pyspellchecker installed), the model's text normalisation, and a duration in latent frames (rates.json gives frames per character per
language). The package does all of this; see ghana_audiodit/tts.py.
from transformers import AutoModel, AutoTokenizer
model = AutoModel.from_pretrained("ghanaopenai/ghana-audiodit", trust_remote_code=True).cuda().eval()
tokenizer = AutoTokenizer.from_pretrained(model.config.text_encoder_model)
Languages
| Language | language= |
ISO 639-3 |
|---|---|---|
| Akuapem Twi | Akuapem_Twi_twi |
twi |
| Anyin | Anyin_any |
any |
| Asante Twi | Asante_Twi_twi |
twi |
| Avatime | Avatime_avn |
avn |
| Bimoba | Bimoba_bim |
bim |
| Bissa | Bissa_bib |
bib |
| Buli | Buli_bwu |
bwu |
| Chumburung | Chumburung_ncu |
ncu |
| Dagaare | Dagaare_dga |
dga |
| Dagbani | Dagbani_dag |
dag |
| Dangme | Dangme_ada |
ada |
| Deg | Deg_mzw |
mzw |
| English (Ghanaian) | English_eng |
eng |
| Ewe | Ewe_ewe |
ewe |
| Fante | Fante_fat |
fat |
| Fulfulde (Maasina) | Fulfulde_Maasina_ffm |
ffm |
| Gikyode | Gikyode_acd |
acd |
| Gonja | Gonja_gjn |
gjn |
| Hausa | Hausa_hau |
hau |
| Kabiye | Kabiye_kbp |
kbp |
| Kasem | Kasem_xsm |
xsm |
| Konkomba | Konkomba_xon |
xon |
| Konni | Konni_kma |
kma |
| Kusaal | Kusaal_kus |
kus |
| Lelemi | Lelemi_lef |
lef |
| Mampruli | Mampruli_maw |
maw |
| Nawuri | Nawuri_naw |
naw |
| Ninkare (Gurenɛ) | Ninkare_gur |
gur |
| Nkonya | Nkonya_nko |
nko |
| Ntcham (Bassar) | Bassar_Ntcham_bud |
bud |
| Ntrubo | Ntrubo_ntr |
ntr |
| Nzema | Nzema_nzi |
nzi |
| Paasaal | Paasaal_sig |
sig |
| Sehwi | Sehwi_sfw |
sfw |
| Sekpele | Sekpele_lip |
lip |
| Selee | Selee_snw |
snw |
| Sisaala (Tumulung) | Sisaala_Tumulung_sil |
sil |
| Siwu | Siwu_akp |
akp |
| Southern Birifor | Birifor_Southern_biv |
biv |
| Tampulma | Tampulma_tpm |
tpm |
| Tem | Tem_kdh |
kdh |
| Tuwuli | Tuwuli_bov |
bov |
| Vagla | Vagla_vag |
vag |
Evaluation
Character error rate (CER) of omniASR CTC-300M transcribing the model's speech: 12 held-out sentences (Asante Twi, Ewe and Dagbani, 4 each), generated without a voice prompt with 3 seeds each (36 clips, three different voices per sentence), averaged. Lower is better.
The model reads universal spelling, and omniASR writes what it hears in normal spelling, so both the transcription and the reference are converted to universal spelling before comparing; this ignores spelling-only differences such as ɔ vs o, which universal spelling merges anyway. The CER on normal spelling, without conversion, is shown for reference.
| CER (universal spelling) | CER (normal spelling) | |
|---|---|---|
| Real recordings (omniASR's own error rate on the human speech) | 19.8 % | 22.2 % |
| LongCat-AudioDiT-1B (base) | 35.5 % | 43.3 % |
| This model | 14.6 % | 21.3 % |
By language (universal): Asante Twi 10.8 %, Dagbani 16.2 %, Ewe 16.8 %. The model's speech is at least as intelligible to omniASR as the real recordings. The set is small, so treat differences of a couple of points as noise. Flow-matching validation loss fell from 1.208 (base) to 0.933 without a prompt.
Training
- Data: about 10 hours per language from ghananlpcommunity/ghana-speech (Hausa: 3 hours) — 199k clips, 407 hours, 43 languages. Most languages are read Bible text; English is conversational Ghanaian English.
- Text: transcripts converted to africa-g2p universal spelling (English left as is).
- Audio: resampled to 24 kHz and RMS-normalised to −23 dBFS. Many recordings were mastered loud (−13 to −15 dBFS), which made the base model's latents several times their normal size; normalising fixed it.
- Method: conditional flow matching with LoRA (rank 64, attention and feed-forward) plus full training of the small embedding, AdaLN and output layers (136M trainable parameters). 85 % of examples had no voice prompt, so no-prompt generation is trained directly; text and prompt were dropped together 10 % of the time for classifier-free guidance, built the same way as inference builds its unconditional input.
- Schedule: AdamW (β 0.9/0.95), learning rate 1e-4 with 500 warm-up steps and cosine decay over 40,000 steps, batch 64, one H200. This checkpoint is step 28,000 (about 9.6 epochs), chosen for the lowest no-prompt CER among checkpoints scored every 4,000 steps. Later checkpoints had slightly lower validation loss but were less intelligible (mainly in Ewe).
Fine-tuning
lora/ holds the adapter and the fully trained layers of this checkpoint, so you can continue
training (for a new language, domain or voice) instead of starting from the base model. The
training pipeline (latent caching, universal-spelling manifests, the trainer, side CER
evaluation) is in GhanaOpenAI/ghana-audiodit
under training/; see training/README.md.
Files
| Path | Contents |
|---|---|
model.safetensors, config.json |
Merged model: base + this fine-tune (fp32) |
configuration_audiodit.py, modeling_audiodit.py |
Model code for trust_remote_code |
rates.json |
Speaking rate per language (latent frames per character) |
lora/ |
LoRA adapter + fully trained layers, for further fine-tuning |
samples/ |
The audio on this page |
Limitations
- Most training text is Bible readings, so the reading style leans formal.
- Write numbers as words: clips containing digits were excluded from training.
- No voice cloning: prompted generation was not reliable enough to ship (see above).
- The voice comes from the seed: the same seed and text give the same voice, but one seed is not guaranteed to be the same voice across different texts, and long passages can drift between chunks.
- English words inside Ghanaian text are pronounced from their English spelling, which the model saw mostly in the English portion of the data.
- Do not use this model to imitate a real person without their consent.
License
The weights are released under CC-BY-NC-4.0, following the training data. The base model
LongCat-AudioDiT and its code are MIT-licensed (LICENSE-longcat-audiodit).
Acknowledgements
Built by Ghana Open AI, supported by Ghana NLP. Thanks to Meituan for LongCat-AudioDiT, the africa-g2p project for the universal spelling, Meta for Omnilingual ASR, and the zjubinchen fork whose LoRA trainer this training code started from.
- Downloads last month
- 29
Model tree for ghananlpcommunity/ghana-audiodit
Base model
meituan-longcat/LongCat-AudioDiT-1B