Instructions to use olaverse/lid-lite-608 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- fastText
How to use olaverse/lid-lite-608 with fastText:
from huggingface_hub import hf_hub_download import fasttext model = fasttext.load_model(hf_hub_download("olaverse/lid-lite-608", "model.bin")) - Notebooks
- Google Colab
- Kaggle
lid-lite-608
lid-lite-608 Highlights
lid-lite-608 is a fast, CPU-only language identification model for 608 languages, built African-first, with the following key features:
- 608 languages in 37 MB. About 3x the languages of Meta's fastText LID (218) at 1/32 of its size, and 1/45 the size of GlotLID. Runs at ~6,800 texts/s on a single CPU thread.
- Best-in-class on short text. 0.672 accuracy on 1-5 word inputs across all 608 languages, vs 0.486 for GlotLID and 0.218 for Meta's fastText LID.
- Strong on African languages. Average F1 0.839 across 20 major African languages vs 0.757 for GlotLID, with large gains on Kinyarwanda, Xhosa, Wolof and Lingala.
- Two modes in one model: coverage mode treats every language equally (corpus building, low-resource mining); traffic mode calibrates to real-world language frequency (user input, routing) and reaches 0.954 web-weighted accuracy, the best of all models tested.
- Robust to messy input. Trained with typos, missing diacritics, URLs, emoji and everyday conversational sentences, plus a dedicated noise class for numbers, URLs and code (F1 0.961).
- Commercial-friendly. Trained only on sources that allow commercial use; Apache-2.0.
Model Overview
lid-lite-608 has the following features:
- Type: fastText supervised classifier (char n-grams 1-5 + word bigrams,
dim=128, quantized) - Languages: 608 (567 distinct ISO 639-3 codes, 36 scripts) plus a noise class
zxx_Zxxx - Labels: ISO 639-3 + ISO 15924 script, e.g.
yor_Latn,srp_Cyrl(full list inlabels.json) - Size: 37 MB (
model.ftz) - Training data: 7.86M samples from commercially usable sources (see below)
- Input: any length; 1 word to full documents
For the highest accuracy, especially on short and conversational input, see the neural sibling lid-neural-608 (mmBERT, same 608 languages, labels and modes).
Quickstart
pip install fasttext huggingface_hub
import json, math, fasttext
from huggingface_hub import hf_hub_download
repo = "olaverse/lid-lite-608"
model = fasttext.load_model(hf_hub_download(repo, "model.ftz"))
calib = json.load(open(hf_hub_download(repo, "priors.json")))
traffic_bias = {l: calib["alpha"] * v for l, v in calib["log_prior_ratio"].items()}
def identify(text, mode="coverage"):
text = " ".join(text.split()) # fastText needs a single line
k = 40 if mode == "traffic" else 1
labels, probs = model.predict([text], k=k) # list input: works with NumPy 1 and 2
labels, probs = labels[0], probs[0]
scores = {l[9:]: math.log(max(p, 1e-12)) + (traffic_bias.get(l[9:], 0.0) if mode == "traffic" else 0.0)
for l, p in zip(labels, probs)}
return max(scores, key=scores.get)
identify("Ẹ kú àárọ̀, ṣé dáadáa ni?") # 'yor_Latn'
identify("Habari za asubuhi") # 'swh_Latn'
identify("Bonjour mon ami", mode="traffic") # 'fra_Latn'
With the olaverse library
The olaverse library (v0.4.1+) wraps the model, the priors and both modes:
pip install "olaverse[lid]"
from olaverse import LIDLite608
lid = LIDLite608() # mode="coverage" (default)
lid.predict("Ẹ kú àárọ̀, ṣé dáadáa ni?") # 'yor_Latn'
lid.predict_proba("Habari za asubuhi", top_k=3) # {'swh_Latn': 0.916, 'swc_Latn': 0.084, 'hau_Latn': 0.0}
lid.predict_batch(["Habari za asubuhi", "Mo fẹ́ lọ sí ọjà lónìí"]) # ['swh_Latn', 'yor_Latn']
lid.predict("Bonjour mon ami") # 'dhv_Latn' (coverage: every language equally likely)
LIDLite608(mode="traffic").predict("Bonjour mon ami") # 'fra_Latn' (traffic: leans to common languages)
Newlines are collapsed for you, empty text raises ValueError, and it works with NumPy 2. Its predictions match
the code above exactly (checked on 20,000 held-out texts in both modes).
Switching Between Coverage and Traffic Mode
One model, two ways to read its output:
mode="coverage" (default)
Every language is treated as equally likely. Use it when every language matters, such as mining low-resource text from a web crawl or building corpora. Highest per-language accuracy (macro-F1 0.856).
mode="traffic"
Scores are adjusted by each language's real-world frequency (priors.json), so short or ambiguous
input leans toward the languages that actually dominate everyday traffic. Use it for user input,
search queries, chat and request routing: +2.2 pts web-weighted accuracy and +6.6 pts on 1-5 word
inputs compared with coverage mode.
Choosing Between lid-lite-608 and lid-neural-608
Both models share the same 608 languages, labels and coverage/traffic modes.
| lid-lite-608 | lid-neural-608 | |
|---|---|---|
| Size | 37 MB | 140M parameters (~560 MB) |
| Speed | ~6,800 texts/s on one CPU thread | |
| Held-out test, accuracy | 0.858 | 0.867 |
| 1-5 word inputs, traffic mode (web-weighted) | 0.829 | 0.906 |
| Everyday phrases, traffic mode | 0.838 | 0.943 |
| Best for | Bulk filtering, CPU-only and edge deployments | User input, short text, highest accuracy |
Use both: run lid-lite-608 as a fast first pass and send only short or uncertain inputs to lid-neural-608.
Evaluation
All models scored on the same data. Other models' language codes are mapped to ours and they are scored on their best guess among our 608 languages, so naming differences never count as errors.
Show table
| lid-lite-608 (coverage) | lid-lite-608 (traffic) | GlotLID v3 | Meta fastText LID-218 | |
|---|---|---|---|---|
| Model size | 37 MB | 37 MB | 1,687 MB | 1,176 MB |
| Licence | Apache-2.0 | Apache-2.0 | Apache-2.0 | CC-BY-NC-4.0 |
| Held-out test, accuracy (240k) | 0.858 | 0.843 | 0.783 | 0.280 |
| Held-out test, macro-F1 | 0.856 | 0.843 | 0.795 | 0.182 |
| Held-out test, 1-5 words | 0.672 | 0.656 | 0.486 | 0.218 |
| Web-weighted accuracy | 0.932 | 0.954 | 0.946 | 0.901 |
| Web-weighted, 1-5 words | 0.763 | 0.829 | 0.808 | 0.764 |
| Tatoeba (conversational, 34k) | 0.922 | 0.914 | 0.868 | 0.557 |
| UDHR-LID (291 languages) | 0.910 | 0.907 | 0.908 | 0.567 |
| FLORES+ devtest (202 languages) | 0.934 | 0.928 | 0.957 | 0.874 |
| FLORES+ first 1-5 words | 0.708 | 0.708 | 0.692 | 0.618 |
On the 208 languages Meta's fastText LID covers (table)
| lid-lite-608 (coverage) | lid-lite-608 (traffic) | GlotLID v3 | Meta fastText LID-218 | |
|---|---|---|---|---|
| Held-out test | 0.900 | 0.897 | 0.885 | 0.822 |
| Held-out test, 1-5 words | 0.749 | 0.750 | 0.681 | 0.628 |
| Tatoeba | 0.937 | 0.939 | 0.911 | 0.858 |
| FLORES+ first 1-5 words | 0.732 | 0.735 | 0.724 | 0.684 |
By input length (held-out test, accuracy)
Show table
| Words | lid-lite-608 | GlotLID v3 |
|---|---|---|
| 1-5 | 0.672 | 0.486 |
| 6-20 | 0.871 | 0.796 |
| 21-80 | 0.924 | 0.902 |
| 81-256 | 0.944 | 0.942 |
| Noise / junk (F1) | 0.961 | 0.620 |
African languages (held-out test, F1)
Show table
| Language | lid-lite-608 | GlotLID v3 | Language | lid-lite-608 | GlotLID v3 | |
|---|---|---|---|---|---|---|
Yoruba yor |
0.906 | 0.827 | Kinyarwanda kin |
0.758 | 0.371 | |
Hausa hau |
0.881 | 0.806 | Kirundi run |
0.687 | 0.724 | |
Igbo ibo |
0.896 | 0.853 | Lingala lin |
0.814 | 0.642 | |
Nigerian Pidgin pcm |
0.776 | 0.749 | Twi twi |
0.785 | 0.732 | |
Swahili swh |
0.808 | 0.719 | Ewe ewe |
0.753 | 0.738 | |
Amharic amh |
0.985 | 0.939 | Fon fon |
0.904 | 0.841 | |
Zulu zul |
0.743 | 0.761 | Wolof wol |
0.809 | 0.600 | |
Xhosa xho |
0.840 | 0.602 | Somali som |
0.910 | 0.826 | |
Oromo gaz |
0.761 | 0.741 | Shona sna |
0.882 | 0.855 | |
Luganda lug |
0.906 | 0.867 | Tigrinya tir |
0.984 | 0.956 | |
| Average | 0.839 | 0.757 |
F1 across all input lengths, including 1-5 words.
About the test sets. Held-out test: 100 samples per length bucket per language, from websites
never seen in training. Web-weighted: the same test with each language weighted by its share of
web text, approximating real-world traffic. Tatoeba: held-out everyday sentences. FLORES+ and
UDHR-LID: public benchmarks not used to build our data. Full tables and per-language results are
in eval/.
Best Practices
- Pick the mode for your traffic: coverage mode for corpus building and low-resource mining, traffic mode for user-facing input.
- Long documents: split into ~256-word chunks and aggregate; this also catches pages that switch language partway through.
- Single-line input: fastText reads one line at a time, so collapse newlines first (the Quickstart does this).
- Junk filtering:
zxx_Zxxxmarks numbers, URLs, code and other non-language text; drop or route it separately.
Training Data and Licence
Trained from scratch on 7.86M samples from FineWeb-2 and MADLAD-400 (ODC-By 1.0), Aya Dataset and WURA (Apache-2.0), MasakhaNews (AFL-3.0) and Tatoeba (CC-BY 2.0 FR, sentences © Tatoeba contributors, https://tatoeba.org). Religious text, wrong-script text, mislabelled documents and foreign boilerplate were filtered out. Released under Apache-2.0.
Citation
@misc{lid-lite-608,
title = {lid-lite-608},
author = {Olaverse},
year = {2026},
url = {https://huggingface.co/olaverse/lid-lite-608}
}
- Downloads last month
- -



