lid-lite-608

lid banner

lid-lite-608 Highlights

lid-lite-608 is a fast, CPU-only language identification model for 608 languages, built African-first, with the following key features:

  • 608 languages in 37 MB. About 3x the languages of Meta's fastText LID (218) at 1/32 of its size, and 1/45 the size of GlotLID. Runs at ~6,800 texts/s on a single CPU thread.
  • Best-in-class on short text. 0.672 accuracy on 1-5 word inputs across all 608 languages, vs 0.486 for GlotLID and 0.218 for Meta's fastText LID.
  • Strong on African languages. Average F1 0.839 across 20 major African languages vs 0.757 for GlotLID, with large gains on Kinyarwanda, Xhosa, Wolof and Lingala.
  • Two modes in one model: coverage mode treats every language equally (corpus building, low-resource mining); traffic mode calibrates to real-world language frequency (user input, routing) and reaches 0.954 web-weighted accuracy, the best of all models tested.
  • Robust to messy input. Trained with typos, missing diacritics, URLs, emoji and everyday conversational sentences, plus a dedicated noise class for numbers, URLs and code (F1 0.961).
  • Commercial-friendly. Trained only on sources that allow commercial use; Apache-2.0.

Model Overview

lid-lite-608 has the following features:

  • Type: fastText supervised classifier (char n-grams 1-5 + word bigrams, dim=128, quantized)
  • Languages: 608 (567 distinct ISO 639-3 codes, 36 scripts) plus a noise class zxx_Zxxx
  • Labels: ISO 639-3 + ISO 15924 script, e.g. yor_Latn, srp_Cyrl (full list in labels.json)
  • Size: 37 MB (model.ftz)
  • Training data: 7.86M samples from commercially usable sources (see below)
  • Input: any length; 1 word to full documents

For the highest accuracy, especially on short and conversational input, see the neural sibling lid-neural-608 (mmBERT, same 608 languages, labels and modes).

Quickstart

pip install fasttext huggingface_hub
import json, math, fasttext
from huggingface_hub import hf_hub_download

repo = "olaverse/lid-lite-608"
model = fasttext.load_model(hf_hub_download(repo, "model.ftz"))
calib = json.load(open(hf_hub_download(repo, "priors.json")))
traffic_bias = {l: calib["alpha"] * v for l, v in calib["log_prior_ratio"].items()}

def identify(text, mode="coverage"):
    text = " ".join(text.split())                      # fastText needs a single line
    k = 40 if mode == "traffic" else 1
    labels, probs = model.predict([text], k=k)         # list input: works with NumPy 1 and 2
    labels, probs = labels[0], probs[0]
    scores = {l[9:]: math.log(max(p, 1e-12)) + (traffic_bias.get(l[9:], 0.0) if mode == "traffic" else 0.0)
              for l, p in zip(labels, probs)}
    return max(scores, key=scores.get)

identify("Ẹ kú àárọ̀, ṣé dáadáa ni?")                   # 'yor_Latn'
identify("Habari za asubuhi")                          # 'swh_Latn'
identify("Bonjour mon ami", mode="traffic")            # 'fra_Latn'

With the olaverse library

The olaverse library (v0.4.1+) wraps the model, the priors and both modes:

pip install "olaverse[lid]"
from olaverse import LIDLite608

lid = LIDLite608()                                     # mode="coverage" (default)
lid.predict("Ẹ kú àárọ̀, ṣé dáadáa ni?")               # 'yor_Latn'
lid.predict_proba("Habari za asubuhi", top_k=3)        # {'swh_Latn': 0.916, 'swc_Latn': 0.084, 'hau_Latn': 0.0}
lid.predict_batch(["Habari za asubuhi", "Mo fẹ́ lọ sí ọjà lónìí"])   # ['swh_Latn', 'yor_Latn']

lid.predict("Bonjour mon ami")                         # 'dhv_Latn'  (coverage: every language equally likely)
LIDLite608(mode="traffic").predict("Bonjour mon ami")  # 'fra_Latn'  (traffic: leans to common languages)

Newlines are collapsed for you, empty text raises ValueError, and it works with NumPy 2. Its predictions match the code above exactly (checked on 20,000 held-out texts in both modes).

Switching Between Coverage and Traffic Mode

One model, two ways to read its output:

mode="coverage" (default)

Every language is treated as equally likely. Use it when every language matters, such as mining low-resource text from a web crawl or building corpora. Highest per-language accuracy (macro-F1 0.856).

mode="traffic"

Scores are adjusted by each language's real-world frequency (priors.json), so short or ambiguous input leans toward the languages that actually dominate everyday traffic. Use it for user input, search queries, chat and request routing: +2.2 pts web-weighted accuracy and +6.6 pts on 1-5 word inputs compared with coverage mode.

Choosing Between lid-lite-608 and lid-neural-608

Both models share the same 608 languages, labels and coverage/traffic modes.

lid-lite-608 lid-neural-608
Size 37 MB 140M parameters (~560 MB)
Speed ~6,800 texts/s on one CPU thread 1,100 texts/s on an A100 GPU (40-60/s on an Apple M4 laptop)
Held-out test, accuracy 0.858 0.867
1-5 word inputs, traffic mode (web-weighted) 0.829 0.906
Everyday phrases, traffic mode 0.838 0.943
Best for Bulk filtering, CPU-only and edge deployments User input, short text, highest accuracy

Use both: run lid-lite-608 as a fast first pass and send only short or uncertain inputs to lid-neural-608.

Evaluation

All models scored on the same data. Other models' language codes are mapped to ours and they are scored on their best guess among our 608 languages, so naming differences never count as errors.

lid-lite-608 vs GlotLID and Meta

Show table
lid-lite-608 (coverage) lid-lite-608 (traffic) GlotLID v3 Meta fastText LID-218
Model size 37 MB 37 MB 1,687 MB 1,176 MB
Licence Apache-2.0 Apache-2.0 Apache-2.0 CC-BY-NC-4.0
Held-out test, accuracy (240k) 0.858 0.843 0.783 0.280
Held-out test, macro-F1 0.856 0.843 0.795 0.182
Held-out test, 1-5 words 0.672 0.656 0.486 0.218
Web-weighted accuracy 0.932 0.954 0.946 0.901
Web-weighted, 1-5 words 0.763 0.829 0.808 0.764
Tatoeba (conversational, 34k) 0.922 0.914 0.868 0.557
UDHR-LID (291 languages) 0.910 0.907 0.908 0.567
FLORES+ devtest (202 languages) 0.934 0.928 0.957 0.874
FLORES+ first 1-5 words 0.708 0.708 0.692 0.618
On the 208 languages Meta's fastText LID covers (table)
lid-lite-608 (coverage) lid-lite-608 (traffic) GlotLID v3 Meta fastText LID-218
Held-out test 0.900 0.897 0.885 0.822
Held-out test, 1-5 words 0.749 0.750 0.681 0.628
Tatoeba 0.937 0.939 0.911 0.858
FLORES+ first 1-5 words 0.732 0.735 0.724 0.684

By input length (held-out test, accuracy)

Accuracy by input length

Show table
Words lid-lite-608 GlotLID v3
1-5 0.672 0.486
6-20 0.871 0.796
21-80 0.924 0.902
81-256 0.944 0.942
Noise / junk (F1) 0.961 0.620

African languages (held-out test, F1)

African languages F1

Show table
Language lid-lite-608 GlotLID v3 Language lid-lite-608 GlotLID v3
Yoruba yor 0.906 0.827 Kinyarwanda kin 0.758 0.371
Hausa hau 0.881 0.806 Kirundi run 0.687 0.724
Igbo ibo 0.896 0.853 Lingala lin 0.814 0.642
Nigerian Pidgin pcm 0.776 0.749 Twi twi 0.785 0.732
Swahili swh 0.808 0.719 Ewe ewe 0.753 0.738
Amharic amh 0.985 0.939 Fon fon 0.904 0.841
Zulu zul 0.743 0.761 Wolof wol 0.809 0.600
Xhosa xho 0.840 0.602 Somali som 0.910 0.826
Oromo gaz 0.761 0.741 Shona sna 0.882 0.855
Luganda lug 0.906 0.867 Tigrinya tir 0.984 0.956
Average 0.839 0.757

F1 across all input lengths, including 1-5 words.

About the test sets. Held-out test: 100 samples per length bucket per language, from websites never seen in training. Web-weighted: the same test with each language weighted by its share of web text, approximating real-world traffic. Tatoeba: held-out everyday sentences. FLORES+ and UDHR-LID: public benchmarks not used to build our data. Full tables and per-language results are in eval/.

Best Practices

  • Pick the mode for your traffic: coverage mode for corpus building and low-resource mining, traffic mode for user-facing input.
  • Long documents: split into ~256-word chunks and aggregate; this also catches pages that switch language partway through.
  • Single-line input: fastText reads one line at a time, so collapse newlines first (the Quickstart does this).
  • Junk filtering: zxx_Zxxx marks numbers, URLs, code and other non-language text; drop or route it separately.

Training Data and Licence

Trained from scratch on 7.86M samples from FineWeb-2 and MADLAD-400 (ODC-By 1.0), Aya Dataset and WURA (Apache-2.0), MasakhaNews (AFL-3.0) and Tatoeba (CC-BY 2.0 FR, sentences © Tatoeba contributors, https://tatoeba.org). Religious text, wrong-script text, mislabelled documents and foreign boilerplate were filtered out. Released under Apache-2.0.

Citation

@misc{lid-lite-608,
  title  = {lid-lite-608},
  author = {Olaverse},
  year   = {2026},
  url    = {https://huggingface.co/olaverse/lid-lite-608}
}
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including olaverse/lid-lite-608