Instructions to use mistralai/LIDstral-Arabic with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- fastText
How to use mistralai/LIDstral-Arabic with fastText:
from huggingface_hub import hf_hub_download import fasttext model = fasttext.load_model(hf_hub_download("mistralai/LIDstral-Arabic", "model.bin")) - Notebooks
- Google Colab
- Kaggle
LIDstral Arabic
LIDstral Arabic is a fast language and dialect classifier for text written in Arabic script. It identifies Modern Standard Arabic, all Arabic dialects, and non-Arabic languages that use the same script.
We include these non-Arabic languages in training to help the model distinguish languages that share a script and avoid classifying text as Arabic based on its script alone.
Method
The model combines a fastText one-vs-all (OVA) classifier, an MLP stacker, and per-class isotonic calibration:
- fastText OVA: produces independent scores for each class, preserving evidence for competing languages and dialects.
- MLP stacker: A 64-unit LayerNorm MLP maps these scores to probabilities over the classes. It uses 144 features derived from raw scores, prior corrections, confidence margins, entropy, and regional summaries. The checkpoint stores feature normalization and label ordering.
- Per-class isotonic regressors: fitted on development predictions calibrate the MLP probabilities. The pipeline then renormalizes them across classes. Calibration can change both confidence and the predicted label.
The full pipeline requires three files:
lid_ova_full.binmlp_stacker_ln.ptisotonic_calibration.pkl
classify.py runs the pipeline on CPU, including preprocessing and feature extraction. Each input that remains nonempty after preprocessing receives a predicted label and calibrated probabilities over all 51 classes.
Evaluation
We evaluate on 84,870 Arabic examples drawn from various benchmarks including NADI, MADAR, QADI, Casablanca, FLEURS, SMOL, Omnilingual ASR, and Hassaniya. On Moroccan Darija, LIDstral Arabic achieves 88.67% F1, compared with 72.65% for LahjatBERT ALDi CL, the strongest tested baseline for this dialect, and 71.28% for GlotLID v3.
Usage
Download the private repository with a Hugging Face account that has access:
pip install huggingface_hub
hf auth login
hf download mistralai/LIDstral-Arabic --local-dir LIDstral-Arabic
cd LIDstral-Arabic
pip install -r requirements.txt
python classify.py --text "ياك نتا بخير؟ آش خبارك مع الخدمة؟ نتمنى تكون الأمور كلها مزيانة من جيهتك."
Illustrative output, formatted for readability. Only selected probabilities are shown here; the full output includes all classes.
{
"label": "Morocco",
"confidence": 0.9523,
"probabilities": {
"Morocco": 0.9523,
"Algeria": 0.0125,
"Mauritania": 0.0115,
"MSA": 0.0062,
"Libya": 0.0048,
"..."
}
}
Keep all three model files in the same directory. The script loads them from its own directory by default. Use --model-dir to select another local bundle.
For JSONL input, each line must contain a string text field:
python classify.py --input input.jsonl > predictions.jsonl
Use --input - to read from standard input. The script returns one JSON object per input, with label, confidence, and probabilities. Empty inputs and text removed entirely by preprocessing raise an error.
To load the model once and classify multiple texts, use the Python API:
from classify import ArabicLID
model = ArabicLID(".").load()
predictions = model.predict([
"ياك نتا بخير؟ آش خبارك مع الخدمة؟ نتمنى تكون الأمور كلها مزيانة من جيهتك.",
"هذا نص باللغة العربية الفصحى.",
])
Illustrative batch output, with selected probabilities shown for each input:
[
{
"label": "Morocco",
"confidence": 0.9466,
"probabilities": {
"Morocco": 0.9466,
"MSA": 0.0185,
"Algeria": 0.0125,
"..."
}
},
{
"label": "MSA",
"confidence": 0.7295,
"probabilities": {
"MSA": 0.7295,
"Iraq": 0.1203,
"Egypt": 0.0266,
"..."
}
}
]
Preprocessing removes web and ASR artifacts, diacritics, decorative elongation, and repeated punctuation. It preserves alef variants.
Limitations
The model always assigns one of its 51 supported classes. It may therefore misclassify unsupported languages or mixed-language text.
License
This model is licensed under the Apache 2.0 License.
You must not use this model in a manner that infringes, misappropriates, or otherwise violates any third party’s rights, including intellectual property rights.
- Downloads last month
- 30
