Confirmation: Can I use HuggingFace datasets for academic ML training?

Hello HuggingFace community,

My name is Mohammed Talbani, a Computer Science student at ESTS in Safi, Morocco. I am completing an AI internship where I need to build an image dataset of traditional Arabic footwear (babouche, balgha, na’al, markub, etc.) for a machine learning classification model.

I found several relevant shoe/footwear datasets on HuggingFace Hub. I would like to confirm:

  1. Can I download and use datasets published under open licenses (CC0, CC BY, Apache 2.0) on HuggingFace as training data for my ML model?
  2. Is bulk downloading via the HuggingFace API permitted for academic research?
  3. Are there any additional terms beyond the individual dataset licenses that I should be aware of?

This project is strictly non-commercial and academic.

Thank you for any guidance!

Mohammed Talbani
Computer Science Student — ESTS Safi, Morocco

Short answer: yes to all three, with a couple of things to check.

  1. Each dataset’s own license is what governs your use. CC0, CC BY and Apache 2.0 all allow training a model, academic or not. CC BY needs attribution, so cite the dataset in your report. Avoid anything tagged NC if you ever go beyond the internship, and read gated datasets’ access terms before accepting them. The license tag on the Hub is declared by the uploader and can be wrong, especially for scraped image sets, so check the dataset card for the real source of the images.

  2. Bulk downloading is fine through the official tools (datasets or huggingface_hub, e.g. snapshot_download). Log in with a token, since anonymous traffic gets lower rate limits, and avoid hammering the API with custom scrapers.

  3. Beyond the dataset licenses, the only general rules are the Hugging Face Terms of Service and the Hub rate limits. If any images show people, also check the dataset card for privacy or consent notes.