Best Multilingual Training Data Sources: 20 Datasets and Vendors Compared

Best Multilingual Training Data Sources: 20 Datasets and Vendors Compared

Run your multilingual eval and the pattern is familiar. English holds up, and the language your expansion plan depends on falls apart. A model is judged on its worst language, and most multilingual datasets are sold on their best ones.

The web explains why. In Common Crawl's latest crawl, CC-MAIN-2026-39, English is the primary language of 41.86% of pages, and 122 of the 161 detected languages each make up less than 0.1% (Common Crawl's own language statistics). Maltese is 0.0032% of pages and Quechua 0.0009%. Those figures only cover what the CLD2 detector can see (about 160 languages), so the real tail is longer. For a multilingual LLM in a world of more than 7,000 languages (Glottolog lists 7,674 spoken first languages), "100+ languages" says little about yours.

So we did the tedious part. We pulled the per-language tables ourselves on 2026-10-08, read every source's own card, and printed the same two rows for each: how deep it goes past the top 30 languages, and what its licence, and the licence of whatever it was built from, actually allows. By the end, you will have a shortlist that fits your languages, modality and licence needs, plus a check you can run on any source we missed. If nothing here covers your language, the last band shows how teams doing multilingual AI work build from in-language web sources.

Last updated October 2026. Facts read from each publisher's page on 2026-10-08.

Quick Digest
  • Count is not coverage: FineWeb-2 lists 1,868 language-script pairs, but only 203 of them have more than 10,000 documents.
  • Read the licence text: corpora built from Common Crawl inherit its Terms of Use, and none of them is public domain.
  • Speech is thin past the top: 108 of Common Voice's 295 languages have under 10 validated hours.
  • When open data runs out: license from a language-data vendor, or build a corpus from in-language web sources.

How did we judge these multilingual datasets?

Five factors decide whether a source fits your languages, and the headline language count is the weakest of them. They are ranked in the order they tend to eliminate candidates, and the rubric works on sources that are not on this list.

Factor What we checked Why it matters Where it shows
1. Languages (as stated) The publisher's count, in its own unit The headline number, and the least informative "Languages (as stated)" row
2. Tail depth past the top ~30 The per-language table, or "not published by publisher" Averages hide the tail "Low-resource depth" row
3. Licence + built-from The licence text in the card body, plus upstream terms Hub licence fields are often missing or wrong "Licence (and built-from)" row
4. Native vs translated Whether the card says human-written, templated or machine translated Translated text dominates parts of the tail "Modality & size" and "Watch-out" rows
5. Access Download, gated, suspended, membership or contract, plus public price A dataset you cannot get is not a source "Access & cost" row

87 of 205 audited language corpora had under 50% usable data, and 15 had no usable in-language text at all. Source: the 2022 audit of web-crawled multilingual datasets (Kreutzer et al., TACL 2022), covering the then-current releases of CCAligned, ParaCrawl, WikiMatrix, OSCAR and mC4.

That audit is why the depth row exists. Fluent annotators judged 100 lines per corpus and deliberately targeted the smallest languages, because averages would have hidden the result: English was 43% of all sentences in OSCAR, and a random mC4 sample had over a 63% chance of coming from one of the 8 largest languages. A dataset-level quality score mostly describes the head.

  • What "depth" means: volume past the top ~30 languages in the publisher's unit, and diversity. The FineWeb2 paper reports that 70% of its language-script pairs (1,320 of 1,868) get more than half their documents from Bible- or Wikipedia-related domains.
  • Why the hub licence field is not used: a Data Provenance Initiative audit of 1,800+ text datasets found licences omitted more than 70% of the time on major hosting sites, and wrong more than 50% of the time when present.
  • How we flag translated data: native, templated or machine translated, as the card states it. Described, not judged.
  • How we treat vendor claims: labelled "vendor claim", dated 2026-10-08, not audited by us.

A language count still works as a filter: it tells you a source attempted your language. No source on this list paid to be included; Forage AI appears once, in the build-your-own band, because it is a collection partner, not a data source.

Multilingual sources at a glance

A language count is not coverage; check the depth column for your language. The roster lists all 20 in band order with one licence flag each. "Yes" never means "free for any use": each entry quotes the licence text, and the flags save you a wasted week on sources that were never an option.

# Source Band Best for Licence flag
1 Common Crawl Pretraining text Raw material for your own pipeline Yes, CC Terms of Use apply
2 FineWeb-2 Pretraining text Broadest filtered long-tail text Attribution
3 HPLT 3.0 Pretraining text Large-scale tail-aware text with annotations Packaging only
4 MADLAD-400 Pretraining text Audited document-level text, 419 languages Attribution
5 CulturaX Pretraining text Cleaned mC4 + OSCAR in one place Attribution, inherited; gated
6 Common Corpus Pretraining text Per-document open licences Yes, per document
7 Wikipedia dumps Pretraining text Clean attributable text Share-alike
8 OPUS Parallel Any parallel data for a pair Per corpus
9 NLLB mined bitext Parallel Non-English-centric pairs at scale Attribution + source terms
10 Mozilla Common Voice v27.0 Speech Open read speech, most languages Yes, CC0 with conditions
11 FLEURS Speech Parallel speech for fine-tune/eval Attribution
12 Multilingual LibriSpeech Speech Deep read speech, 8 European languages Attribution
13 Aya Dataset Instruction & evaluation Native human-written instruction data Yes, Apache 2.0
14 FLORES+ Instruction & evaluation Measuring translation quality Eval only
15 Hugging Face Hub Repositories Searching and streaming open sets Per dataset
16 LDC (with ELRA/ELDA) Repositories Contract-licensed curated corpora Contract
17 Appen Licensed language data Off-the-shelf datasets + custom collection Contract
18 Defined.ai Licensed language data Marketplace datasets + collection Contract
19 LXT Licensed language data Custom collection, wide locale list Contract
20 Forage AI, managed extraction partner Build your own A corpus built from in-language web sources Contract, client-defined sources

Status notes, as of October 2026: OSCAR 23.01 shows "access suspended" on Hugging Face. mC4 and CC-100 are covered under CulturaX. Emilia is non-commercial only (CC-BY-NC). ParaCrawl's last release was September 2021. Common Voice moved to Mozilla Data Collective in October 2025.

Best multilingual training data sources by data type

The 20 sources sit in seven bands, and the bands follow one question: what do you need to train? Pretraining text first, then parallel, speech, instruction and evaluation, repositories, licensed language data, and building your own. Open sources come before paid ones in each band. Every entry prints the publisher's own unit, because pages, words, tokens and hours are not interchangeable measures of LLM training data.

Web-scale pretraining text

Tails are thinner than the counts suggest. FineWeb-2 is the broadest at 1,868 language-script pairs, and Common Corpus is the only one licensed per document, but only 33 of its languages pass 1B tokens. Every corpus built from Common Crawl inherits its Terms of Use (ToU).

To make depth comparable, we checked five probe languages (counts as of October 2026) that span scripts and regions: Maltese, Yoruba, Amharic, Khmer (no spaces between words) and Quechua (a macro-language).

Source (version) Unit Maltese Yoruba Amharic Khmer Quechua
Common Crawl CC-MAIN-2026-39 pages 69,827 36,897 99,183 231,875 18,577
FineWeb-2 v2.1.1, kept docs 489,190 79,999 428,373 1,586,460 2,903 (quy)
HPLT 3.0 docs 752,735 171,248 571,240 1,323,660 20,199 (quy)
MADLAD-400 clean docs 265.4K 52.1K 106.3K 285.7K 2.4K (qu)
CulturaX docs 151,320 192 243,349 1,013,181 1,202 (qu)
Wikipedia articles 8,023 41,357 15,733 12,892 24,641

Pages, documents and articles are not interchangeable. Compare within a row (which language is thin in this source) and within a column (which source has more documents). Never add across rows. Quechua is also coded differently everywhere (FineWeb-2 splits it into 11+ dialect subsets, HPLT 3.0 uses quy, MADLAD-400 uses qu), so search every ISO code your language goes by before concluding a source lacks it.

Every corpus above lists Yoruba as covered. CulturaX holds 192 Yoruba documents; HPLT 3.0 holds 171,248. Filtering widens the gap: computed from FineWeb-2's own CSV, languages with 1M+ post-dedup documents kept 34.9% of them, while languages in the 10K to 100K range kept 6.9%. The FineWeb-2 card even reports that for Swahili, filtered data (around 1B tokens) performed worse than deduplicated data (around 3B tokens). In the tail, filtering is a trade-off, not a free quality win.

Horizontal bar chart of Yoruba documents held by four multilingual pretraining corpora that all list Yoruba as covered: HPLT 3.0 171,248, FineWeb-2 kept 79,999, MADLAD-400 clean 52.1K, CulturaX 192. Source: each publisher's per-language table, read 2026-10-08.
Same language, same "covered" label: Yoruba documents per corpus, in each publisher's own count.

1. Common Crawl

Attribute Detail
Best for Raw material for your own language-ID, dedup and filtering pipeline. Not for: a ready training set
Languages (as stated) About 160 detectable with CLD2; 161 identified in CC-MAIN-2026-39
Low-resource depth English 41.86% of pages; 122 of 161 languages under 0.1%; Yoruba 36,897 pages
Modality & size Raw HTML (WARC), metadata (WAT), text (WET); one crawl = 2.17B pages, 361.4 TiB uncompressed
Licence (and built-from) Terms of Use (7 Mar 2024): limited licence to the service; content stays its owners' responsibility; legal advice recommended; AI use named in the indemnity clause
Access & cost Free; s3://commoncrawl/ or data.commoncrawl.org
Reviews (n) No G2 listing (open dataset). Practitioner signal: HN user ks2048, 2026-03-21
Watch-out C4, The Pile and RefinedWeb are built from it but English-only; 3.06% of pages have no identified language

Common Crawl is the raw material behind most of this band. If you run your own pipeline, it gives you control over language-ID and filtering thresholds, which matters because published defaults are not always right for a small language.

The cost is that all of that work becomes yours: language ID, dedup, boilerplate removal and legal review. A founder working on low-resource data gathering said in March 2026 that corpora like Common Crawl and FineWeb are "really lacking quality data sources if you know where to look".

Watch out

Common Crawl is not public domain. Its Terms of Use grant a limited licence to the service, leave copyright with content owners, and name AI training as a use the user must indemnify Common Crawl for. Every corpus built from it inherits those terms.

2. FineWeb-2

Attribute Detail
Best for The broadest open pretraining text across the long tail. Not for: English (excluded)
Languages (as stated) 1,868 language-script pairs ("over 1000 languages")
Low-resource depth 474 pairs >1K docs; 203 >10K docs; 56 >1B words. Probe: Yoruba 80K docs; Ayacucho Quechua 2.9K
Modality & size Text; 3.34T words, 5.02B docs; 96 Common Crawl snapshots (2013 to April 2024)
Licence (and built-from) ODC-BY 1.0: "also subject to CommonCrawl's Terms of Use"
Access & cost Free; Hugging Face per subset (e.g. por_Latn), streaming supported
Reviews (n) No G2 listing (open dataset)
Watch-out GlotLID can mistake close languages (Croatian and Bosnian); ablations on 9 languages only

The breadth is real, and so is the thinness: the 80th-largest pair, Central Kurdish, holds about 236M words, against 588.6B for Russian. Filtering kept 4.8% of post-dedup Maltese documents, and several Quechua subsets are over 98% Bible text by the publisher's own CSV.

We still start every tail-language check here. FineWeb-2's per-language CSV publishes both kept and removed counts, so you can see what the filters threw away before downloading. Each language also ships a small test split you should not train on.

3. HPLT 3.0

Attribute Detail
Best for Large-scale, tail-aware web text with quality annotations. Not for: teams that need a licence on the text itself
Languages (as stated) 198
Low-resource depth Gemma-3 tokens: 82 languages >1B; 75 <100M; 36 <10M. Probe: Yoruba 171K docs; Quechua 20K
Modality & size Text; about 50 TB compressed; 13.5T tokens excluding English
Licence (and built-from) CC0 on the packaging only: "We do not own any of the text"; notice-and-takedown
Access & cost Free; zstd JSONL via per-language map files
Reviews (n) No G2 listing (open dataset)
Watch-out CC0 is often misread as covering the text; users carry EU copyright and privacy obligations

In our probe, HPLT 3.0 held the most Yoruba and Quechua documents of any open text corpus. The July 2025 release ships annotations most corpora skip: web-register labels for 104 languages, a document-quality score, PII annotation and robots.txt opt-out filtering.

Depth is not the weak point here. The licence on the underlying text is, and that question goes back upstream to the original websites.

4. MADLAD-400

Attribute Detail
Best for Audited, document-level text for 419 languages. Not for: quality filtering in Chinese, Thai or Burmese
Languages (as stated) 419 (after dropping 79 of 498 in a self-audit)
Low-resource depth Clean whitespace tokens: 42 languages >1B; 291 of 419 under 10M. Probe: Quechua 2.4K docs
Modality & size Text; clean 3.7B docs / 2.6T tokens; noisy 7.2B docs; snapshots to 1 Aug 2022
Licence (and built-from) The card text says CC-BY-4.0, the hub metadata says ODC-BY; both require attribution; the text stays bound by the Common Crawl ToU
Access & cost Free; Hugging Face (allenai/madlad-400), clean and noisy
Reviews (n) No G2 listing (open dataset)
Watch-out Certain tail languages contain "only Bible text, or in some cases jw.org data"

MADLAD-400 is the only corpus here whose builders published a per-language self-audit and dropped the languages that failed. The Google team found significant Bible data in 141 of 498 languages and cut 79 to avoid "Representation washing". That work dates from 2023, so treat it as context, not current state.

5. CulturaX

Attribute Detail
Best for Cleaned, deduplicated mC4 + OSCAR in one place. Not for: low-resource tail work
Languages (as stated) 167
Low-resource depth Probe: Yoruba 192 docs; Quechua 1,202 docs. English is 45.13% of tokens
Modality & size Text; 6.3T tokens, 3.2B+ docs
Licence (and built-from) "Strictly follows" mC4 (ODC-BY + Common Crawl ToU) and OSCAR (CC0 packaging over text it does not own)
Access & cost Free but gated (Hugging Face login, accept terms)
Reviews (n) No G2 listing (open dataset). Practitioner signal: Reddit, 2026-06, calls its Maltese coverage "basically unusable" (opinion)
Watch-out Likely inherits upstream language-ID errors (our inference); may contain personal information

CulturaX gets the sharpest depth row on this list: its "167 languages" includes Yoruba at 192 documents. For head languages, though, it is a clean, ready-merged corpus, and if you already work under mC4 and OSCAR terms it saves you the merge.

Three older corpora sit underneath it. The mC4 dataset has 108 subsets under ODC-BY plus the Common Crawl terms (not CC BY-SA, whatever one ranking page says). CC-100 is a static 2018 set whose hub licence field reads "unknown". The OSCAR corpus, version 23.01, currently says access is suspended until "the situation has been clarified".

6. Common Corpus

Attribute Detail
Best for Commercial pretraining where every document carries its own open licence. Not for: contemporary web language or tail languages
Languages (as stated) Mostly English and French; 33 languages >1B tokens
Low-resource depth Probe languages not published by publisher at this granularity
Modality & size Text, including OCR'd books and newspapers, government, science and code; 2.27T tokens (October 2026)
Licence (and built-from) Per document: "uncopyrighted or freely licensed and may be used for both commercial and non-commercial purposes"
Access & cost Free on Hugging Face (PleIAs/common_corpus)
Reviews (n) No G2 listing (open dataset)
Watch-out More than half "predates the 21st century" (historical language, OCR errors)

This is the cleanest licence chain on the list. If your languages are among the 33 and a historical register is acceptable, it beats every Common Crawl derivative on licence, because each document carries its own license field.

The trade is language and era. A model pretrained heavily on old newspapers will sound like one, and the v3 expansion to Chinese, Japanese, Arabic, Korean and Hindi is still "ongoing".

7. Wikipedia dumps

Attribute Detail
Best for Clean, attributable encyclopedic text. Not for: volume in the tail, or conversational text
Languages (as stated) 348 active editions
Low-resource depth 11 editions >1M articles. Probe: Yoruba 41,357 articles; Quechua 24,641; Maltese 8,023
Modality & size Text, plus Wikidata structured data
Licence (and built-from) Text: GFDL and CC BY-SA 4.0 (attribution + share-alike); Wikidata: CC0
Access & cost Free; dumps.wikimedia.org or the Hugging Face mirror
Reviews (n) No G2 listing (open dataset). Practitioner signal: Reddit r/MLQuestions, 2025-10-18, advises contributing to Wikipedia
Watch-out Share-alike obligations; the Hugging Face mirror lags (2023 snapshot, still says 3.0)

Article count does not track web presence: Maltese has the most web text of our five probes and the fewest articles. For several tail languages, Wikipedia is bigger than what the filtered corpora keep, with Quechua at 24,641 articles against FineWeb-2's 2,903 kept Ayacucho documents. Check it first.

The obligation is share-alike. If you redistribute derived text, the licence travels with it, so it belongs in your licence review before Wikipedia goes into a commercial mix.

Parallel and translation corpora

OPUS is the widest index of parallel corpora (1,038 languages, 102.9B sentence pairs), and each corpus inside it carries its own licence. Mined bitext like NLLB reaches pairs no one translated by hand, and its card admits part of it is machine translation.

That matters in the tail. In a 6.4-billion-sentence study, 57.1% of translated sentences on the web appeared in three or more languages, and the authors found this content is mostly machine translated and makes up a large fraction of all web text in lower-resource languages (Thompson et al., Findings of ACL 2024). The 2022 audit of then-current releases found CCAligned averaging 29.25% correct sentences and WikiMatrix 23.74%, plus 83 corpora with mislabelled language codes.

The head is the opposite case. For EU languages paired with English, any machine translation dataset question is about licence, not supply: ParaCrawl alone gives 278.3M English-German sentences.

8. OPUS

Attribute Detail
Best for Finding any available parallel data for a language pair. Not for: a single licence you can sign off once
Languages (as stated) 1,038 languages; 1,214 corpora
Low-resource depth English-X pairs: Maltese 31.66M; Yoruba 1.79M; Khmer 8.62M; Quechua 2.88M (99.6% from NLLB)
Modality & size Parallel text; 102.9B sentence pairs
Licence (and built-from) Per corpus; no central licence statement
Access & cost Free; website, OPUS API, OpusTools, OPUS Explorer
Reviews (n) No G2 listing (open dataset)
Watch-out Includes ParaCrawl (last release Sept 2021) and CCMatrix (no stated licence)

OPUS, run by the University of Helsinki, is the first stop for any pair: its API returns per-pair counts by corpus in seconds, which is how we built our parallel probe. It does not settle the licence. Among its largest corpora, CCMatrix and OpenSubtitles state no licence on their OPUS page, and ParaCrawl and HPLT waive rights only in the packaging.

"OPUS has 2.9M English-Quechua pairs" in practice means one mined, partly machine-translated corpus has them. That is three orders of magnitude above FineWeb-2's kept Ayacucho documents. Different units, but a clear prompt to sample first.

9. NLLB mined bitext

Attribute Detail
Best for Non-English-centric pairs at scale. Not for: clean human translation in the tail
Languages (as stated) 148 English-centric + 1,465 non-English-centric pairs
Low-resource depth Not printed in the card. Via OPUS: en-yo 1.51M, en-am 16.14M, en-qu 2.87M pairs
Modality & size Mined parallel text, about 450 GB (LASER3 encoders)
Licence (and built-from) ODC-BY, plus "the respective Terms of Use and License of the original source"
Access & cost Free on Hugging Face (allenai/nllb)
Reviews (n) No G2 listing (open dataset)
Watch-out Card: "Some of the translations are in fact machine translations"

For a large share of tail pairs, NLLB is the only parallel data at scale: it supplies 95.8% of OPUS's English-Amharic pairs and 84.2% of English-Yoruba. The card is candid that machine-translated content stayed in because raw HTML was not available to filter it.

So use it, and sample it. A fluent reader going through 100 pairs will tell you more than the pair count.

Multilingual speech datasets

For any multilingual speech dataset, the language count and the hours tell different stories. Common Voice lists 295 languages, but 108 have under 10 validated hours. FLEURS has about 10 hours per language by design, and Multilingual LibriSpeech is deep but covers 8 languages.

Bar chart of Common Voice v27.0 validated hours for five probe languages against a 10-hour marker: Maltese 8.76, Yoruba 5.96, Amharic 1.94, Khmer 0.28, Ayacucho Quechua 0.05. 108 of Common Voice's 295 languages have under 10 validated hours. Source: Common Voice cv-dataset v27.0.
All five probe languages are listed in Common Voice; none reaches 10 validated hours.

Our probe languages make the point. In Common Voice v27.0, Maltese has 8.76 validated hours, Yoruba 5.96, Amharic 1.94, Khmer 0.28 from 15 contributors, and Ayacucho Quechua 0.05 from 8. All five are "covered".

The licence chain applies here too: an audit of nearly 4,000 datasets found "over 80% of the source content in widely-used text, speech, and video datasets, carry non-commercial restrictions" (Longpre et al., ICLR 2025).

10. Mozilla Common Voice (Scripted Speech v27.0)

Attribute Detail
Best for Open read speech across the most languages, with commercial use allowed. Not for: tail languages that need 100+ hours
Languages (as stated) 295
Low-resource depth Validated hours: ≥100 h 39 languages · <10 h 108 · <1 h 31. Probe: Yoruba 5.96 h, Khmer 0.28 h
Modality & size Read speech; 29,295 validated hours (cut-off 2026-09-11)
Licence (and built-from) CC0-1.0 with conditions: no attempts to identify speakers; no re-hosting or re-sharing
Access & cost Free; since October 2025, "exclusively available through Mozilla Data Collective"
Reviews (n) No G2 listing (open dataset)
Watch-out Volunteer read speech; uneven hours; old "100+ languages (v11)" figures are outdated

Hours follow contributors, not speaker population. Kinyarwanda has 2,002 hours and Luganda 437, while Hausa sits at 4. Community campaigns can make a "low-resource" language deep.

The practical 2026 change is access. Pipelines that pulled Common Voice from Hugging Face need to move to Mozilla Data Collective, and the no-re-sharing condition rules out mirroring it inside a shared dataset.

11. FLEURS

Attribute Detail
Best for n-way parallel speech for fine-tuning or evaluation. Not for: pretraining
Languages (as stated) 102
Low-resource depth About 10 h of training data per language, flat by design
Modality & size Read speech of 2,009 parallel FLoRes sentences, 1-3 recordings each
Licence (and built-from) CC-BY-4.0
Access & cost Free on Hugging Face (google/fleurs)
Reviews (n) No G2 listing (open dataset)
Watch-out Card notes a "known mismatch" between read-speech and noisier settings

Google built FLEURS for comparison, not volume. Because every language reads the same sentences, it is one of the few ways to measure one model across 102 languages on identical content, which is why it anchors multilingual speech recognition benchmarks.

Ten hours helps a fine-tune and will not carry pretraining. Use it as your measuring stick.

12. Multilingual LibriSpeech

Attribute Detail
Best for Deep read speech in 8 European languages. Not for: anything outside those 8
Languages (as stated) 8
Low-resource depth Train hours: German 1,966.51 · Dutch 1,554.24 · French 1,076.58 · Polish 103.65; no probe language
Modality & size Read audiobooks (LibriVox); about 88% of hours are English
Licence (and built-from) CC-BY-4.0 (treat it as an attribution licence)
Access & cost Free on Hugging Face (facebook/multilingual_librispeech)
Reviews (n) No G2 listing (open dataset)
Watch-out European only, read speech, English-dominated

For its seven non-English languages, Meta's MLS gives roughly 100 to 2,000 hours each, comparable to or deeper than Common Voice for most of them. If your languages are on its list, it is a strong base.

Three bigger releases belong in your licence notes, not your shortlist. YODAS (369,510 h) is YouTube-sourced with captions "not necessarily transcribed by a human", so platform terms apply. Emilia is CC-BY-NC, "only for non-commercial purposes". MLCommons' Unsupervised People's Speech (1M+ hours) is unlabelled and, by its own card, almost all American-accented English.

Instruction and evaluation data

The Aya Dataset is native, human-written instruction data in 65 languages under Apache 2.0. The larger Aya Collection is mostly translated. FLORES+ and Global-MMLU exist to measure, and the FLORES+ card says not to train on it.

We checked the Aya Collection's composition through the Hugging Face datasets-server on 2026-10-08: about 96% of its 513M rows are machine translations (NLLB 3.3B), and the human-written Aya Dataset is 0.04% of it. That is a description, not a verdict. The Aya paper notes that "training models with translated data can yield significant benefits".

13. Aya Dataset

Attribute Detail
Best for Native, human-written instruction data for multilingual fine-tuning. Not for: volume
Languages (as stated) 65 (71 with dialects and scripts)
Low-resource depth Per-language counts not pulled for this edition; the card warns of "a few dominant annotators" in certain languages
Modality & size 204,114 human-annotated prompt-completion pairs
Licence (and built-from) Apache 2.0: "any purpose, whether academic or commercial"
Access & cost Free on Hugging Face (CohereLabs/aya_dataset)
Reviews (n) No G2 listing (open dataset)
Watch-out The Aya Collection (115 languages) is 95.9% machine translated and 4.1% templated; do not mistake it for this dataset

The Aya Dataset was written by people in their own languages and released under a licence that allows commercial use. Native instruction data at this quality rarely exists outside English, and multilingual LLM training teams feel that gap first in fine-tuning.

Expert Insights

Sara Hooker, then VP of Research at Cohere and head of Cohere For AI, described a data "cliff" outside English fine-tuning data when Aya launched, which made its human-written data "incredibly rare". As she put it: "These models have been used all over the world and so people want it to work for them." (VentureBeat, 13 February 2024)

At 204K pairs, plan it as a fine-tuning set, and read the per-language counts on the card before assuming your language is well represented.

14. FLORES+ (eval only)

Attribute Detail
Best for Measuring translation quality across 230 language varieties. Not for: training, ever
Languages (as stated) 230 language varieties (v4.6)
Low-resource depth 997 dev + 1,012 devtest sentences per language, flat
Modality & size Parallel evaluation sentences from Wikinews, Wikijunior and Wikivoyage
Licence (and built-from) CC BY-SA 4.0
Access & cost Free, gated (accept conditions)
Reviews (n) No G2 listing (open dataset)
Watch-out Global-MMLU (42 languages, Apache 2.0, 13 of them low resource) is also eval only

FLORES+ is maintained by the Open Language Data Initiative as the successor to FLORES-200, so if you search "flores 200", this is the current version (230 varieties, not the 222 on older pages).

Its value is that it stays out of training. A clean shared eval set is what makes the depth claims on this list testable for your language.

Watch out

FLORES+ and Global-MMLU are evaluation sets. The FLORES+ card says it "should not be used as training data". Train on them and your scores stop meaning anything.

Dataset repositories and catalogues

You search for multilingual datasets on the Hugging Face Hub and read the licence in each card. LDC and ELRA are the contract-licensed route. There, commercial training rights are written down per corpus. Self-serve marketplaces are covered in our AI dataset marketplaces compared; this piece covers only their multilingual holdings.

What we've seen on this list: sidebar licence fields contradicted the card text three times, on MADLAD-400, CC-100 ("unknown") and the Wikipedia mirror (3.0 vs 4.0). So we built the table we wanted when we started.

Source Where per-language numbers live Unit
Common Crawl cc-crawl-statistics languages.csv pages
FineWeb-2 GitHub language-distribution CSV (with removed counts) words, docs
HPLT 3.0 / MADLAD-400 README tables tokens, docs
Common Voice cv-dataset GitHub JSON (validHrs) validated hours
NLLB / OPUS OPUS API per pair sentence pairs
Aya Collection Hugging Face datasets-server /size rows

15. Hugging Face Hub

Attribute Detail
Best for Searching and streaming open multilingual sets. Not for: trusting the sidebar licence or the "multilingual" tag
Languages (as stated) Per dataset; 12,175 datasets tagged multilingual of 1,082,877 total (2026-10-08)
Low-resource depth Per dataset card; the tag says nothing about volume
Modality & size All modalities; per dataset
Licence (and built-from) Per dataset; read the card body, not the sidebar
Access & cost Free browsing; gated sets need login and acceptance
Reviews (n) Reviews not pulled for this edition
Watch-out Hub licence fields are unreliable; tags are self-reported
Platform Rating n Date read
G2 / Capterra Not pulled for this edition n/a 2026-10-08

The Hub is the fastest way to find and stream a candidate, as long as you treat it as an index, not evidence. Gated datasets at least make you read and accept terms.

Our working habit: search the tag, open the card, find the per-language table, read the licence paragraph. If a card has no per-language table, that absence is your depth answer.

16. LDC (with ELRA/ELDA folded in)

Attribute Detail
Best for Contract-licensed, curated speech and text corpora. Not for: quick, free experiments
Languages (as stated) Not stated on the pages read (LDC); ELRA lists 1,598 language resources
Low-resource depth Per corpus (IARPA Babel and LORELEI corpora target low-resource languages)
Modality & size Speech, text, lexicons; per corpus
Licence (and built-from) Membership or non-member agreement; corpus-specific licences supersede. ELRA: END USER (research), VAR (commercial)
Access & cost LDC for-profit membership $34,000 (not-for-profit $2,400); one ELRA database lists €20,000 for a non-member commercial licence
Reviews (n) No verified reviews
Watch-out Research licences do not carry over to commercial training
Platform Rating n Date read
G2 / Capterra No verified reviews (not pulled) n/a 2026-10-08

The Linguistic Data Consortium at the University of Pennsylvania and ELRA in Europe hold decades of curated corpora; Joshi et al. built their language-resource classes from the LDC catalog and the ELRA Map.

Price is the barrier, and LDC says most corpora reach non-members "under research-only licenses". In return, the commercial training licence is written down per corpus, which is exactly the licence row this article asks for.

Forage AI promotional banner reading
When no catalogue holds your language, Forage AI runs managed collection from in-language web sources you define.

Licensed and commissioned language data

When a tail language or a commercial licence rules out open data, a vendor is the fastest licensed route. AI training data companies sell by locale count. So every number here is a vendor claim, read on 2026-10-08. Welo Data, TELUS Digital, DataForce and Shaip are covered in our AI training data providers guide.

Only one of these three publishes per-dataset language and volume publicly (Defined.ai), none publish prices, and Appen offers samples under NDA. Capterra, Trustpilot and Gartner Peer Insights returned access errors to our research pass, so the reviews tables are G2-only. For an honest comparison, ask four questions before you sign:

  1. Volume per language: hours or tokens for your language, not a locale count.
  2. Native vs translated: what share is human-written in-language.
  3. Sample: a representative sample before purchase.
  4. Licence scope: perpetual or term, exclusive or not, and whether it covers commercial model training.

17. Appen

Attribute Detail
Best for Off-the-shelf multilingual datasets with commercial training rights, plus custom collection. Not for: buyers who need per-language volumes up front
Languages (as stated) Vendor claim: "500+ global locales covered"
Low-resource depth Not published by publisher
Modality & size Vendor claim: 597 datasets across nine categories; text, speech, RLHF/SFT data
Licence (and built-from) Contract; "perpetual, non-exclusive commercial training rights" on off-the-shelf data (vendor wording)
Access & cost Quote-based; samples under NDA for most catalogued datasets
Reviews (n) G2 4.2 (n=35), read 2026-10-08
Watch-out Locale counts are not language depth
Platform Rating n Date read
G2 4.2 35 2026-10-08
Capterra / Trustpilot / Gartner Not available (access blocked) n/a 2026-10-08

Appen's commercial-training-rights wording is clearer than most open-dataset licence chains here, which helps when legal wants one document to sign off.

The public catalogue tells you to filter "by locale for current coverage" and shows no per-dataset volumes, so expect depth for your language only through sales and an NDA sample.

18. Defined.ai

Attribute Detail
Best for Buying off-the-shelf speech and text datasets from a marketplace. Not for: decisions based on its review count
Languages (as stated) Vendor claim: "500+ languages, dialects, and locales"
Low-resource depth Listed per dataset on the marketplace (e.g. "696 hours of podcast videos in Hindi")
Modality & size Off-the-shelf datasets, collection, annotation, evaluation, machine translation (MT)
Licence (and built-from) Contract; terms vary per dataset
Access & cost Quote-based; no prices on the marketplace
Reviews (n) G2 4.5 (n=2), read 2026-10-08; too small to rely on
Watch-out A marketplace moves the licence question to each dataset rather than resolving it
Platform Rating n Date read
G2 4.5 2 2026-10-08
Capterra / Trustpilot / Gartner Not available (access blocked) n/a 2026-10-08

Defined.ai is the only vendor of the three whose public catalogue lets you check per-language volume before you call sales. You can confirm your language has hours before the first meeting.

A 4.5 rating from two reviewers tells you almost nothing, so weight references and samples over the score.

19. LXT

Attribute Detail
Best for Custom collection across a wide locale list. Not for: buyers who need third-party reviews first
Languages (as stated) Vendor claim: "1,000+ language locales covered"
Low-resource depth Not published by publisher
Modality & size Collection, annotation, evaluation, transcription; vendor claim "10M+ vetted contributors in 150+ countries"
Licence (and built-from) Contract
Access & cost Not public
Reviews (n) No verified reviews
Watch-out No independent reviews; ask for references and a per-language sample
Platform Rating n Date read
G2 No listing found n/a 2026-10-08

LXT makes the widest locale claim of any vendor here. If your locale is unusual, that alone justifies a sample request.

We found no G2 listing for LXT itself; its crowd runs through clickworker, whose reviews belong to that brand. Ask for references from a buyer in your language family.

Build your own multilingual corpus from the web

This article is for informational purposes only and does not constitute legal advice. Consult a qualified attorney for legal guidance specific to your situation.

Building a corpus is a sourcing project, not a download. When open sources run out for your language, the route is non-English training data from in-language web sources that a general crawl under-samples: news, government, forums and educational websites written in that language.

Seven-step route for building a multilingual corpus from in-language web sources: in-language source register, language ID at collection, dedup and boilerplate removal, robots.txt and terms review, access within rate limits and terms, provenance log, and ongoing maintenance that loops back to step 1.
The build route is a sourcing project, and step 7 loops back to step 1.

The best public worked example is MaCoCu (EAMT 2022). Instead of filtering Common Crawl, a four-institution, EU-funded consortium crawled national top-level domains such as .hr for Croatian. Its first release collected 347.9M Maltese words, comparable to or more than the 286.5M words FineWeb-2 keeps from 96 snapshots. A targeted crawl can beat a filtered general crawl, and it took a consortium.

Targeting beats filtering for a reason. In the 2022 audit, language-ID filtering raised median precision "from 13.8% pre-filter to 43.9% post-filter" at "a steep cost of 77.5% loss in recall". Filtering does not create data you do not have. The route has seven steps, and the last never finishes:

  1. In-language source register: named sources per language (news, government, forums, subtitles, national-domain websites), not a crawl of everything. Measure coverage against it.
  2. Language ID at collection: the language-ID (LID) model sets the ceiling. In Google's test, language-ID models that scored well on clean test sets produced web corpora that were only about 5% in the right language for many low-resource languages (Caswell et al., COLING 2020).
  3. Dedup and boilerplate removal: tuned per language, since aggressive filters can remove most of a small one. Our operational playbook for training data covers this at scale.
  4. Robots.txt and terms review: a floor, not a legal answer. Where personal data is in scope, France's data-protection authority, the CNIL, has said ignoring robots.txt or CAPTCHAs defeats the legitimate-interest basis.
  5. Access: within rate limits and site terms; logins and CAPTCHAs are never bypassed. Since July 2025, Cloudflare, which says it handles traffic for about 20% of the web, asks every new domain whether to allow AI crawlers and defaults toward blocking them.
  6. Provenance log: per document, the source URL, collection date, terms snapshot and what was checked. It is a record, not a guarantee; our piece on public vs private training data covers why it matters for licensing.
  7. Maintenance (ongoing): sources change and close. In 2023-24, more than 28% of the most actively maintained sources in C4 became fully restricted (Longpre et al., NeurIPS 2024). C4 is English, so read it as a sign that web supply is narrowing. Then loop back to step 1.
Step DIY Managed partner
Source register Your team researches sources per language Defined with the partner; you co-own the list
Language ID You choose and validate a LID model (e.g. the GlotLID/OpenLID class) Agreed with the partner, with validation samples you sign off
Dedup / boilerplate Your pipeline Partner's pipeline; you review samples
Robots + terms review Your review plus legal counsel Scoped with the partner; legal advice remains yours
Access Your engineering, within site terms Partner operates collection within site terms
Provenance log You build it What you receive is agreed in the project scope
Maintenance Your on-call Partner maintains as sources change

If your language is in the top 30, building is usually unnecessary for volume. The route earns its cost when the open tail fails you, and in our probe it did for four of five languages in speech, and for Quechua and Yoruba in text.

20. Forage AI, managed extraction partner

Attribute Detail
Best for Teams whose languages, domains or licence needs are not met by open sets or vendor catalogues. Not for: anyone who wants a ready-made dataset today
Languages (as stated) Defined per project with the client
Low-resource depth Depends on the sources you and Forage AI select; no pre-built corpus
Modality & size Custom collection from web sources defined with the client; no size claim
Licence (and built-from) Contract; the source and provenance records you receive are agreed in the project scope
Access & cost Custom quote
Reviews (n) No public reviews listed
Watch-out Does not sell or resell datasets; timelines depend on source access and volume

We are not a data source, and we do not hold a multilingual corpus to sell you. Forage AI is a managed extraction partner: we build and run a custom collection against sources you and we define together, including the long-tail, frequently changing websites a general crawl under-samples. We own pipeline design, extraction, QA (human review plus algorithmic checks), maintenance and delivery, in CSV, JSON or XML by API or cloud integration.

In practice, we take the collection, QA, maintenance and delivery work off your team, with the provenance records you need agreed in the scope up front. If your language is in the top 30, you do not need this route. The same holds if a dataset above already covers it. For how an engagement runs, see how managed web data extraction works, or talk to us about data for AI.

Multilingual datasets compared

Here are all 20 sources side by side, on the rows that decide fit. Read the depth column for your language first. "Allowed" in the licence column means under the stated licence text, eval-only rows are not for training, and every figure is as of October 2026 in the publisher's own unit.

Source Band Languages (as stated) Tail depth Licence / commercial Best for
Common Crawl Pretraining ~160 detectable (CLD2) 122 of 161 langs <0.1% of pages CC ToU; legal advice advised Your own pipeline
FineWeb-2 Pretraining 1,868 language-script pairs 203 pairs >10K docs; Yoruba 80K docs ODC-BY + CC ToU Broadest filtered tail
HPLT 3.0 Pretraining 198 75 langs <100M tokens; Yoruba 171K docs CC0 packaging only Tail-aware scale
MADLAD-400 Pretraining 419 291 of 419 <10M tokens CC-BY-4.0 (card) / ODC-BY (hub) Audited text
CulturaX Pretraining 167 Yoruba 192 docs; Quechua 1,202 Inherits mC4 + OSCAR; gated Head-language merge
Common Corpus Pretraining Mostly EN/FR; 33 >1B tokens Probes not published Per document, commercial allowed Clean licence chain
Wikipedia dumps Pretraining 348 editions Maltese 8,023 articles CC BY-SA 4.0 + GFDL Attributable floor
OPUS Parallel 1,038 en-qu 2.88M pairs, 99.6% from NLLB Per corpus Any pair
NLLB mined bitext Parallel 148 + 1,465 pairs en-yo 1.51M pairs; contains MT ODC-BY + source terms Non-English pairs
Common Voice v27.0 Speech 295 108 of 295 <10 h CC0 + conditions Open read speech
FLEURS Speech 102 ~10 h/lang by design CC-BY-4.0 Fine-tune / eval
Multilingual LibriSpeech Speech 8 Polish 104 h (smallest) CC-BY-4.0 Deep EU speech
Aya Dataset Instruction & eval 65 Per-language counts not pulled Apache 2.0 Native instructions
FLORES+ Instruction & eval 230 varieties Eval only CC BY-SA 4.0; eval only Measurement
Hugging Face Hub Repository 12,175 tagged multilingual Per dataset Per dataset Search
LDC (+ELRA) Repository 1,598 resources (ELRA) Per corpus Contract; for-profit / VAR Contract-licensed corpora
Appen Licensed data 500+ locales (vendor claim) Not published Contract Off-the-shelf + custom
Defined.ai Licensed data 500+ (vendor claim) Per dataset on marketplace Contract Marketplace
LXT Licensed data 1,000+ locales (vendor claim) Not published Contract Wide locale list
Forage AI Build your own Defined per project Depends on sources selected Contract Corpus built for you
Forage AI promotional banner reading
Forage AI is a managed extraction partner: collection from in-language web sources, with QA and maintenance included.

How do you choose a multilingual training data source?

Answer three questions in order, then run four checks. We call it "three questions, four checks": the two rows from our rubric, depth and licence chain, turned into a decision you can make in an afternoon.

The first question is your language tier, because it decides which bands are in play for low resource languages. Joshi et al. (Microsoft Research India, ACL 2020) sorted 2,485 languages into resource classes: 2,191 sit in Class 0, "The Left-Behinds", and only 7 in Class 5, "The Winners". Maltese sits in Class 2.

Tier Roughly What exists Where to look
Top ~30 Head languages Billions of words or tokens Pretraining text; check licence and quality
Mid Below the head Thinner, uneven per source Pretraining text + parallel; check depth per language
Tail Joshi classes 0-2 Thousands of documents; minutes to hours of speech Licensed language data, or build your own

The second question is modality: pretraining text, parallel, speech, instruction or evaluation. The third is licence need, research only or commercial, and then you run four checks on every candidate that survives:

  1. Depth for your language: open the per-language file (the table in the repositories band shows where). Search every ISO code your language goes by.
  2. Licence chain: the dataset licence plus upstream terms, such as the Common Crawl ToU or NLLB's "original source" terms.
  3. Native vs translated: the Aya Collection is 95.9% translated; the Aya Dataset is native. Know which you have.
  4. Eval contamination: check whether FLORES+, Global-MMLU or your own test sets are inside your training mix.

The depth check can be a few lines of code, and here's how it actually works on a tail language. This streams FineWeb-2's Maltese subset and spot-checks language-ID scores without a full download. The 0.8 threshold is illustrative, not the dataset authors' recommendation.

from datasets import load_dataset
from itertools import islice

# Maltese (a Joshi "Class 2" language) from FineWeb-2, streamed: no full download
ds = load_dataset("HuggingFaceFW/fineweb-2", name="mlt_Latn",
                  split="train", streaming=True)

sample = list(islice(ds, 200))
low = [r for r in sample if r["language_score"] < 0.8]
print(f"{len(low)}/{len(sample)} docs below 0.8 language-ID confidence")
for r in sample[:5]:
    print(r["url"], "|", r["text"][:120].replace("\n", " "))
# Then do what Kreutzer et al. did: have a reader judge ~100 lines by hand.

The last comment is the step teams skip. A score is a model's opinion; 100 lines read by a fluent speaker is evidence. If your languages are all top-30 and research licences suffice, the four checks shrink to one: the licence chain.

Verdict: the best multilingual data source for your languages

The best multilingual training data source is the one with real depth in your languages and a licence chain you can read end to end; the language count is the weakest signal.

  • Top-30 language, pretraining: FineWeb-2 or HPLT 3.0.
  • Licence-clean text: Common Corpus (per-document open licences) or Wikipedia (CC BY-SA 4.0).
  • Speech: Common Voice for training, FLEURS for fine-tuning and evaluation; check hours first.
  • Translation pairs: OPUS, checking each corpus's licence.
  • Instruction tuning: the Aya Dataset (native, human-written).
  • Evaluation: FLORES+ and Global-MMLU, never in training.
  • Tail language with commercial use: a licensed language-data vendor, or a corpus built from in-language web sources with a managed extraction partner such as Forage AI.

Re-read the card at every release, because counts and licences change, and we re-check this list each edition.

Forage AI promotional banner reading
Forage AI owns pipeline design, extraction, QA, maintenance and delivery for a corpus built from web sources.

Frequently asked questions

Which multilingual datasets cover low-resource languages?

FineWeb-2, HPLT 3.0 and MADLAD-400 reach furthest, but thinly: only 203 of FineWeb-2's 1,868 pairs have more than 10K documents, and 291 of MADLAD-400's 419 languages have under 10M tokens. Coverage varies widely, with Yoruba ranging from 192 documents in CulturaX to 171,248 in HPLT 3.0.

Are mC4 and OSCAR licensed for commercial training?

Neither is public domain, and neither gives a clean commercial grant on its own. mC4 is ODC-BY, and you are also "bound by the Common Crawl terms of use". OSCAR 23.01 waives rights only on the packaging and is currently suspended on Hugging Face. Take legal advice before commercial use, as Common Crawl's own terms recommend.

Where can I find multilingual datasets on Hugging Face?

Filter by the multilinguality:multilingual tag, which returned 12,175 datasets on 8 October 2026. Then read each card's licence text and per-language table. Tags and sidebar licences are self-reported, and audits found hub licence fields missing more than 70% of the time.

What are the best multilingual speech datasets for ASR?

Common Voice v27.0 covers the most languages (295, CC0 with conditions, via Mozilla Data Collective), but 108 have under 10 validated hours. FLEURS gives about 10 hours in each of 102 languages for fine-tuning or evaluation, and Multilingual LibriSpeech is deep in 8 European languages. Watch Emilia's non-commercial licence.

What is a parallel corpus, and when do you need one?

A parallel corpus is sentence-aligned text in two or more languages, used for machine translation and cross-lingual alignment. OPUS indexes 1,038 languages and 102.9B sentence pairs with a licence per corpus, and NLLB mined bitext covers non-English-centric pairs, with part of them machine translated.

Should I use an open dataset or an AI training data company?

Start with open datasets when your language is in the top 30 and research or attribution licences fit. Go to a vendor when you need a commercial licence, a tail language, or speech hours open sets lack. Either way, prioritize two asks: per-language volumes and a sample. Of the three vendors here, only Defined.ai publishes per-dataset volumes publicly.

What if no public dataset covers my language?

Search every ISO code and script your language goes by, then check Wikipedia, which for several tail languages holds more text than the filtered corpora keep (24,641 Quechua articles against 2,903 kept FineWeb-2 Ayacucho documents). If that still falls short, ask a licensed language-data vendor for per-language volumes and a sample, or build a corpus from in-language web sources. The MaCoCu project's national-domain crawl collected 347.9M Maltese words in its first release, comparable to or more than the 286.5M words FineWeb-2 keeps.

Can I scrape the web to build a multilingual corpus?

Whether collection is lawful depends on jurisdiction, the type of data and each site's terms, and it is not settled, so take legal advice for your case. Treat robots.txt and site terms as a floor, not a legal answer. Where personal data is in scope, France's CNIL has said that ignoring robots.txt or CAPTCHAs defeats the legitimate-interest basis, and since July 2025 Cloudflare asks every new domain whether to allow AI crawlers. Keep a provenance log per document so you can show what you collected and under which terms.

S
Written by
Sai Subramaniam
Data Infrastructure Enthusiast, Forage AI

Sai is a data infrastructure enthusiast who has spent the past two to three years following the AI space closely, from the infrastructure layer to the fast-growing world of data for AI. He is genuinely curious about how modern data pipelines get built and where the data industry is heading, and he writes insightful pieces on the core topics that shape this niche.

Reviewed by the team of experts at Forage AI for accuracy and clarity.