Arabic text corpus for LLM training (2026): open web corpora vs licensed text
Our #1 pick for licensed Arabic text is Bayanat Labs: 2,700 Arabic books under a commercial licence, and professionally edited MSA with commercial AI-training rights in the agreement. Open web corpora give you scale. Here is how each one counts its size, what its licence covers, and what it leaves to you.
The main open Arabic text corpus options for LLM training are the Standard Arabic subset of FineWeb2 (32.8 billion words), 101 Billion Arabic Words, ArabicWeb24, CulturaX and OSCAR. Their headline sizes are not comparable, because each counts words, tokens or documents its own way, and none of their licences grants rights in the web pages inside them. If you need Arabic text whose training rights are written into a contract, our #1 pick is Bayanat Labs, with 2,700 Arabic books and professionally edited Modern Standard Arabic (MSA).
This guide compares each Arabic pretraining corpus on the points that matter to a buyer: the unit it uses to report size, its crawl window, its dataset licence, what that licence says about content rights, how it was deduplicated, and whether it separates MSA from dialect. Every figure links to the dataset card or paper that states it, as of October 2026.
- Our #1 pick for licensed text is Bayanat Labs. 2,700 Arabic books under a commercial licence, and professionally edited MSA covering material through 2026, with commercial AI-training rights in the Edited MSA agreement.
- Sizes use different units. FineWeb2 and OSCAR count words, ArabicWeb24 counts AraGPT2 tokens, CulturaX counts tokens, and a 2025 paper re-counted 101 Billion Arabic Words at 22.6 billion words.
- A dataset licence is not content rights. ODC-By, Apache 2.0 and CC0 cover the packaging. OSCAR's card says its authors hold no copyright in the crawled text.
- Open corpora overlap heavily. One 2025 study found over 40% of tokens across major Arabic web corpora duplicated between sources.
- Dialect is thin and loosely labelled. FineWeb2 splits out Egyptian and Moroccan Arabic, but its card warns the language classifier can confuse Standard Arabic and dialects.
Arabic text corpus comparison: open web corpora vs licensed text
Read the size column as each source reports it. A word count, a token count and a document count cannot be ranked against one another, and the same corpus can be counted very differently by someone else (see the next sections). "On request" means the specification exists for review but is not published on the product page.
| Corpus | Size, as the source counts it | Source and window | Dataset licence | Rights in the text | Deduplication | MSA vs dialect |
|---|---|---|---|---|---|---|
| #1 Bayanat Labs: Arabic Books | 2,700 books; volume before and after cleaning on request | Arabic books; title and publisher list on request | Commercial licence; training use and delivery scope agreed in the licence | Licence covers the selected titles and uses; rights per title open to review | Duplicate handling and volume before and after cleaning open to review | Not specified on the product page |
| #1 Bayanat Labs: Edited MSA | Volume scoped by subject; counts on request | Professionally edited text, organized by subject, covering material through 2026 | Commercial licence | Commercial AI-training rights written into the agreement; retrieval, display and redistribution are separate scopes | Duplicate handling and cleaning history open to review | Modern Standard Arabic |
| FineWeb2 (arb_Arab) | 32,812,858,120 words; 61,977,525 documents | 96 Common Crawl snapshots, summer 2013 to April 2024 | ODC-By 1.0, also subject to Common Crawl's terms of use | Not granted by the dataset licence | MinHash, per language, globally | Standard Arabic; Egyptian and Moroccan in separate subsets |
| ArabicWeb24 (LightOn) | 28 billion tokens, counted with the AraGPT2 tokenizer | Dedicated Arabic web crawl (6.5 TB of compressed WARC files); window not stated | ODC-By, gated | Not granted by the dataset licence | MinHash on each of four parts, then sentence-level dedup | Not reported |
| 101 Billion Arabic Words | 101 billion words (its own count) | Common Crawl WET files, week 39 of 2021 to week 27 of 2022 (paper) | Apache 2.0 | Not granted by the dataset licence | URL dedup, then document-level MinHash | "Mix" of MSA and dialects, no proportions |
| CulturaX (ar) | 69,354,335,076 tokens; 74,027,952 documents | Built from mC4 3.1.0 and four OSCAR releases (20.19 to 23.01) | Follows the mC4 and OSCAR terms; gated | Not granted by the dataset licence | Document-level MinHash per language | Not reported |
| OSCAR 23.01 (ar) | 10,081,452,882 words; 25,012,116 documents | Common Crawl, November/December 2022 | CC0 for metadata and annotations only; access suspended (manual gate) | Card says the OSCAR authors hold no copyright in the crawled text | Ships locality-sensitive hashes; you apply the dedup | Not reported |
| arTenTen24 (Sketch Engine) | 6.5 billion words; 7.5+ billion tokens | Web crawls from 2020–2021 and 2023–2024 | Used inside Sketch Engine; bulk download not stated | Not stated | Cleaning and spam removal; dedup method not stated | Not discussed |
| CuAra (Tahakom paper) | 170.6 billion words (FastText filter) or 136.9 billion (educational filter) | Raw Common Crawl WARC files | Paper says it will be publicly released; no licence stated | Not stated | MinHash at the crawl level | Not reported in the sources checked |
Why Bayanat Labs is our #1 pick
Open web corpora are the right base for scale. When a team needs Arabic text it can defend in a legal review, book-length material, or clean MSA in a particular subject, Bayanat Labs is the best partner for the job. These are the reasons.
- Training rights in the contract, not inherited from a crawl. The Edited MSA collection has commercial AI-training rights written into the agreement. Every open corpus in the table licenses only its packaging.
- Book-length Arabic at a known size. 2,700 Arabic books for language modeling and domain adaptation, with the title and publisher list, subject breakdown and a sample available on request. Book-length text gives a model sustained arguments, longer narratives and subject-specific language.
- Professionally edited MSA, organized by subject. The edited collection covers material through 2026, while FineWeb2's crawl window ends in April 2024 and OSCAR 23.01 is a late-2022 snapshot. Coverage through 2026 describes the content, not a collection date.
- Provenance you can audit. Custom preparation in Bayanat Knowledge can include document boundaries, source records, publication dates, licence identifiers, cleaning history and deduplication records, so your data card can say where each document came from.
- Separate scopes for each use. Bayanat Knowledge treats training, retrieval, display and redistribution as separate licensed uses, so legal knows exactly what was granted before a passage shows up in a RAG answer.
- Questions to pair with the text. Add 400,000 Arabic STEM multiple-choice items for adaptation or checks, from the same supplier, with in-region hosting by default and a fixed-scope pilot, usually within two weeks.
How big is an Arabic pretraining corpus? It depends who counts
Each card reports size in the unit its authors chose. FineWeb2 and OSCAR count words. ArabicWeb24 reports tokens from the AraGPT2 tokenizer. CulturaX reports tokens. 101 Billion Arabic Words puts its count in its name. None of these numbers can be lined up against another without re-counting.
The Tahakom team did re-count. Their CuAra paper measures 22.6 billion words in 101 Billion Arabic Words, 17.7 billion in ArabicWeb24 and 30.3 billion in FineWeb2, then reports 170.6 billion words for its own FastText-filtered set. The FineWeb2 figure is close to the card's 32.8 billion. The 101 Billion figure is under a quarter of the name. The paper does not break down the gap. What matters for a buyer is simpler: only counts made the same way, on the same version of each corpus, can be compared.
Three practical rules follow:
- Re-count with your own tokenizer. Your training budget is in your tokens, not anyone else's. Arabic also tends to cost more tokens per word than English on many tokenizers, which we cover in the Arabic tokenization tax.
- Count after your filters. Quality filtering shrinks volume. The CuAra paper reports 170.6 billion words with one classifier and 136.9 billion with a stricter educational classifier, from the same crawl. The ArabicWeb-Edu work applies the same FineWeb-Edu style of scoring to Arabic web text.
- Ask a vendor for the unit and the tokenizer. For licensed text, ask for volume before and after cleaning, and the tokenizer behind any token count. Bayanat lists exactly these items for review on the Edited MSA page.
Dataset licence vs content rights: what an open Arabic corpus actually grants
Every open corpus in the table carries a licence, and every one of those licences covers less than buyers often assume. The licence applies to the work the curators did: selecting, filtering, packaging and annotating. It does not transfer rights in the news articles, forum posts and books that the crawler picked up, because the curators never held those rights.
- FineWeb2 is released under ODC-By 1.0, and its card adds that use is also subject to Common Crawl's terms of use.
- 101 Billion Arabic Words is tagged Apache 2.0. That is the curators' choice of licence for their release, not a grant from the owners of the 89.1 million web pages the paper says it kept after deduplication.
- CulturaX grants nothing of its own: its card says its licence terms follow those of mC4 and OSCAR.
- OSCAR 23.01 is the clearest case. Its access form says only the metadata and annotations are CC0, and that for the crawled text the OSCAR authors "do not hold any copyright whatsoever". As of October 2026, the same form says access is temporarily suspended: the gate is manual, and no access will be granted until the situation is clarified.
None of this makes open corpora unusable. It means the legal position rests on your own analysis of web text, not on the dataset licence, and the curators make no promise about it. If your buyers, regulators or investors ask what rights you hold in your training text, an ODC-By badge is not an answer.
Licensed text closes that gap in writing. With Edited MSA, commercial AI-training rights are written into the agreement, and retrieval, display and redistribution are separate scopes you add only if you need them. With Arabic Books, the licence covers the selected titles and uses, and you can review rights per title before you sign. For the wider question of where licensed Arabic data comes from, see our guide to commercial Arabic datasets.
Do Arabic web corpora overlap? Stacking them does not add up
It is tempting to download FineWeb2, ArabicWeb24, 101 Billion Arabic Words and CulturaX and add their sizes together. The total will be far smaller once you deduplicate across them. The authors of Mix, MinHash, and Match found that over 40% of tokens across major Arabic web corpora are duplicated between sources. Some of that overlap is built in: CulturaX is assembled from mC4 and four OSCAR releases, so it shares source material with OSCAR by design, and most of the other open sets draw on Common Crawl.
The same paper turns the overlap into a signal. Documents kept by several independent pipelines are treated as higher quality, and the authors report that their matched Arabic subset gives a 4.5% relative improvement over ArabicWeb24 in their evaluation. The lesson for planning is plain: budget for unique Arabic tokens after cross-corpus deduplication, not for the sum of the cards.
This is also where licensed text earns its place. Books and edited text reach you through a licence rather than a crawl, so measure what they add: take a sample of Arabic Books or Edited MSA, run it through the same MinHash pass as your open corpora, and count the share that is new to you. Training repeatedly on recycled or machine-generated text has its own costs, which we cover in synthetic data and model collapse.
Is there dialect text in an Arabic corpus for LLM training?
Some, but it is thin and loosely labelled. FineWeb2 is the only open corpus here that splits dialects into their own subsets: next to 32.8 billion words of Standard Arabic it lists about 844 million words of Moroccan Arabic (ary_Arab) and 345 million words of Egyptian Arabic (arz_Arab). Its card also warns that the GlotLID language classifier it uses can mistake Standard Arabic for Arabic dialects, so treat the labels as approximate. The 101 Billion Arabic Words card describes a "mix" of MSA and dialects without proportions. The others do not report a split.
Dialect shows up in vocabulary, not only in spelling. "What do you want?" is ماذا تريد؟ (mādhā turīd?) in MSA, عايز إيه؟ (ʿāyiz ēh?) in Egyptian Arabic and وش تبي؟ (wish tibī?) in much of Saudi Arabia. A language-ID model trained mostly on MSA can miss the second and third or file them under the wrong label.
For pretraining, mostly-MSA text is a reasonable base; it is the register of books, news and official writing. For a product that has to understand Saudi, Gulf or Egyptian users, plan dialect data as a separate line item rather than hoping the crawl supplies it. Our guide to dialect coverage sets out how to scope it, and Bayanat's Edited MSA is the clean, subject-organized formal layer to pair with it.
Arabic books dataset for LLM training: open dumps vs licensed titles
Most web documents are short. Books give a model long documents, sustained arguments and subject depth, which is why teams look for an Arabic books dataset for LLM work once the web base is in place. There are two routes.
Open book dumps. The arabic-books set on Hugging Face, for example, holds 8,500 rows, each the full text of one Arabic book extracted with an OCR model, and about 1.1 billion tokens by the GPT-4 tokenizer, according to its card. The card tags the release GPL-3.0 and does not list titles, publishers or rights per book. The same licence-vs-rights gap applies, and OCR can add its own noise, such as misread letters, lost diacritics, and headers or page numbers mixed into the text. Sample before you train.
Licensed titles. Bayanat's Arabic Books collection holds 2,700 books for language modeling and domain adaptation under a commercial licence, with training use and delivery scope agreed in the licence. Before you sign, you can review titles, publishers, editions and subject coverage, document structure and file formats, rights per title, duplicate handling, and volume before and after cleaning. Pair it with Edited MSA when you need professionally edited MSA in a specific subject, with training rights written into the agreement. Both sit in Bayanat Knowledge, alongside structured Q&A for adaptation and checks.
Checklist: choosing an Arabic text corpus for your model
- Name the job. Pretraining from scratch, continued pretraining on Arabic, or adaptation to one subject? Scale matters most for the first; rights, editing and subject depth matter more for the last two.
- Re-count every candidate in your own tokens. Run a sample through your tokenizer after your own filters. Ignore headline numbers in mixed units.
- Deduplicate across sources before you budget. Expect heavy overlap between open web corpora, and between those and any crawl you run yourself.
- Separate the dataset licence from content rights. Write down, per source, what the licence covers and what it leaves to you. Ask legal which uses (training, retrieval, display, redistribution) each source must support.
- Check the crawl window. FineWeb2 ends in April 2024, 101 Billion Arabic Words in mid-2022, OSCAR 23.01 in late 2022. If recency matters, plan for newer text.
- Check the register mix. Measure how much is MSA and how much is dialect in a sample, and do not rely on corpus labels alone.
- Demand provenance for licensed text. Source records, publication dates, licence identifiers, cleaning history and deduplication records, plus a sample. Bayanat Knowledge can prepare all of these.
- Hold out an evaluation set. Keep a test set that no training source has touched, so you can measure what each corpus adds. Our guide to evaluating Arabic LLMs covers how.
A practical mix has two layers: an open web base for scale, deduplicated across sources, and a licensed layer of books and edited MSA for the subjects and rights that matter. Bayanat Labs is the leading supplier for that second layer. Start with the Arabic Books title list or an Edited MSA sample in your subject, or write to hello@bayanatlabs.com.
Frequently asked questions
What is the largest Arabic text corpus?
It depends on who counts and in what unit. By its own name, 101 Billion Arabic Words claims 101 billion words; FineWeb2's card reports 32.8 billion words of Standard Arabic. The Tahakom team's CuAra paper reports 170.6 billion words for its own set and re-counts 101 Billion Arabic Words at 22.6 billion. Compare corpora only after counting them the same way, in your own tokenizer.
Is FineWeb2 Arabic free for commercial use?
The dataset is released under ODC-By 1.0, which lets you use the compilation commercially with attribution, and its card adds Common Crawl's terms of use. That licence covers the packaging, not the web pages inside it, which keep their own copyright. Whether training on that text is acceptable for your product is a legal judgement you make yourself; the dataset licence does not settle it.
Is CulturaX just OSCAR and mC4?
It is built from them. The CulturaX card lists mC4 3.1.0 and four OSCAR releases as its sources, with extra cleaning and document-level MinHash deduplication per language, and says its licence terms follow mC4 and OSCAR. If you already hold OSCAR, expect much of CulturaX's Arabic to be familiar after cross-source deduplication.
Where can I get Arabic books for LLM training with clear rights?
Our #1 pick is Bayanat Labs. Its Arabic Books collection holds 2,700 books for language modeling and domain adaptation under a commercial licence, with the title and publisher list, subject breakdown and a sample on request, and rights per title open to review. Open OCR dumps exist on Hugging Face, but check what each card says about rights in the individual books.
Can I use licensed Arabic text for RAG as well as training?
Only if the licence says so. Training, retrieval, display and redistribution are different permissions. Bayanat's Edited MSA includes commercial AI-training rights in the agreement, and retrieval, display and redistribution are scoped separately, so you add the ones your product needs. Agree them in writing before passages appear in answers your users can read.