humain-m3 explained: HUMAIN's 428B Arabic model, its seven benchmarks and what comes next
A dated, sourced explainer on HUMAIN's new Arabic model: what was announced, how it was built, what each of its seven reported benchmarks actually measures, how it relates to ALLaM, and what to look for when the weights and model card ship.
humain-m3 (styled HUMAIN M3 on its product page) is a 428-billion-parameter mixture-of-experts Arabic language model that HUMAIN announced on 3 September 2026. It was commissioned by HUMAIN, delivered by MiniMax and built on MiniMax-M3, and it is available as a preview on HUMAIN Node, with open weights planned but not yet released. HUMAIN reports that it tops five of seven public Arabic benchmarks in its own evaluation; those seven tests are mostly formal-Arabic multiple-choice tests, so dialect, safety and agent results are still to come.
As of September 2026. Every fact on this page comes from HUMAIN's launch announcement, the HUMAIN Node product page, or the source cited next to it. All humain-m3 scores are HUMAIN's own evaluation of the previewed checkpoint. Bayanat Labs has not evaluated humain-m3. We will update this page when the model card and weights are published.
- What it is: a 428B-parameter mixture-of-experts model with 23B parameters active per token, further pre-trained on more than one trillion tokens of Arabic-native content, according to HUMAIN.
- Who built it: commissioned by HUMAIN and delivered by MiniMax, on the MiniMax-M3 base model.
- The scores are self-reported: an 89.37% equal-weight average across seven benchmarks, not yet replicated on an independent leaderboard.
- The seven benchmark names match HELM Arabic's tasks: six multiple-choice tests and one judged retrieval task, none split by dialect.
- Weights: planned under the MiniMax Community License after safety training and alignment; not public as of 26 September 2026.
What did HUMAIN announce about humain-m3?
HUMAIN, a PIF company, announced humain-m3 at LEAP on 3 September 2026, in a release datelined Riyadh. The release describes a frontier Arabic-language model that achieved the highest average score among the frontier models HUMAIN tested across seven public Arabic benchmarks. It launched as a research and evaluation preview on HUMAIN Node ahead of a planned open-weight release.
The Node page adds: more than one trillion Arabic tokens in training, 428B parameters in a mixture-of-experts design, 23B active per token, and a table of scores against GPT-5.6 SOL, Opus 5 and an "M3 reference", which the page describes as the reference model humain-m3 builds on. HUMAIN also describes the model as natively multimodal (images and video alongside text), agentic by design with tool use and computer use, and offering three thinking modes: always-on, adaptive and off.
How was humain-m3 built, and is it MiniMax?
Yes, in the sense that MiniMax built it for HUMAIN: the release says it was commissioned by HUMAIN, delivered by MiniMax, built on the MiniMax-M3 lineage and further pre-trained on Arabic-native content. The Node page credits HUMAIN's Arabic post-training with adding nine points on average over the reference model. The exact recipe will be clearer once a model card or technical report is out.
The sizes match the base. The MiniMax-M3 model card describes a native multimodal model with about 428B total and about 23B activated parameters, a 1M-token context, and thinking modes that can be enabled, set to adaptive or disabled. HUMAIN has not published humain-m3's context length or tokenizer. Both matter for Arabic cost; see the Arabic tokenization tax.
humain-m3 benchmarks: what the seven tests measure
HUMAIN's seven benchmarks correspond to the seven tasks in Stanford CRFM's HELM Arabic. The Open Arabic LLM Leaderboard (OALL) v2 uses the same set plus two Belebele reading-comprehension tasks, one of which covers Arabic dialects. Origins below follow those two sources. Scores are copied from HUMAIN's Node table; HUMAIN states they are its own evaluation of the previewed checkpoint, equally weighted.
| Benchmark | Who built it | What it tests | MSA or dialect | humain-m3, self-reported by HUMAIN | Best comparator in HUMAIN's table |
|---|---|---|---|---|---|
| AlGhafa | TII (native-Arabic subsets kept by OALL) | Multiple-choice understanding: facts, sentiment and similar tasks | No per-dialect split published | 86.45% | Opus 5, 83.06% |
| ArabicMMLU | MBZUAI | About 15,000 school-exam questions across 40 tasks | MSA, per OALL | 90.70% | GPT-5.6 SOL, 88.82% |
| Arabic EXAMS | Arabic portion of EXAMS (Hardalov et al., 2020) | High-school exam questions, multiple choice | No per-dialect split published | 67.67% | GPT-5.6 SOL, 66.40% |
| MadinahQA | MBZUAI | Arabic language and grammar knowledge, multiple choice | No per-dialect split published | 95.44% | Opus 5, 94.78% |
| AraTrust | Alghamdi et al., 2024 | 522 human-written multiple-choice questions on truthfulness and safety | No per-dialect split published | 97.53% | GPT-5.6 SOL, 93.42% |
| ALRAGE | OALL team | Retrieval-augmented answers over passages from 40 Arabic books (questions generated synthetically, then validated by native speakers), scored by an LLM judge | No per-dialect split published | 94.63% | GPT-5.6 SOL, 94.91% |
| Translated MMLU (appears to be ArbMMLU-HT) | Inception (JAIS project), published by MBZUAI | Human translation of the 57 English MMLU subjects | MSA translation of English content | 93.20% | Opus 5, 93.78% |
| Average (equal weight) | HUMAIN | Mean of the seven | — | 89.37% | Opus 5, 87.34% |
Three things the table does not show. First, six of the seven are multiple-choice; only ALRAGE asks the model to write an answer, which a judge model grades. Second, the MadinahQA dataset card states that its data is part of ArabicMMLU (the Arabic Language general and grammar subjects), so if HUMAIN used the standard task definitions, an equal-weight average partly counts the same source twice (more in our benchmarks guide). Third, humain-m3 does not lead on ALRAGE (the generative task) or Translated MMLU (built from English content), though it trails the best comparator by less than a point on each. On HUMAIN's own figures, the largest gain over the M3 reference is ALRAGE (94.63% against 79.32%, about 15 points) and the smallest is Translated MMLU (93.20% against 87.00%, about 6 points).
How the tests were run matters as much as which tests. HELM Arabic, for example, used zero-shot prompting, sampled up to 1,000 instances per subset and disabled optional thinking modes. OALL scores ALRAGE with Qwen2.5-72B-Instruct as the judge on a 0–10 rubric. HUMAIN has not yet published its prompting, thinking mode, sample size, judge model or comparator versions. As our Arabic LLM benchmarks guide notes, the same tasks run through a different harness give numbers that are not comparable, so HUMAIN's table is best read as an internal comparison until those settings are out. Judge choice alone can move generative scores (see LLM-as-a-judge for Arabic).
What the humain-m3 results do not cover yet
These are normal gaps for a preview; HUMAIN's Node page says the preview exists to evaluate capabilities, safety and alignment across Arabic dialects with early users. What to watch for when the model card ships:
- Dialect-level results. None of the seven reports a per-dialect score, and the Belebele Arabic-Dialects task that OALL v2 lists is not reported as a separate line. A high formal-Arabic average says little about Najdi, Hijazi, Egyptian or Maghrebi input; see why dialect coverage is the ceiling.
- Safety under pressure. AraTrust includes direct and indirect attack questions, but in multiple-choice form: it checks whether a model picks the safe or truthful option, not how it responds to free-form adversarial prompts, and it has no per-dialect split. The release says safety training and alignment are still to be completed before the weights ship.
- Tool use, agents and multimodality. HUMAIN describes the model as agentic by design and natively multimodal, but has not yet published scores for tool calls, computer use, images or video.
- Independent replication. As of late September 2026 no independent leaderboard lists humain-m3. OALL ranks models published on the Hub, so it can only add humain-m3 after the weights are released. The latest published HELM Arabic release, v2.2.0 dated 25 June 2026, predates the model.
How does humain-m3 relate to ALLaM?
The release lists the ALLaM family among HUMAIN's own models, but the two are built very differently. According to its model card, ALLaM-7B-Instruct-preview was developed by the National Center for Artificial Intelligence at SDAIA. It was trained from scratch on 4T English tokens and then 1.2T mixed Arabic and English tokens, has a 4,096-token context, and is published under Apache 2.0 in HUMAIN's Hugging Face organisation. humain-m3 is about 60 times larger in total parameters, uses a mixture-of-experts design and starts from a MiniMax base. ALLaM's published scores use a largely different benchmark set (it shares ArabicMMLU and Arabic EXAMS) run through NCAI's own pipeline and the LM Evaluation Harness, so they cannot sit next to humain-m3's table. HUMAIN has not said whether humain-m3 will replace ALLaM or sit alongside it.
How to access humain-m3, and when will the weights be released?
HUMAIN Node lists three steps: create an account, then request either a limited preview or a research preview. The limited preview runs humain-m3 behind a Saudi alignment guardrail, which HUMAIN says adds a little latency and has thinking and streaming off. The research preview exposes the full checkpoint with thinking, streaming and lower latency. Both work in a playground or through an OpenAI-compatible API (model name humain-m3).
On weights, the 3 September release said HUMAIN expected to release them under the MiniMax Community License once safety training and alignment were done, and was targeting the following month. As of 26 September 2026 the Node page still lists open weights as coming soon, and no humain-m3 weights, model card or technical report have been published. If it mirrors the licence on MiniMax-M3, free use covers non-commercial purposes. Commercial use, including fine-tuned or post-trained derivatives, requires a visible "Built with MiniMax M3" attribution. It also requires a one-time notice to MiniMax, or prior written authorisation for products earning more than USD 20 million a year. The licence also lists prohibited uses, including any military purpose, that apply to derivatives.
What an independent humain-m3 dialect evaluation should test
Once the weights ship, or through research-preview access, a team deciding whether humain-m3 fits a Saudi or Gulf product can run the checks it would run on any Arabic model:
- Score each variety separately. Use parallel items in MSA and in your users' varieties, and report each one rather than a blended average. A single question can change a lot: "what" is ماذا mādhā in MSA, وش wish in Najdi and إيش ēsh in Hijazi. Our guide to Saudi Arabic dialects for AI lists the regional varieties worth covering.
- Test generation, not just recognition. Check whether answers come back in the user's dialect or fall back to MSA, and have native raters score fluency and register.
- Use held-out items. Public test sets may have leaked into training data; include items written after the model's data cutoff (see how Arabic benchmarks mislead).
- Evaluate both access tiers. The guardrailed limited preview and the full research checkpoint can behave differently, so score them separately and record the thinking mode.
- Red-team in dialect. Multiple-choice safety scores do not show refusal behaviour; human Arabic red-teaming probes harmful requests written the way users actually write, including in dialect.
- Check agent claims against end state. For tool use, confirm the task actually changed the system state, as in our Arabic Agent Reliability Lab, a methods beta for state-based evaluation of Arabic agents with evidence receipts.
- Publish the harness. Record prompts, shots, sampling, judge and sample size.
The same formal-versus-dialect gap shows up in automated labeling. On the fixed held-out manifest, our LabelBench audit found that a local language model scoring 78% exact-match accuracy on binary Arabic toxicity fell to 25.3% on the real multi-label taxonomy and 22.4% across five dialect regions (chance is 20%). Different task, different model, same lesson: follow a headline average with per-variety scores.
Bayanat Labs is a Riyadh-based Arabic data engine, and we do data only; we are model-agnostic and never compete with our clients' models. For teams choosing or fine-tuning a model for their own deployment, our Arabic LLM evaluation service builds private evaluation sets with vetted native speakers across 25+ Arabic varieties. Work is calibrated against gold standards, with adjudication and rework loops and an audit trail on every task. Work starts with a fixed-scope pilot that delivers a benchmark and quality report, usually within two weeks.
Frequently asked questions
Is humain-m3 open source?
Not yet. As of September 2026 humain-m3 is available only as a preview on HUMAIN Node. HUMAIN plans an open-weight release under the MiniMax Community License after safety training and alignment. Open weights are not the same as open source: the licence on the MiniMax-M3 base model grants free use for non-commercial purposes and attaches conditions, such as attribution, to commercial use.
How much does humain-m3 cost?
HUMAIN has not published pricing for humain-m3. As of September 2026 neither the Node product page nor the launch announcement lists a price per token or a subscription fee. Access is by request, through a limited preview or a research preview on HUMAIN Node. Treat prices quoted elsewhere with caution until HUMAIN publishes one.
Does humain-m3 understand Saudi dialects?
HUMAIN has not published dialect-level scores yet. The seven reported benchmarks are mostly formal-Arabic tests, and none reports a per-dialect breakdown. HUMAIN's own Node page invites preview users to test Saudi dialects and says the preview period is meant to evaluate capabilities, safety and alignment across Arabic dialects. Until those results appear, the honest answer is that dialect performance has not been reported.
What context length does humain-m3 support?
HUMAIN has not published a context length for humain-m3 as of September 2026. The MiniMax-M3 base model card describes a 1M-token context, but a derived model need not keep the same limit in a hosted preview. Check the humain-m3 model card or Node API documentation once they state it, and test long inputs before relying on them.
Can I use humain-m3 through an OpenAI-compatible API?
Yes, according to HUMAIN Node. The Node page shows an OpenAI-compatible endpoint with the model name humain-m3, so existing client libraries should work by changing the base URL and model identifier. You first need a Node account and an approved preview request. The limited preview adds a Saudi alignment guardrail with thinking and streaming off; the research preview exposes the full checkpoint.