Explainer · Agents

Arabic function-calling and agent datasets (2026): what is public, and what is missing

Most public Arabic tool-use data teaches a model to choose one function and fill its arguments. Here is what each set contains, where Arabic argument values break, and what a multi-step agent trajectory needs that these sets do not provide.

Bayanat Labs Research· · 10 min read

In short: yes, public Arabic function calling datasets exist, and the large Hugging Face sets were all first uploaded between November 2025 and May 2026. The biggest, HeshamHaroon/Arabic_Function_Calling (50,810 GPT-4o-generated samples) and the AISA-ArabicFC shared-task set (10,550 training queries), teach one step: decide whether a tool is needed, then choose a function and fill its JSON arguments. Where these sources describe multi-turn or multi-step Arabic tool-call data, it is machine-translated from English or a small hand-written test set, and none of the sources below describes native Arabic agent trajectories that record tool observations and user corrections, then score success by checking the resulting state.

Key takeaways
  • The field is young. A 2025 review of Arabic post-training data on the Hub called public function-calling resources nonexistent; the first large set arrived in November 2025.
  • One dataset, many listings. At least four other Hub listings name HeshamHaroon's GPT-4o set as their source or reuse its card. Count them once.
  • Arguments are the hard part. In the shared task's pilot baselines, function accuracy reaches 0.982 but argument exact match tops out at 0.541.
  • Trajectories are the gap. No source here describes native Arabic multi-step runs with observations, corrections and a state check.

Is there an Arabic function calling dataset?

Yes, but only recently. A review of Arabic post-training datasets on the Hugging Face Hub (first titled "Mind the Gap"; v2 30 September 2025) analysed 366 datasets across 12 domains and described function calling as "currently nonexistent in public Arabic post-training resources."

Since then (as of October 2026):

  • September 2025 (arXiv). QCRI researchers (Ersoy et al., ArabicNLP 2025) machine-translated two English tool-calling sets into Arabic. The paper links a QCRI collection for its training set; the collection lists no items.
  • 27 November 2025. HeshamHaroon/Arabic_Function_Calling is first committed to the Hub.
  • February–March 2026. AISA-AR-FunctionCall restructures that data and trains a 270M FunctionGemma model on it (Nacar et al., arXiv 2603.16901).
  • May–July 2026. Tuwaiq Academy launches the AISA-ArabicFC shared task, with train and dev data released on 1 June and a blind test on 20 July.

Besides AISA-AR-FunctionCall, Sa74ll/arabic-mobile-actions (45,729 samples) and metga97/arabic-tooluse-functiongemma-v1 name HeshamHaroon's set as their source, and Sellopale/Arabic_Function_Calling carries largely the same card text.

Function call vs agent trajectory: what is the difference?

A function-calling record maps one request to one expected call. A trajectory records everything the agent did until the task was done, then checks the result. The example below is illustrative, not drawn from any dataset. A Saudi-style booking: أبي أحجز موعد عند الدكتور بكرة الساعة ٤ العصر (abi aḥjiz mawʿid ʿind ad-diktōr bukra as-sāʿa arbaʿ al-ʿaṣr, "I want to book a doctor's appointment tomorrow at 4 in the afternoon").

FieldSingle-turn function-call record (illustrative)
queryأبي أحجز موعد عند الدكتور بكرة الساعة ٤ العصر
functionbook_appointment
arguments{"date": "2026-10-02", "time": "16:00"}

بكرة ("tomorrow") only becomes 2026-10-02 if the record carries a reference time and time zone. The AISA-ArabicFC prompt includes the current time; the HeshamHaroon schema has no such field. Now the same request as a trajectory, reference time 1 October 2026, 20:00, Asia/Riyadh:

  1. Request as above.
  2. Tool call: search_slots(date="2026-10-02", after="15:00"). Observation: 16:00 is taken; 16:30 and 18:00 are free.
  3. Agent message: offers 16:30 or 18:00. The task rules require confirmation, so nothing is booked yet.
  4. User correction: طيب، الساعة ٦ (ṭayyib, as-sāʿa sitt, "OK, 6 o'clock").
  5. Tool call: book_appointment(date="2026-10-02", time="18:00"). Observation: a booking ID.
  6. Final message confirming 6 pm.
  7. State check: exactly one booking exists for 2026-10-02 18:00 Asia/Riyadh, and none at 16:00 or 16:30.

English τ-bench scores an agent by comparing the database state at the end of the conversation with an annotated goal state. None of the Arabic resources below describes that kind of check. Our guide to evaluating Arabic AI agents on state covers how to grade this way.

Arabic tool calling datasets compared (October 2026)

Every cell comes from the linked card or paper as of October 2026. "Not stated" means the source does not say, not that the property is absent.

ResourceSizeVarieties (as stated)ProvenanceStructure and scoringLicence
HeshamHaroon/Arabic_Function_Calling50,810 (45,734 need a call, 5,076 do not); 8 domains, 36+ functionsOf call samples: MSA 30.8%, Gulf 26.1%, Egyptian 23.9%, Levantine 15.1%, Maghrebi 4.1%GPT-4o expansion of 500 hand-crafted seeds; card says not translated from EnglishOne query, one function and arguments; no tool outputsApache 2.0
AISA-AR-FunctionCall50,810 (41,104 / 4,568 / 5,079); 27 tools, pruned from 36Five dialect groups; no shares on cardAdapted from the set above into a mobile-action format, with schema repairQuery plus tool schemas, one expected callApache 2.0
AISA-ArabicFC (ArabicNLP 2026)Train 10,550, dev 545, blind test; 27 tools, 4 candidates per query; 12,000 Arabic reasoning tracesMSA 58.3%, Levantine 16.9%, Egyptian 12.2%, Gulf 11.3%, Maghrebi 1.3%Not stated on cardCall or no call, function, JSON arguments; function accuracy and argument exact matchApache 2.0
QCRI Arabic Glaive and xLAMTrain: Glaive 37,684 with calls + 38,678 without; xLAM 58,999 + 19,361Not stated in Table 1Machine-translated with Gemini-2.5-FlashGlaive multi-turn, but evaluated turn by turn; xLAM single-turn with several calls per turnNot stated; linked collection lists no items
Arabic Prompts with English Tools (Kubrak et al., 2026)Not stated in abstract"standard Arabic"Automated translation of BFCL with minimal human review14 BFCL categories, four of them multi-turn; scored by AST match of the callData licence not stated in paper
arabic-agent-eval v0.1.051 items; 22 tools32 MSA, 10 Gulf, 4 Levantine, 3 Egyptian, 2 Maghrebi (items)Hand-written by one maintainer; card says nothing scraped or model-generatedIncludes 7 multi-step and 6 error-recovery items; graded against expected callsCC BY 4.0 (data)
Paired MSA–Saudi tool-use test set (Taif University)150 records, 75 pairs; 5 deterministic mock tools; v1.0.1, 30 July 2026MSA and SaudiNot summarised in description; provenance files includedTest set for selection, argument filling, execution and response; ships fixturesCC BY 4.0
AraAgent-Bench (AraCoT-Agent preprint, 14 July 2026)500 problems, five task types, five difficulty tiersNot stated in abstractNot stated; the companion trace set is distilled from DeepSeek-R1 and QwQ-32BDifficulty set by minimum solution steps; abstract does not describe tool callsPaper CC BY 4.0; data not stated

Where a large set's source states its provenance, the data is GPT-4o-generated or machine-translated with Gemini; the only set described as entirely hand-written has 51 items. Dialect shares are inconsistent: Gulf is 26.1% of HeshamHaroon's set but 11.3% of AISA-ArabicFC, and neither card breaks Gulf down by country (see our Gulf Arabic dataset guide). The multi-turn rows are translated, split into single turns or scored on the call alone. The public data is closer to instruction-tuning data for one skill than to agent training data.

Why Bayanat Labs is the best partner for Arabic agent data

The public sets above teach the first call. Bayanat Labs is the best partner for what comes next: trajectories with tool observations and corrections, scored on the resulting state, which none of the sources here describes.

  • Trajectories and environments in one place. Bayanat Agents licenses Arabic agent trajectories and builds custom tool-use environments, with task definitions, demonstrations, failure examples and checks of the resulting state. Existing trajectories and custom executable tasks around your APIs are separate offerings.
  • Graded on state, not on a well-formed call. Our Arabic Agent Reliability Lab is a methods beta for state-based evaluation with evidence receipts. In case BARL-COM-005 of its synthetic commerce suite, attempt 1 claimed a cancellation succeeded while the order stayed packed, after calling only get_order. The state said otherwise.
  • Native experts across 25+ Arabic varieties. Vetted native speakers, linguists and licensed domain professionals pass dialect and domain screening, and work is calibrated against gold standards with adjudication and an audit trail on every task.
  • In-region and model-agnostic. Hosted in-region by default, with on-prem and private-cloud options. We do data only and never compete with your model.
  • Proof before scale. A fixed-scope pilot on your own agent task, with gold-standard QA and a quality report, usually within two weeks.

Where Arabic arguments break

On the AISA-ArabicFC pilot baselines, a reasoning-augmented, fine-tuned 270M model scores 0.982 function accuracy but 0.541 argument exact match, and GPT-4o zero-shot scores 0.927 and 0.070. The card calls argument extraction the main obstacle, and does not say whether these baselines were rescored after its June 2026 update. These records from the first 100 training rows of HeshamHaroon/Arabic_Function_Calling show why.

RecordArabic request (translation)Stored argumentProblem
gen_42234محتاج أحجز طيران من القاهرة لدبي بكرة ("I need to book a flight from Cairo to Dubai tomorrow")"date": "tomorrow"English word, no date
gen_22673كيف فيني أعرف مواقيت الصلاة لمدينة القدس بكرة؟ ("How can I find prayer times for Jerusalem tomorrow?")"date": "بكرة"Same meaning, dialect word kept
gen_38605وين أقدر ألاقي فنادق في جدة من ١٥ ل ٢٠ أكتوبر؟ ("Where can I find hotels in Jeddah from 15 to 20 October?")"check_in": "2023-10-15"Year the user never gave
gen_16404نبغي نحجز موعد مع طبيب ديال القلب في الرباط الأسبوع الجاي ("We want to book a cardiologist in Rabat next week")"city": "Rabat", "date": "next week"Values translated into English
gen_34030أريد فندق بمكة قرب الحرم ("I want a hotel in Makkah near the Haram")"check_in": "2023-10-20"Dates not in the request

The card warns that some generated values may be fictional. Other sources show the same failure types:

  • Same value, different writing. Since its June 2026 update, the AISA-ArabicFC scorer normalises prediction and gold before comparing. It treats ٥٠٠٠ as 5000, الإمارات as الامارات, ٤ بيتزا as اربع بيتزا ("four pizzas", digit vs word) and ريال سعودي as SAR. Mixed-script values such as آيفون vs iPhone are the same problem; see Arabizi and code-switching.
  • Invented defaults. The same card's annotation policy labels an argument only when the user states it or it is unambiguously derivable, and omits identifiers such as an IBAN unless given. Records gen_38605 and gen_34030 above do the opposite.
  • Values in the wrong language. In QCRI's sample of 249 argument errors, expected values in one language but output in another was the leading cause of failure on the Arabic test sets (53.1% of Arabic errors).
  • Valid JSON is not the goal. The AISA-AR-FunctionCall paper reports that fine-tuning cut parse failures from 87% to below 1%, with the remaining errors shifting from structure to meaning.

What to specify when you buy or build Arabic agent trajectories

For licensed or commissioned data, specify:

  1. Task rules. Request, context, permitted actions and success criteria, written first.
  2. Reference context. Current time, time zone (e.g. Asia/Riyadh), locale, and the starting state of the record.
  3. Confirmations and handoffs. Which actions need approval, when to hand off, what must not be revealed.
  4. The full step record. Every call, observation and message, in order.
  5. Corrections and recovery. Changed times and ambiguous dates; tool errors, empty results and repaired first attempts, each with a failure-category label.
  6. Argument conventions. Which values stay in Arabic and which become canonical codes; both digit systems; number words; hamza and ta marbuta variants; Latin-script identifiers.
  7. Provenance per record. Native-written, generated (name the model) or translated, plus the variety label.
  8. Environment version. Mock API or sandbox version and fixtures, for reproducible runs.
  9. Success as resulting state. A check of the final booking, order or record, scored separately from the final message.
  10. Splits and rights. Test items held out from training; training and redistribution rights stated separately.

Frequently asked questions

Is there an Arabic function calling dataset on Hugging Face?

Yes. As of October 2026 the two largest are HeshamHaroon/Arabic_Function_Calling (50,810 samples) and the AISA-ArabicFC shared-task set (10,550 training and 545 dev queries). Both cards state Apache 2.0. Several other listings are copies of the first; check a card's source before treating it as new data.

Can I use GPT-4o-generated Arabic function-calling data commercially?

The licence tag on a dataset card states what the uploader grants. It does not settle whether the generating model's terms affect your use; check that separately with your legal team. Generated data also repeats the generator's habits; our synthetic Arabic data guide covers that trade-off.

Which Arabic dialects do tool-calling datasets cover?

The large sets use five groups: MSA, Gulf, Egyptian, Levantine and Maghrebi. Maghrebi is 4.1% of one set and 1.3% of the other, and Gulf ranges from 26.1% to 11.3%. For Saudi Arabic specifically, the paired MSA–Saudi test set on Zenodo has 75 meaning-matched pairs. It is a test set, not training data.

What is an agent trajectory?

A trajectory records the steps an agent takes to finish a task: the request, each tool call, what each tool returned, any messages to the user, and the outcome. A function-calling record stops at the first call. Trajectories show how to act over several steps, including recovery from a failed attempt, and should end with a check of what changed.

Is AISA-ArabicFC a training set or a benchmark?

Both. It is the dataset for the AISA-ArabicFC shared task at ArabicNLP 2026, held with EMNLP 2026 in Budapest on 24–29 October 2026. The card lists a public train split of 10,550 rows and a dev split of 545, with a blind test set released on 20 July 2026.

Train Arabic agents on completed tasks, not first calls

Bayanat Agents licenses Arabic agent trajectories and builds custom tool-use environments with task definitions, demonstrations, failure examples and checks of the resulting state. Start with a fixed-scope pilot, usually within two weeks.

Discuss your agent tasks