Gulf customers abandon most bots because the bot misunderstands them, answers generically, or blocks the way to a person. With an AI agent, Arabic dialects make all three worse: customers write in Saudi or Emirati dialect, mix in English and Arabizi, and send voice notes, while most language models perform best in English and Modern Standard Arabic.
This guide is for business owners in Saudi Arabia and the UAE evaluating an AI agent for WhatsApp. It explains what the research actually says about dialect performance, where conversations break in practice, and gives you a test plan to run before you sign anything. For how we approach this in our own builds, see the NourSky AI Employee.
Why do Gulf customers abandon most bots?
The global numbers are sobering before dialect even enters the picture. In a Gartner survey of 3,566 B2B and B2C customers (February–March 2026):
- 49% said they would have been willing to use a chatbot if the company provided one, but only 7% actually used a chatbot or digital assistant in their most recent service interaction.
- Only 27% would try a chatbot again after a negative experience.
- 87% said access to a human agent is essential when companies use GenAI for customer service.
Gartner lists the reasons customers give up: the bot misunderstands the issue, gives generic responses, or makes it hard to reach a person. Eric Keller, Sr Director Analyst at Gartner, drew the design conclusion:
“Service leaders should not use GenAI as a mandatory first step for every issue.”
Now add Arabic. Every one of those failure modes (misunderstanding, generic answers, no way out) gets more likely when the customer writes the way people in Riyadh, Jeddah or Dubai actually write on WhatsApp.
What makes Arabic dialects hard for AI agents?
Arabic is not one written language in practice. Formal writing uses Modern Standard Arabic (MSA). Everyday chat uses dialect, and on WhatsApp that means informal spelling, dialect vocabulary, English words in Latin script, Arabizi (Arabic written in Latin letters and numbers), and voice notes.
Research on large language models confirms the gap. The AraDiCE benchmark, published at COLING 2025 by researchers including Basel Mousi, Nadir Durrani and Fatema Ahmad, states in its abstract:
“Arabic, with its rich diversity of dialects, remains significantly underrepresented in Large Language Models, particularly in dialectal variations.”
The authors found that Arabic-specific models outperformed multilingual ones on dialect tasks, but that “significant challenges persist in dialect identification, generation, and translation”. Their cultural-awareness benchmark specifically covered the Gulf, Egypt and the Levant.
A second benchmark, DialectalArabicMMLU (submitted October 2025, authors from IBM Research, NYU Abu Dhabi and MBZUAI including Nizar Habash), translated 3,000 questions into each of five dialects, including Saudi and Emirati, and tested 19 open-weight models from 1B to 13B parameters. Average accuracy was 62.8% in English and 51.9% in MSA, and the dialect averages fell below both.
Two honest caveats. First, these are academic benchmarks of knowledge questions, not customer-service conversations. Second, they tested mostly smaller open models; large commercial models may do better. But the direction is consistent: the further the input moves from English and formal Arabic, the more room there is for error, and customer service is exactly where input is least formal.
Where do AI agents break in a real Gulf WhatsApp conversation?
| Message type | Example | What typically breaks | What good design does |
|---|---|---|---|
| Dialect vocabulary | «ابي أحجز موعد بكرة» | Intent matching trained on formal phrasing | Tested on real customer phrasing, not rewritten FAQs |
| Mixed Arabic and English | «كم السعر للـ package الكاملة؟» | Product names and prices lost in the switch | Product catalogue includes the English and Arabic names customers use |
| Arabizi | “abi mow3ed bukra” | Treated as English or as noise | Either handled or recognised and handed to a person |
| Very short replies | «اوكي» · «تمام» · «لا» | Loses the thread of the conversation | Keeps context of the previous question |
| Voice notes | A 20-second voice message | Ignored, or an irrelevant text reply | Transcribed if supported, otherwise an honest reply and a handoff |
| Local references | Neighbourhood names, local holidays | Unknown places, wrong dates | Service areas and calendar in the agent’s data |
Notice that the right-hand column is mostly about data and rules, not about the model. A strong model with no local data still fails; a model with honest limits and a clean handoff fails gracefully.
How do you test an AI agent’s Arabic before you buy?
Don’t accept a demo written by the vendor. Run this test on a trial or sandbox:
- Export 50 real customer messages from your own WhatsApp, with personal details removed. These are your test set.
- Include every style your customers use: dialect, MSA, English, mixed, Arabizi, one-word replies.
- Send each message cold, as a first message, and record whether the intent was understood.
- Run 10 full conversations, not single messages. Many failures appear on the third or fourth turn, when context matters.
- Ask something the business doesn’t offer. The correct result is an honest “we don’t offer that”, not an invented answer.
- Ask for a person in dialect, for example «ابي أكلم موظف». It must work on the first try.
- Send a voice note and see what happens.
- Check the CRM record after each conversation: are the name, need and stage captured correctly?
- Have a native speaker from your market read the replies. Correct is not enough; they should sound natural for your brand.
- Score it: understood, answered correctly, handed off correctly. Decide your pass mark before you start.
If a vendor won’t let you run your own messages through the system before you commit, treat that as information.
Should an AI agent reply in dialect or in Modern Standard Arabic?
There is no single right answer; it is a brand decision with trade-offs.
| Approach | Strength | Risk |
|---|---|---|
| Formal MSA | Safe, clear, suits clinics, legal, finance | Can feel stiff on WhatsApp |
| Neutral “white” Gulf Arabic | Friendly and widely understood across the Gulf | Needs careful review to avoid sounding forced |
| Mirroring the customer’s dialect | Feels most natural when it works | Mistakes in dialect sound worse than correct MSA |
Many businesses settle on the middle option: understand every style the customer uses, and reply in a neutral, friendly register that fits the brand. Understanding the dialect matters more than imitating it.
How does NourSky approach Arabic dialects?
We will not promise that any AI agent, ours included, understands every dialect and every spelling. What NourSky does is make the limits visible and safe:
- Test on your real messages before launch, using the plan above.
- Build the agent’s data from the words your customers actually use, including local names for products, areas and services.
- Set handoff rules for anything the agent doesn’t understand with confidence, so the customer reaches a person instead of a loop.
- Review live conversations after launch and refine where understanding broke.
To understand the concept end to end, read What Is an AI Employee? A Practical Definition for Business Owners, and to estimate what slow or failed replies cost you, try the lost-leads calculator.
Test an AI agent on your customers’ real Arabic
Bring real WhatsApp messages from your customers to a live demo. We run them through the AI employee in front of you, including the ones you expect to fail, and show you what happens when it doesn’t understand.
FAQ
Can AI agents understand Saudi and Gulf Arabic dialects?
Often, but not reliably across every phrasing. Research such as AraDiCE (COLING 2025) and DialectalArabicMMLU (2025) shows language models still perform worse on dialects than on MSA and English. Test any agent on your own customers’ messages before you rely on it.
Why do customers give up on chatbots?
According to Gartner’s 2026 research, customers abandon chatbots when they misunderstand the issue, give generic responses, or make it hard to reach a person. Only 27% would try a chatbot again after a negative experience.
What should an AI agent do when it doesn’t understand an Arabic message?
Say so honestly and hand the conversation to a person, with the conversation history attached, rather than guessing or looping the customer through the same question.
Sources
- Gartner: Only 27% of Customers Would Try a Chatbot Again After a Negative Experience (2026)
- Gartner: 87% of Customers Say Companies Using GenAI Must Provide Access to a Human Agent (2026)
- Mousi et al. — AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs, COLING 2025
- Altakrori et al. — DialectalArabicMMLU: Benchmarking Dialectal Capabilities (arXiv, 2025)