Best AI models for Arabic
Arabic text uses more tokens than English, and the difference depends on the model. We measured it ourselves, then combined it with independent quality scores and current prices to show which models give the best results for your money in Arabic.
The cost of Arabic, by model family
| Tokenizeri | Tokens (English) | Tokens (Arabic) | Extra vs English |
|---|---|---|---|
| Mistral (Nemo) | 220 | 277 | +26% |
| OpenAI (GPT-4o) | 219 | 295 | +35% |
| Google (Gemma 3) | 218 | 320 | +47% |
| Meta (Llama 4) | 219 | 347 | +58% |
| Alibaba (Qwen 3) | 219 | 360 | +64% |
| DeepSeek (V3.1) | 219 | 368 | +68% |
| Z.ai (GLM-4.5) | 219 | 409 | +87% |
Anthropic (Claude) and xAI (Grok) do not publish their tokenizers, so they cannot be measured; for their models we use the average of the measured families (marked ≈). Google’s Gemini is estimated with Google’s open Gemma tokenizer.
Most capable models and their real cost in Arabic
Rankings →| Model | Overall indexi | Arabic factor | Inputi | 1,000 pages |
|---|---|---|---|---|
GPT-6 AstraOpenAI | 166.6 | ×1.35 | $10 | $4.56 |
Claude Fable 5.1Anthropic | 165.0 | ≈ ×1.55 | $10 | $5.25 |
Claude Fable 5Anthropic | 163.6 | ≈ ×1.55 | $10 | $5.25 |
Claude Opus 5Anthropic | 162.7 | ≈ ×1.55 | $5 | $2.62 |
GPT-5.5 ProOpenAI | 162.5 | ×1.35 | $30 | $13.69 |
GPT-5.6 SolOpenAI | 162.0 | ×1.35 | $2 | $0.912 |
GPT-5.6 TerraOpenAI | 159.3 | ×1.35 | $2 | $0.912 |
GPT-5.5OpenAI | 159.3 | ×1.35 | $5 | $2.28 |
GPT-5.4 ProOpenAI | 159.1 | ×1.35 | $30 | $13.69 |
Claude Opus 4.8Anthropic | 158.3 | ≈ ×1.55 | $5 | $2.62 |
Gemini 3.7 FlashGoogle DeepMind | 157.7 | ×1.47 | $0.75 | $0.371 |
Kimi K3Moonshot AI | 157.7 | ≈ ×1.55 | $3 | $1.57 |
Gemini 3.8 FlashGoogle DeepMind | 157.1 | ×1.47 | $0.75 | $0.371 |
Muse Spark 1.3Meta | 156.9 | ×1.58 | $1.25 | $0.671 |
GPT-5.4OpenAI | 156.9 | ×1.35 | $2.50 | $1.14 |
Cost to read 1,000 pages of text (about 300,000 words in English) written in Arabic, as input. Each model gets the factor measured on its company’s latest public tokenizer, so treat it as an estimate: newer models may use a different tokenizer. ≈ marks companies that publish no tokenizer (average of measured families).
Best value for Arabic
Strong models (overall index 140 or more), cheapest first for processing Arabic text.
- $0.01 / 1,000 pages
- $0.017 / 1,000 pages
- $0.036 / 1,000 pages
- $0.038 / 1,000 pages
- $0.045 / 1,000 pages
- $0.084 / 1,000 pages
- $0.09 / 1,000 pages
- $0.091 / 1,000 pages
Models built for Arabic
Specialised models are not always available through the big APIs above, but they are worth testing, especially for local dialects and cultural knowledge.
- Jais ↗Arabic–English models from the UAE, published with open weights.
- ALLaM ↗Arabic-focused models from Saudi Arabia’s data and AI authority.
- Fanar ↗Arabic-centric models and assistant from Qatar.
- Falcon Arabic ↗Arabic versions of the Falcon open models from Abu Dhabi’s Technology Innovation Institute.
- SILMA ↗Small open Arabic models designed to run at low cost.
Independent Arabic leaderboards
These projects test models directly in Arabic. Use them alongside the overall scores above.
- Open Arabic LLM Leaderboard (OALL) ↗Benchmarks of open models on Arabic knowledge and reasoning tests.
- HELM Arabic ↗Transparent, reproducible evaluation of open and commercial models in Arabic.
- Arabic Broad Leaderboard ↗Broad Arabic benchmark covering many task types, with speed measurements.
Tips for using AI in Arabic
- Test with your own dialect. Most models are strongest in Modern Standard Arabic; results in Egyptian, Gulf, Levantine or Maghrebi dialects vary much more from one model to another.
- Ask for Arabic explicitly. Add "answer in Arabic" to your instructions, especially when your question contains English terms, otherwise some models switch language.
- Check right-to-left display. Mixed Arabic and Latin text (product names, code, numbers) can appear in the wrong order if your app does not handle direction properly.
- Budget for extra tokens. Depending on the model, the same content costs 26% to 87% more in Arabic than in English: use the table above or the token counter to estimate.
- Review names, numbers and diacritics. Machine-generated Arabic can be fluent yet get proper names, dates or vowel marks wrong, so proofread anything you publish.
How we measured
We wrote the same four texts (a news item, a customer email, a technical explanation and travel tips) in 15 languages, then counted the tokens that each public tokenizer produces. Quality scores come from Epoch AI and are measured mostly in English, so always test the finalists on your own Arabic content.
Get a personal recommendationCount the tokens of your own text