Token counter
Paste a text to see how many tokens it uses and what it costs with each model family. Your text stays in your browser: nothing is sent to our servers.
| Model familyi | Tokens | As input | As output |
|---|---|---|---|
| Type or paste a text to see the results. | |||
The OpenAI count is exact (GPT-4o tokenizer). Other families are estimated from our measurements of their public tokenizers on the same texts in each language. Claude and Grok do not publish their tokenizers, so they use the average of the measured families.
The language tax: same text, different token counts
Tokens needed for the same four texts in each language, measured with each public tokenizer.
| Tokenizeri | English | Français | Español | العربية |
|---|---|---|---|---|
| Mistral (Nemo) | 220 | 280 +27% | 276 +25% | 277 +26% |
| OpenAI (GPT-4o) | 219 | 273 +25% | 256 +17% | 295 +35% |
| Google (Gemma 3) | 218 | 287 +32% | 256 +17% | 320 +47% |
| Meta (Llama 4) | 219 | 279 +27% | 261 +19% | 347 +58% |
| Alibaba (Qwen 3) | 219 | 325 +48% | 302 +38% | 360 +64% |
| DeepSeek (V3.1) | 219 | 327 +49% | 318 +45% | 368 +68% |
| Z.ai (GLM-4.5) | 219 | 304 +39% | 281 +28% | 409 +87% |
Last update: Sep 25, 2026
Why do other languages use more tokens than English?
Tokenizers are built mostly from English text, so common English words are usually a single token, while French and Spanish words, and Arabic words even more, are often split into several pieces. Arabic also attaches prefixes and suffixes to words, which multiplies the pieces. Tokenizers with a more multilingual vocabulary close part of the gap. Glossary →