Every price on this site is given per million tokens, and every model’s memory limit is counted in tokens too. Yet the word is rarely explained. A token is a piece of text: a whole word when the word is common, part of a word when it is rare, a space or a punctuation mark. Models never see letters or words, only these pieces, which a tool called a tokenizer cuts out before the text reaches the model. A common word like "house" is usually a single token; a long or rare word is split into two, three or more pieces.

How many words is a token?

We counted it on the same short texts in 15 languages, with the tokenizer of OpenAI’s GPT-4o, one of the most widely used. In English, 100 words took 113 tokens. So 1,000 tokens hold about 890 words, and a token is 5.2 characters long on average, spaces included. The rule of thumb you often read, 100 tokens for 75 words, is older and more cautious: current tokenizers do a little better on everyday English.

Other languages need more tokens

LanguageSame text vs EnglishTokens per 100 wordsCharacters per token
English—1135.2
Chinese+11%—1.4
Portuguese+15%1274.8
Spanish+17%1224.9
German+17%1325.1
Russian+19%1604.4
French+25%1374.8
Indonesian+31%1624.4
Arabic+35%1883.2
Turkish+35%1993.7
Korean+39%2191.8
Vietnamese+45%1213.6
Italian+46%1604.1
Hindi+57%1553.3
Japanese+61%—1.4

Tokenizers learn from their training text, which is mostly English, so English words are more often kept whole and other languages get cut into more pieces: Arabic takes 188 tokens per 100 words, Korean 219. Words are not the whole story, though, because some languages say the same thing in fewer, longer words. What counts for the bill is the whole message, and that is the first column: the same text needs 61% more tokens in Japanese, but only 11% more in Chinese. Chinese and Japanese do not put spaces between words, so for them only characters are meaningful: about 1.4 characters per token in Chinese.

Different models, different tokenizers

Each model family has its own tokenizer, and we measured 7 of them. For English they are almost identical: between 218 and 220 tokens for the same text. For other languages they can differ widely. The extreme case is Hindi: Gemma needs 284 tokens for our text and GLM 1,123, 4× as many. If you work in a language other than English, the price per token is only half the comparison: the best AI for your language tool and the language tax analysis do the rest.

What it means for your bill

A few everyday examples in English, with the same tokenizer:

  • a 200-word email: about 230 tokens;
  • a 10-page report of 5,000 words: about 5,600 tokens;
  • a 90,000-word novel: about 100,000 tokens. At $1 per million tokens, having a model read it once costs about $0.10.

Three things make real bills bigger than these numbers suggest. What the model writes is billed too, usually at several times the price of what it reads. Chat apps resend the whole conversation with each new message. And reasoning models write hidden "thinking" tokens, billed as output: see what the thinking costs. The cost calculator does the sums for you, and the glossary keeps the short definition.

How we measured

We ran 4 short everyday texts, such as a news item and an email, translated into each language, through the public tokenizers of 7 model families, on September 25, 2026. It is a small sample, about 194 words in English, so treat the figures as a good indication rather than an exact rate: longer, more technical or more casual texts can give somewhat different numbers. Words are counted by the spaces between them.

The figures in this article are recalculated from our data at every update (last: Sep 25, 2026). How we work