The context window is how much text a model can take in at once: your instructions, the conversation so far and any documents, all counted in tokens. It decides whether you can hand a model a whole contract, a codebase or a year of support tickets, or whether you have to cut them into pieces first. Here is where the 328 text models we track stand.
Most models now read a lot
The median context window is 256K tokens, roughly 492 printed pages. 90% of the models read at least 128K tokens, and 32% read a million or more, the length of several long novels.
| Context window | Models | Share |
|---|---|---|
| ≤ 32K | 19 | 6% |
| 32K–128K | 74 | 23% |
| 128K–512K | 130 | 40% |
| ≥ 1M | 105 | 32% |
The biggest windows
| Model | Company | Context window | About |
|---|---|---|---|
| Grok 4.20 Multi-Agent | xAI | 2M | 3,750 pages |
| Grok 4.20 | xAI | 2M | 3,750 pages |
| Llama 4 Scout | Meta | 1.3M | 2,458 pages |
| GPT-6.1 Sol Pro | OpenAI | 1.1M | 1,969 pages |
| GPT-6.1 Sol | OpenAI | 1.1M | 1,969 pages |
Reading is not free
A bigger window lets you send more, and every token you send is billed. Filling the whole window of GPT-6 Astra, the highest-rated model we track, costs about $10.50 in input alone, for a single request. With DeepSeek V4 Flash 0731, the cheapest model we list that reads a million tokens, the same full window costs $0.019. In a chat, where the history is sent again with every message, those amounts come back at every turn.
Bigger is not always better
Research has repeatedly found that models make better use of what sits at the start and at the end of a long input than of what is buried in the middle. For many tasks, a good search step that sends the right ten pages beats pasting in a thousand: it is cheaper, faster and often more accurate. Long windows earn their keep when the whole text really matters at once, such as checking a long contract for contradictions, following a large codebase or summarising a book.
How much do you need?
- Chat and customer support: 32K to 128K tokens is plenty, as long as old messages are summarised.
- Questions on documents: 128K handles a few long documents at once; beyond that, send only the relevant passages.
- Whole codebases, books or case files: look at the models that read a million tokens or more, and check what a full window costs.
The best long-context models page ranks them with their prices, and the glossary explains the term in two lines.
The figures in this article are recalculated from our data at every update (last: Sep 30, 2026). How we work