AI prices span an enormous range. Among the 136 paid models with an independent score, the cheapest costs $0.022 per million tokens and the most expensive $67.50: 3,103× more. Their scores are far closer together. So what does the extra money actually buy?
Two notes on the numbers. Prices are "blended": they assume three tokens read for every token written, a typical ratio for chat and document work. Scores are the overall index from Epoch AI’s independent tests, a single scale where higher means more capable.
The value frontier
Sort every model by price and keep only those that score higher than every cheaper model: that is the value frontier. Each model on it is the most capable you can get at its price; every other model is beaten by something cheaper. Today, 9 models make the list.
| Model | Company | Overall index | Blended price |
|---|---|---|---|
| Mistral Nemo | Mistral AI | 118.6 | $0.022 |
| gpt-oss-20b | OpenAI | 137.8 | $0.036 |
| Qwen3.7 Flash | Alibaba Qwen | 144.6 | $0.055 |
| DeepSeek V4 Flash 0731 | DeepSeek | 154.5 | $0.094 |
| GPT-5.6 Luna | OpenAI | 156.3 | $0.45 |
| Gemini 3.7 Flash | Google DeepMind | 157.7 | $1.50 |
| GPT-5.6 Sol | OpenAI | 162.0 | $4 |
| Claude Opus 5 | Anthropic | 162.7 | $10 |
| GPT-6 Astra | OpenAI | 166.6 | $20 |
Diminishing returns
The first dollars buy a lot. For $0.25 or less per million tokens, the best model is DeepSeek V4 Flash 0731 (154.5); under $1, GPT-5.6 Luna (156.3); under $5, GPT-5.6 Sol (162.0). The top model, GPT-6 Astra, scores 166.6 for $20.
In other words, the last 10.3 points of the index cost 44× as much as the best model under $1. For most everyday work that gap is hard to notice; on the hardest problems, it can be the difference between a right and a wrong answer.
A higher price is no guarantee either: the most expensive model we track, GPT-5.4 Pro ($67.50), scores 159.1, below GPT-5.6 Sol, which costs 17× less.
When paying more is worth it
- Hard reasoning, advanced math and complex code, where the best models are clearly ahead.
- Low volumes: at a few hundred requests a day, even a premium model costs little.
- Costly mistakes: legal, medical or financial drafts, where one wrong answer outweighs any saving.
When it is not
- High-volume, simple tasks: sorting, extraction, short answers, first drafts that a person checks anyway.
- Mixed workloads: send the easy requests to a model from the cheap end of the frontier and keep a top model for the rest. The model mix simulator shows the saving.
What the index does not tell you is how a model does on your own task, in your language, with your data. Use it to build a shortlist, then test the finalists. The rankings break the scores down by domain: coding, math, science and facts.
The figures in this article are recalculated from our data at every update (last: Sep 30, 2026). How we work