AI prices span an enormous range. Among the 136 paid models with an independent score, the cheapest costs $0.022 per million tokens and the most expensive $67.50: 3,103× more. Their scores are far closer together. So what does the extra money actually buy?

Two notes on the numbers. Prices are "blended": they assume three tokens read for every token written, a typical ratio for chat and document work. Scores are the overall index from Epoch AI’s independent tests, a single scale where higher means more capable.

The value frontier

Sort every model by price and keep only those that score higher than every cheaper model: that is the value frontier. Each model on it is the most capable you can get at its price; every other model is beaten by something cheaper. Today, 9 models make the list.

ModelCompanyOverall indexBlended price
Mistral NemoMistral AI118.6$0.022
gpt-oss-20bOpenAI137.8$0.036
Qwen3.7 FlashAlibaba Qwen144.6$0.055
DeepSeek V4 Flash 0731DeepSeek154.5$0.094
GPT-5.6 LunaOpenAI156.3$0.45
Gemini 3.7 FlashGoogle DeepMind157.7$1.50
GPT-5.6 SolOpenAI162.0$4
Claude Opus 5Anthropic162.7$10
GPT-6 AstraOpenAI166.6$20

Diminishing returns

The first dollars buy a lot. For $0.25 or less per million tokens, the best model is DeepSeek V4 Flash 0731 (154.5); under $1, GPT-5.6 Luna (156.3); under $5, GPT-5.6 Sol (162.0). The top model, GPT-6 Astra, scores 166.6 for $20.

In other words, the last 10.3 points of the index cost 44× as much as the best model under $1. For most everyday work that gap is hard to notice; on the hardest problems, it can be the difference between a right and a wrong answer.

A higher price is no guarantee either: the most expensive model we track, GPT-5.4 Pro ($67.50), scores 159.1, below GPT-5.6 Sol, which costs 17× less.

When paying more is worth it

  • Hard reasoning, advanced math and complex code, where the best models are clearly ahead.
  • Low volumes: at a few hundred requests a day, even a premium model costs little.
  • Costly mistakes: legal, medical or financial drafts, where one wrong answer outweighs any saving.

When it is not

  • High-volume, simple tasks: sorting, extraction, short answers, first drafts that a person checks anyway.
  • Mixed workloads: send the easy requests to a model from the cheap end of the frontier and keep a top model for the rest. The model mix simulator shows the saving.

What the index does not tell you is how a model does on your own task, in your language, with your data. Use it to build a shortlist, then test the finalists. The rankings break the scores down by domain: coding, math, science and facts.

The figures in this article are recalculated from our data at every update (last: Sep 30, 2026). How we work