"How much will it cost?" is the first question before putting an AI model in front of customers. The honest answer: it depends on the model far more than on anything else. Using the live prices on this site, we priced three common workloads with three models:
- GPT-6 Astra (OpenAI), the highest-rated model we track (overall index 166.6);
- GPT-5.6 Sol (OpenAI), the best-rated model costing a fifth of that or less (162.0);
- DeepSeek V4 Flash 0731 (DeepSeek), the cheapest model that still scores above the median of all rated models (154.5, against a median of 145.8).
The three workloads
- Customer support bot: 1,000 conversations a day of about four messages. Each message resends the instructions and the conversation so far, so one conversation reads about 6,000 tokens and writes about 800.
- Questions on company documents: 300 questions a day, each answered from about 12,000 tokens of retrieved passages, with an answer of 400 tokens.
- Coding assistant for a team of about ten developers: 1,500 requests a day, 8,000 tokens of code in and 1,000 out.
The monthly bill
| Workload | GPT-6 Astra (166.6) | GPT-5.6 Sol (162.0) | DeepSeek V4 Flash 0731 (154.5) |
|---|---|---|---|
| Customer support bot | $3,000 | $600 | $10.92 |
| Questions on company documents | $1,260 | $252 | $3.10 |
| Coding assistant (team of ~10) | $5,850 | $1,170 | $20.88 |
Monthly cost in US dollars (30 days), with the overall index of each model in brackets.
The same support bot costs $3,000 a month with GPT-6 Astra, $600 with GPT-5.6 Sol and $10.92 with DeepSeek V4 Flash 0731: 275× between the first and the last. Prices per token differ that much, and a busy product multiplies every difference by millions of tokens.
Where the money goes
Two things drive the bill. The first is what the model reads. Chat apps resend the instructions and the whole conversation with every message, so a long conversation costs far more than its last message suggests, and document tools read thousands of tokens of context for every question. The second is what the model writes: with GPT-6 Astra, an output token costs 5× as much as an input token, and reasoning models also write hidden "thinking" tokens that are billed as output (see what the thinking costs).
Three ways to cut it
- Prompt caching. If 50% of what each support conversation reads comes from the cache, the GPT-6 Astra bill falls from $3,000 to $2,190. Instructions, examples and documents that repeat are the easy candidates.
- A model mix. Send the easiest 70% of conversations to DeepSeek V4 Flash 0731 and keep GPT-6 Astra for the rest: $908 a month instead of $3,000. Try your own split in the model mix simulator.
- Less context. Summarise long conversations instead of resending them in full, and retrieve fewer, better passages from your documents.
So which model should you pick?
Price is only half of the question. The cheaper models score lower in independent tests (154.5 against 166.6 on the overall index), and whether that gap matters depends entirely on your task. Run 30 of your real conversations through two or three candidates before deciding, and use the cost calculator to price your own volumes for every model.
The prices behind these figures update several times a day, and this page is recalculated with them.
The figures in this article are recalculated from our data at every update (last: Sep 30, 2026). How we work