AI API pricing explained: tokens, input, output and caching

How AI model pricing works, why output costs more than input, and how to estimate your monthly bill with a worked example.

Every AI model page on this site shows two prices, like "$3 / $15". Here is what they mean and how to turn them into a real monthly budget.

You pay per token

Models read and write text in tokens: words or pieces of words. In English, 1,000 tokens is roughly 750 words, or a page and a half. French typically needs 20 to 30% more tokens for the same text, and Arabic often more than that, because tokenizers were historically trained mostly on English. If your app works in French or Arabic, budget accordingly.

Input and output are priced separately

The first price is for input, the tokens you send: instructions, conversation history, documents. The second is for output, the tokens the model writes. Output is usually 3 to 8 times more expensive because each output token requires a full pass through the model. Prices are quoted per million tokens.

A worked example

Say you run a support assistant. Each request sends 1,500 tokens (instructions, the customer's message and some history) and gets a 300-token answer. You handle 2,000 requests a day.

  • Input per month: 1,500 × 2,000 × 30 = 90 million tokens
  • Output per month: 300 × 2,000 × 30 = 18 million tokens

With a model priced at $3 input / $15 output: 90 × $3 + 18 × $15 = $540 per month. With a model at $0.25 / $1.50: 90 × $0.25 + 18 × $1.50 = $49.50 per month. Same traffic, eleven times cheaper. The cost calculator does this sum for every model at once.

Hidden multipliers to watch

  • Conversation history. In a chat, every new message re-sends the whole conversation, so input grows with each turn.
  • Reasoning tokens. Reasoning models "think" before answering, and those thinking tokens are billed as output. A short visible answer can hide thousands of billed tokens.
  • Images and files. An image or a PDF page is converted into tokens too, often several hundred to a few thousand each.

Ways to pay less

  • Prompt caching: if every request starts with the same long instructions or document, cached input is billed at a fraction of the normal price (shown as "Cached input" on model pages when available).
  • Batch processing: most providers offer about 50% off for requests that can wait a few hours.
  • Route by difficulty: send easy requests to a small model and only hard ones to a flagship.
  • Keep prompts lean: trim history and avoid pasting whole documents when a relevant excerpt will do.

Prices change often, usually downwards. Our change log records every price change we detect.