How to cut your AI API bill: nine things that work

Most AI bills are bigger than they need to be. Nine practical ways to pay less, from caching and batch discounts to sending the easy requests to a cheaper model.

Most AI bills are bigger than they need to be, and the fixes are rarely clever. They come down to a few habits: sending too much text, using a top model for easy work, and paying full price for things that repeat. Here are nine ways to cut the bill, roughly in the order I would try them.

1. Measure before you cut

Log the tokens in and out of every request for a few days, with the feature that sent it. Usually one or two features account for most of the spend, and output tokens, which often cost several times more than input, matter more than expected. The cost calculator turns those numbers into a monthly bill for every model.

2. Send the easy requests to a cheaper model

Sorting messages, pulling fields out of a document or answering from a FAQ rarely needs a flagship model. Send those to a cheap model and keep the expensive one for the hard requests. With typical price gaps, moving 70% of the traffic can cut the bill by more than half; the model mix simulator shows it with today’s prices.

3. Cache what repeats

Long instructions, examples and reference documents are often identical from one request to the next. The major providers bill cached input at a fraction of the normal price, often about a tenth, and some charge a little extra the first time a prompt is stored. Put the stable part of the prompt first and the part that changes last, so the cache can match it.

4. Use the batch discount for work that can wait

Reports, tagging an archive, overnight summaries: when the answer can arrive within a day, the batch interfaces of OpenAI, Anthropic and Google cost about half the normal price. Check each provider’s current terms before you rely on it.

5. Send less history

Chat apps resend the whole conversation with every message, so the tenth message pays for the nine before it. Keep the last few turns word for word and replace the older ones with a short summary. For questions about documents, send fewer and better passages rather than everything that looks related.

6. Cap the length of answers

Set a maximum output length and ask for the format you actually need: a yes or no, a JSON object, three bullet points. Output is the expensive side of the bill, and models are happy to write more than you asked for.

7. Keep the thinking in check

Reasoning models write hidden "thinking" tokens that are billed as output, sometimes thousands per answer. Turn thinking down or off for simple requests, and set a thinking budget where the provider offers one. The analysis of what the thinking costs puts numbers on it.

8. Pick a model that suits your language

The same text can need very different numbers of tokens depending on the model, especially outside English. For some languages the gap between model families is bigger than the gap between their prices. Look up yours in the best AI for your language tool.

9. Look again every few months

Prices change often, usually downwards, and last year’s premium model may be cheap today. Follow the change log or the RSS feeds, and re-run your numbers when something moves.

None of this needs a new architecture. The first three steps are usually quick to set up, and they tend to make the biggest difference.