Model mix simulator
Most requests are easy. Send them to a cheap model, keep a top model for the hard ones, and pay a fraction of the bill.
—per month
The mix—per month
You save—
| Model | Overall indexi | The mix | You save |
|---|
Rated models that cost at most a third of the top model per request, best overall index first. The saving uses the share of easy requests above. per 1M tokens · USD, 30 days per month.
How to split the work
Most apps send many simple requests and a few hard ones. Easy: sorting a message, pulling fields out of a document, answering from a FAQ, writing a short reply. Hard: multi-step reasoning, code changes, long analysis, anything where a mistake is costly.
Three ways to route requests
- By task: your app already knows what it is doing (a support ticket or a contract review) and picks the model.
- With a sorter: a cheap model reads each request first and labels it easy or hard.
- By escalation: the everyday model answers first; when it is unsure or a check fails, the request goes to the top model.
Before switching, test the everyday model on a sample of your real requests: the overall index measures general ability, not your task.