When a model’s weights are public, anyone with the hardware can host it and sell access. That creates a real market: the same model is often available from half a dozen companies, at very different prices. We compared the hosts of the 57 rated models that at least three of them offer (a median of 6 hosts per model).
The price gap
For a typical model, the most expensive host charges 2.2× what the cheapest one does (blended price: three tokens read for each token written). For 56% of these models the gap is 2× or more. The widest gaps today:
| Model | Hosts | Cheapest | Most expensive | Gap |
|---|---|---|---|---|
| DeepSeek V3.2 | 13 | AtlasCloud · $0.15 | SambaNova · $3.38 | 22× |
| DeepSeek V4 Flash 0731 | 28 | StreamLake · $0.066 | Cloudflare · $0.66 | 10× |
| Llama 3.1 8B Instruct | 5 | DeepInfra · $0.025 | CoreWeave · $0.22 | 8.8× |
| Mistral Nemo | 6 | DekaLLM · $0.021 | Mistral · $0.15 | 7.1× |
| gpt-oss-120b | 20 | CoreWeave · $0.065 | Cerebras · $0.45 | 6.9× |
| Llama 3.3 70B Instruct | 10 | DeepInfra · $0.16 | Together · $1.04 | 6.7× |
With DeepSeek V3.2, the same request costs 22× as much at SambaNova as at AtlasCloud.
Why the prices differ
- Speed. Some hosts run models on chips built for speed and charge for it; others use slower, cheaper setups. If you need answers in real time, the premium can be worth it.
- Compression. 38% of the offers we see serve a compressed version of the model (for example 8-bit or 4-bit), which is cheaper to run and can be slightly less accurate.
- Limits and extras. The context length, the maximum answer, prompt caching and reliability vary from one host to another.
The maker is not always the cheapest
Among these models, 41 are also sold by the company that made them. For 22 of them, the maker charges more than the cheapest host. Buying from the maker can still make sense: you get the reference version, often the newest features first, and one bill with a company you may already use.
How to choose a host
- Open the model’s page on this site: the "Where to use it" table lists every host with its price, its precision (full or compressed) and its uptime over the last 24 hours.
- Check reliability: 92% of offers answered at least 95% of requests over the last 24 hours, but the rest can fail often enough to matter.
- Before switching for good, run a sample of your real requests on two hosts: compression and settings can change the answers.
The gap between hosts adds to the gap between models: see how much intelligence a dollar buys and the cost calculator.
The figures in this article are recalculated from our data at every update (last: Sep 30, 2026). How we work