AI API price vs measured latency — the same model, different hosts
The same model can cost 11.0× more depending on who hosts it. Google: Gemma 4 31B is $0.090 per million input tokens at deepinfra and $0.990 at cerebras — identical weights, 11.0× the price. Below, every multi-hosted model we track, with what each host charges and how fast we measured it answering.
Which models cost the most and least depending on the host?
| Model | Hosts | Cheapest | Dearest | Spread | Fastest measured |
|---|---|---|---|---|---|
| Google: Gemma 4 31B | 7 | deepinfra $0.090 | cerebras $0.990 | 11.0× | sambanova 20 ms |
| DeepSeek: DeepSeek V4 Flash 0731 | 8 | deepinfra $0.060 | deepseek $0.440 | 7.3× | fireworks 19 ms |
| Z.ai: GLM 5.2 | 9 | deepinfra $0.487 | glm $1.400 | 2.9× | fireworks 19 ms |
| MiniMax: MiniMax M2.7 | 5 | deepinfra $0.250 | sambanova $0.600 | 2.4× | sambanova 20 ms |
| MiniMax: MiniMax M3 | 5 | deepinfra $0.280 | sambanova $0.600 | 2.1× | sambanova 20 ms |
| Google: Gemma 4 26B A4B | 4 | deepinfra $0.070 | google $0.150 | 2.1× | google 44 ms |
| DeepSeek: DeepSeek V4 Flash Vision Exp | 5 | deepinfra $0.216 | deepseek $0.440 | 2.0× | fireworks 19 ms |
| Z.ai: GLM 5.3 Flash | 9 | novita $0.075 | baseten $0.150 | 2.0× | fireworks 19 ms |
| Qwen: Qwen3.5-9B | 3 | siliconflow $0.100 | together $0.170 | 1.7× | siliconflow 79 ms |
| DeepSeek: DeepSeek V4 Flash 0423 | 3 | deepinfra $0.090 | novita $0.140 | 1.6× | novita 76 ms |
| Qwen: Qwen3.5-122B-A10B | 3 | siliconflow $0.260 | novita $0.400 | 1.5× | novita 76 ms |
| DeepSeek: DeepSeek V4 Pro 0423 | 5 | fireworks $1.200 | baseten $1.740 | 1.4× | fireworks 19 ms |
| MoonshotAI: Kimi K2.7 Code | 6 | deepinfra $0.680 | together $0.950 | 1.4× | fireworks 19 ms |
| DeepSeek: DeepSeek V4 Pro 0813 | 7 | novita $0.990 | deepseek $1.320 | 1.3× | fireworks 19 ms |
| Z.ai: GLM 5.1 | 6 | deepinfra $1.050 | glm $1.400 | 1.3× | novita 76 ms |
| MoonshotAI: Kimi K2.6 | 5 | deepinfra $0.750 | fireworks $0.950 | 1.3× | fireworks 19 ms |
| StepFun: Step 3.7 Flash | 3 | deepinfra $0.160 | stepfun $0.200 | 1.2× | novita 76 ms |
| Qwen: Qwen3.5-27B | 3 | siliconflow $0.250 | novita $0.300 | 1.2× | novita 76 ms |
| Z.ai: GLM 5.3 | 9 | reka $1.170 | glm $1.400 | 1.2× | fireworks 19 ms |
| Meta: Muse Glimmer 30B | 3 | deepinfra $0.300 | together $0.350 | 1.2× | fireworks 19 ms |
| Thinking Machines: Inkling Small | 3 | deepinfra $0.450 | together $0.500 | 1.1× | baseten 69 ms |
| Thinking Machines: Inkling | 3 | deepinfra $0.950 | together $1.000 | 1.1× | baseten 69 ms |
| MoonshotAI: Kimi K3 | 4 | deepinfra $2.850 | baseten $3.000 | 1.1× | fireworks 19 ms |
| Qwen: Qwen3.8 2.4T A95B | 4 | novita $2.000 | together $2.000 | 1.0× | novita 76 ms |
Why compare hosts only within one model?
Because across models a price comparison means nothing. A cheap small model and an expensive large one are not competing for the same job, and a table that ranks them together ranks nothing. Held to a single model the comparison becomes exact: identical weights, identical task, and the only variables left are what the host charges and how quickly it answers.
Where do these two numbers come from?
They come from different places, and we keep them apart. Latency and uptime are measured by our own probes in 4 regions, every five minutes, against each host’s real API. Prices are not measured — providers publish them, and we read them from OpenRouter’s public catalogue. We do not blend the two into a single “value score”: that number would be our invention rather than either provider’s fact, and the weighting inside it would quietly decide the ranking. The two columns sit side by side and the trade-off stays yours.
Method: how latency is measured · Prices as published 2026-09-08 via OpenRouter · data & licence