AI API price vs measured latency — the same model, different hosts
The same model can cost 11.6× more depending on who hosts it. DeepSeek: DeepSeek V3.2 is $0.259 per million input tokens at siliconflow and $3.000 at sambanova — identical weights, 11.6× the price. Below, every multi-hosted model we track, with what each host charges and how fast we measured it answering.
Which models cost the most and least depending on the host?
| Model | Hosts | Cheapest | Dearest | Spread | Fastest measured |
|---|---|---|---|---|---|
| DeepSeek: DeepSeek V3.2 | 6 | siliconflow $0.259 | sambanova $3.000 | 11.6× | sambanova 23 ms |
| Google: Gemma 4 31B | 7 | deepinfra $0.120 | cerebras $0.990 | 8.2× | sambanova 23 ms |
| Z.ai: GLM 4.7 | 5 | deepinfra $0.400 | cerebras $2.250 | 5.6× | google 40 ms |
| DeepSeek: DeepSeek V4 Pro | 7 | deepseek $0.435 | fireworks $1.740 | 4.0× | fireworks 17 ms |
| MiniMax: MiniMax M2.7 | 7 | deepinfra $0.250 | groq $0.600 | 2.4× | fireworks 17 ms |
| Google: Gemma 4 26B A4B | 4 | deepinfra $0.070 | google $0.150 | 2.1× | google 40 ms |
| Z.ai: GLM 5.2 | 8 | novita $0.724 | baseten $1.400 | 1.9× | fireworks 17 ms |
| Qwen: Qwen3 VL 30B A3B Instruct | 3 | deepinfra $0.150 | siliconflow $0.290 | 1.9× | novita 71 ms |
| Qwen: Qwen3.5-9B | 3 | siliconflow $0.100 | together $0.170 | 1.7× | siliconflow 81 ms |
| Z.ai: GLM 5 | 4 | deepinfra $0.600 | glm $1.000 | 1.7× | novita 71 ms |
| MoonshotAI: Kimi K2.6 | 6 | deepinfra $0.750 | together $1.200 | 1.6× | fireworks 17 ms |
| DeepSeek: DeepSeek V4 Flash | 5 | deepinfra $0.090 | deepseek $0.140 | 1.6× | fireworks 17 ms |
| Qwen: Qwen3.5-122B-A10B | 3 | siliconflow $0.260 | novita $0.400 | 1.5× | novita 71 ms |
| Z.ai: GLM 5.1 | 7 | deepinfra $1.050 | glm $1.400 | 1.3× | fireworks 17 ms |
| MoonshotAI: Kimi K2.7 Code | 5 | deepinfra $0.740 | fireworks $0.950 | 1.3× | fireworks 17 ms |
| MoonshotAI: Kimi K2.5 | 3 | deepinfra $0.450 | novita $0.570 | 1.3× | novita 71 ms |
| Z.ai: GLM 4.6 | 3 | deepinfra $0.500 | glm $0.600 | 1.2× | novita 71 ms |
| Qwen: Qwen3.5-27B | 3 | siliconflow $0.250 | novita $0.300 | 1.2× | novita 71 ms |
| NVIDIA: Nemotron 3 Ultra | 3 | deepinfra $0.500 | together $0.600 | 1.2× | baseten 54 ms |
| NVIDIA: Nemotron 3 Nano 30B A3B | 3 | deepinfra $0.050 | nebius $0.060 | 1.2× | novita 71 ms |
| MiniMax: MiniMax M2 | 3 | minimax $0.255 | novita $0.300 | 1.2× | google 40 ms |
| Thinking Machines: Inkling | 3 | deepinfra $1.000 | together $1.000 | 1.0× | baseten 54 ms |
| StepFun: Step 3.7 Flash | 3 | stepfun $0.200 | novita $0.200 | 1.0× | novita 71 ms |
| MiniMax: MiniMax M3 | 4 | novita $0.300 | deepinfra $0.300 | 1.0× | novita 71 ms |
| MiniMax: MiniMax M2.5 | 4 | friendli $0.300 | siliconflow $0.300 | 1.0× | novita 71 ms |
Why compare hosts only within one model?
Because across models a price comparison means nothing. A cheap small model and an expensive large one are not competing for the same job, and a table that ranks them together ranks nothing. Held to a single model the comparison becomes exact: identical weights, identical task, and the only variables left are what the host charges and how quickly it answers.
Where do these two numbers come from?
They come from different places, and we keep them apart. Latency and uptime are measured by our own probes in 4 regions, every five minutes, against each host’s real API. Prices are not measured — providers publish them, and we read them from OpenRouter’s public catalogue. We do not blend the two into a single “value score”: that number would be our invention rather than either provider’s fact, and the weighting inside it would quietly decide the ranking. The two columns sit side by side and the trade-off stays yours.
Method: how latency is measured · Prices as published 2026-07-25 via OpenRouter · data & licence