# AI API price vs measured latency — same model, different hosts

Updated July 25, 2026. Latency measured by llmlatency.dev from 4 regions. Prices published by providers, read via OpenRouter on 2026-07-25 — not measured by us.

| Model | Hosts | Cheapest in $/M | Dearest in $/M | Spread | Fastest measured |
| --- | --- | --- | --- | --- | --- |
| Z.ai: GLM 5.2 | 8 | novita $0.724 | baseten $1.400 | 1.9x | fireworks 17ms |
| DeepSeek: DeepSeek V4 Pro | 7 | deepseek $0.435 | fireworks $1.740 | 4.0x | fireworks 17ms |
| Google: Gemma 4 31B | 7 | deepinfra $0.120 | cerebras $0.990 | 8.2x | sambanova 24ms |
| MiniMax: MiniMax M2.7 | 7 | deepinfra $0.250 | groq $0.600 | 2.4x | fireworks 17ms |
| Z.ai: GLM 5.1 | 7 | deepinfra $1.050 | glm $1.400 | 1.3x | fireworks 17ms |
| DeepSeek: DeepSeek V3.2 | 6 | siliconflow $0.259 | sambanova $3.000 | 11.6x | sambanova 24ms |
| MoonshotAI: Kimi K2.6 | 6 | deepinfra $0.750 | together $1.200 | 1.6x | fireworks 17ms |
| DeepSeek: DeepSeek V4 Flash | 5 | deepinfra $0.090 | deepseek $0.140 | 1.6x | fireworks 17ms |
| MoonshotAI: Kimi K2.7 Code | 5 | deepinfra $0.740 | fireworks $0.950 | 1.3x | fireworks 17ms |
| Z.ai: GLM 4.7 | 5 | deepinfra $0.400 | cerebras $2.250 | 5.6x | google 40ms |
| Google: Gemma 4 26B A4B  | 4 | deepinfra $0.070 | google $0.150 | 2.1x | google 40ms |
| MiniMax: MiniMax M2.5 | 4 | friendli $0.300 | siliconflow $0.300 | 1.0x | novita 71ms |
| MiniMax: MiniMax M3 | 4 | novita $0.300 | deepinfra $0.300 | 1.0x | novita 71ms |
| Z.ai: GLM 5 | 4 | deepinfra $0.600 | glm $1.000 | 1.7x | novita 71ms |
| MiniMax: MiniMax M2 | 3 | minimax $0.255 | novita $0.300 | 1.2x | google 40ms |
| MoonshotAI: Kimi K2.5 | 3 | deepinfra $0.450 | novita $0.570 | 1.3x | novita 71ms |
| NVIDIA: Nemotron 3 Nano 30B A3B | 3 | deepinfra $0.050 | nebius $0.060 | 1.2x | novita 71ms |
| NVIDIA: Nemotron 3 Ultra | 3 | deepinfra $0.500 | together $0.600 | 1.2x | baseten 54ms |
| Qwen: Qwen3 VL 30B A3B Instruct | 3 | deepinfra $0.150 | siliconflow $0.290 | 1.9x | novita 71ms |
| Qwen: Qwen3.5-122B-A10B | 3 | siliconflow $0.260 | novita $0.400 | 1.5x | novita 71ms |
| Qwen: Qwen3.5-27B | 3 | siliconflow $0.250 | novita $0.300 | 1.2x | novita 71ms |
| Qwen: Qwen3.5-9B | 3 | siliconflow $0.100 | together $0.170 | 1.7x | siliconflow 82ms |
| StepFun: Step 3.7 Flash | 3 | stepfun $0.200 | novita $0.200 | 1.0x | novita 71ms |
| Thinking Machines: Inkling | 3 | deepinfra $1.000 | together $1.000 | 1.0x | baseten 54ms |
| Z.ai: GLM 4.6 | 3 | deepinfra $0.500 | glm $0.600 | 1.2x | novita 71ms |

Source: https://llmlatency.dev/price-vs-latency · CC BY 4.0
