AI API price vs measured latency — the same model, different hosts

The same model can cost 11.6× more depending on who hosts it. DeepSeek: DeepSeek V3.2 is $0.259 per million input tokens at siliconflow and $3.000 at sambanova — identical weights, 11.6× the price. Below, every multi-hosted model we track, with what each host charges and how fast we measured it answering.

Which models cost the most and least depending on the host?

ModelHostsCheapestDearestSpreadFastest measured
DeepSeek: DeepSeek V3.26siliconflow $0.259sambanova $3.00011.6×sambanova 23 ms
Google: Gemma 4 31B7deepinfra $0.120cerebras $0.9908.2×sambanova 23 ms
Z.ai: GLM 4.75deepinfra $0.400cerebras $2.2505.6×google 40 ms
DeepSeek: DeepSeek V4 Pro7deepseek $0.435fireworks $1.7404.0×fireworks 17 ms
MiniMax: MiniMax M2.77deepinfra $0.250groq $0.6002.4×fireworks 17 ms
Google: Gemma 4 26B A4B 4deepinfra $0.070google $0.1502.1×google 40 ms
Z.ai: GLM 5.28novita $0.724baseten $1.4001.9×fireworks 17 ms
Qwen: Qwen3 VL 30B A3B Instruct3deepinfra $0.150siliconflow $0.2901.9×novita 71 ms
Qwen: Qwen3.5-9B3siliconflow $0.100together $0.1701.7×siliconflow 81 ms
Z.ai: GLM 54deepinfra $0.600glm $1.0001.7×novita 71 ms
MoonshotAI: Kimi K2.66deepinfra $0.750together $1.2001.6×fireworks 17 ms
DeepSeek: DeepSeek V4 Flash5deepinfra $0.090deepseek $0.1401.6×fireworks 17 ms
Qwen: Qwen3.5-122B-A10B3siliconflow $0.260novita $0.4001.5×novita 71 ms
Z.ai: GLM 5.17deepinfra $1.050glm $1.4001.3×fireworks 17 ms
MoonshotAI: Kimi K2.7 Code5deepinfra $0.740fireworks $0.9501.3×fireworks 17 ms
MoonshotAI: Kimi K2.53deepinfra $0.450novita $0.5701.3×novita 71 ms
Z.ai: GLM 4.63deepinfra $0.500glm $0.6001.2×novita 71 ms
Qwen: Qwen3.5-27B3siliconflow $0.250novita $0.3001.2×novita 71 ms
NVIDIA: Nemotron 3 Ultra3deepinfra $0.500together $0.6001.2×baseten 54 ms
NVIDIA: Nemotron 3 Nano 30B A3B3deepinfra $0.050nebius $0.0601.2×novita 71 ms
MiniMax: MiniMax M23minimax $0.255novita $0.3001.2×google 40 ms
Thinking Machines: Inkling3deepinfra $1.000together $1.0001.0×baseten 54 ms
StepFun: Step 3.7 Flash3stepfun $0.200novita $0.2001.0×novita 71 ms
MiniMax: MiniMax M34novita $0.300deepinfra $0.3001.0×novita 71 ms
MiniMax: MiniMax M2.54friendli $0.300siliconflow $0.3001.0×novita 71 ms

Why compare hosts only within one model?

Because across models a price comparison means nothing. A cheap small model and an expensive large one are not competing for the same job, and a table that ranks them together ranks nothing. Held to a single model the comparison becomes exact: identical weights, identical task, and the only variables left are what the host charges and how quickly it answers.

Where do these two numbers come from?

They come from different places, and we keep them apart. Latency and uptime are measured by our own probes in 4 regions, every five minutes, against each host’s real API. Prices are not measured — providers publish them, and we read them from OpenRouter’s public catalogue. We do not blend the two into a single “value score”: that number would be our invention rather than either provider’s fact, and the weighting inside it would quietly decide the ranking. The two columns sit side by side and the trade-off stays yours.

Method: how latency is measured · Prices as published 2026-07-25 via OpenRouter · data & licence