AI API price vs measured latency — the same model, different hosts

The same model can cost 11.0× more depending on who hosts it. Google: Gemma 4 31B is $0.090 per million input tokens at deepinfra and $0.990 at cerebras — identical weights, 11.0× the price. Below, every multi-hosted model we track, with what each host charges and how fast we measured it answering.

Which models cost the most and least depending on the host?

ModelHostsCheapestDearestSpreadFastest measured
Google: Gemma 4 31B7deepinfra $0.090cerebras $0.99011.0×sambanova 20 ms
DeepSeek: DeepSeek V4 Flash 07318deepinfra $0.060deepseek $0.4407.3×fireworks 19 ms
Z.ai: GLM 5.29deepinfra $0.487glm $1.4002.9×fireworks 19 ms
MiniMax: MiniMax M2.75deepinfra $0.250sambanova $0.6002.4×sambanova 20 ms
MiniMax: MiniMax M35deepinfra $0.280sambanova $0.6002.1×sambanova 20 ms
Google: Gemma 4 26B A4B 4deepinfra $0.070google $0.1502.1×google 44 ms
DeepSeek: DeepSeek V4 Flash Vision Exp5deepinfra $0.216deepseek $0.4402.0×fireworks 19 ms
Z.ai: GLM 5.3 Flash9novita $0.075baseten $0.1502.0×fireworks 19 ms
Qwen: Qwen3.5-9B3siliconflow $0.100together $0.1701.7×siliconflow 79 ms
DeepSeek: DeepSeek V4 Flash 04233deepinfra $0.090novita $0.1401.6×novita 76 ms
Qwen: Qwen3.5-122B-A10B3siliconflow $0.260novita $0.4001.5×novita 76 ms
DeepSeek: DeepSeek V4 Pro 04235fireworks $1.200baseten $1.7401.4×fireworks 19 ms
MoonshotAI: Kimi K2.7 Code6deepinfra $0.680together $0.9501.4×fireworks 19 ms
DeepSeek: DeepSeek V4 Pro 08137novita $0.990deepseek $1.3201.3×fireworks 19 ms
Z.ai: GLM 5.16deepinfra $1.050glm $1.4001.3×novita 76 ms
MoonshotAI: Kimi K2.65deepinfra $0.750fireworks $0.9501.3×fireworks 19 ms
StepFun: Step 3.7 Flash3deepinfra $0.160stepfun $0.2001.2×novita 76 ms
Qwen: Qwen3.5-27B3siliconflow $0.250novita $0.3001.2×novita 76 ms
Z.ai: GLM 5.39reka $1.170glm $1.4001.2×fireworks 19 ms
Meta: Muse Glimmer 30B3deepinfra $0.300together $0.3501.2×fireworks 19 ms
Thinking Machines: Inkling Small3deepinfra $0.450together $0.5001.1×baseten 69 ms
Thinking Machines: Inkling3deepinfra $0.950together $1.0001.1×baseten 69 ms
MoonshotAI: Kimi K34deepinfra $2.850baseten $3.0001.1×fireworks 19 ms
Qwen: Qwen3.8 2.4T A95B4novita $2.000together $2.0001.0×novita 76 ms

Why compare hosts only within one model?

Because across models a price comparison means nothing. A cheap small model and an expensive large one are not competing for the same job, and a table that ranks them together ranks nothing. Held to a single model the comparison becomes exact: identical weights, identical task, and the only variables left are what the host charges and how quickly it answers.

Where do these two numbers come from?

They come from different places, and we keep them apart. Latency and uptime are measured by our own probes in 4 regions, every five minutes, against each host’s real API. Prices are not measured — providers publish them, and we read them from OpenRouter’s public catalogue. We do not blend the two into a single “value score”: that number would be our invention rather than either provider’s fact, and the weighting inside it would quietly decide the ranking. The two columns sit side by side and the trade-off stays yours.

Method: how latency is measured · Prices as published 2026-09-08 via OpenRouter · data & licence