Z.ai: GLM 5 — price and measured latency by host

Z.ai: GLM 5 is served by 14 providers; 4 of them are measured here. Input tokens cost $0.600 per million at deepinfra and $1.000 at glm — 1.7× more for the same model. The fastest of them to answer is novita at 71 ms p50 from US (Central).

What does Z.ai: GLM 5 cost at each provider?

HostInput $/MOutput $/MContextFastest p50FromUptime
deepinfra$0.600$2.080202,752226 msUS (Central)99.9%
siliconflow$0.950$2.550204,80082 msUS (Central)100%
novita$1.000$3.200202,80071 msUS (Central)100%
glm$1.000$3.200202,752349 msAsia (Tokyo)100%

Why does the same model cost different amounts?

The weights are identical; what differs is the hardware it runs on, the quantisation the host chose, how much context they will accept, and what margin they take. A host running the model at reduced precision serves it more cheaply and, usually, faster — the trade is accuracy you cannot see in a price table. Where the host published a quantisation tag, it is in the source data linked below.

Are the price and the latency measured the same way?

No, and the difference matters. The latency figures are ours: probes in 4 regions open a real connection to each host’s API and time the response, every five minutes. The prices are not measured — they are what each provider publishes, collected through OpenRouter’s public catalogue on 2026-07-25. Prices in this market change without notice, so treat the table as a pointer and confirm on the provider’s own page before you commit to one.

Does the cheapest host give you the worst latency?

Not reliably, and that is the useful part. Price per token and time-to-first-byte are set by different things — one by GPU economics and margin, the other by where the endpoint sits and how its network is peered. They are worth reading together precisely because they do not move together. Note also that edge latency is the delay before generation starts; it says nothing about tokens per second once the model is running.

Latency: measured by us, see methodology · Prices: published by providers via OpenRouter · All models: price vs latency