Z.ai: GLM 5.2 — price and measured latency by host

Z.ai: GLM 5.2 is served by 33 providers; 8 of them are measured here. Input tokens cost $0.724 per million at novita and $1.400 at baseten — 1.9× more for the same model. The fastest of them to answer is fireworks at 17 ms p50 from Asia (Tokyo).

What does Z.ai: GLM 5.2 cost at each provider?

HostInput $/MOutput $/MContextFastest p50FromUptime
novita$0.724$2.2751,048,57671 msUS (Central)100%
deepinfra$0.930$3.0001,048,576226 msUS (Central)99.9%
siliconflow$1.302$4.0921,048,57682 msUS (Central)100%
glm$1.400$4.4001,048,576349 msAsia (Tokyo)100%
fireworks$1.400$4.4001,048,57617 msAsia (Tokyo)100%
friendli$1.400$4.4001,048,576185 msUS (Central)100%
together$1.400$4.400262,144111 msUS (Central)100%
baseten$1.400$4.400524,28854 msUS (Central)100%

Why does the same model cost different amounts?

The weights are identical; what differs is the hardware it runs on, the quantisation the host chose, how much context they will accept, and what margin they take. A host running the model at reduced precision serves it more cheaply and, usually, faster — the trade is accuracy you cannot see in a price table. Where the host published a quantisation tag, it is in the source data linked below.

Are the price and the latency measured the same way?

No, and the difference matters. The latency figures are ours: probes in 4 regions open a real connection to each host’s API and time the response, every five minutes. The prices are not measured — they are what each provider publishes, collected through OpenRouter’s public catalogue on 2026-07-25. Prices in this market change without notice, so treat the table as a pointer and confirm on the provider’s own page before you commit to one.

Does the cheapest host give you the worst latency?

Not reliably, and that is the useful part. Price per token and time-to-first-byte are set by different things — one by GPU economics and margin, the other by where the endpoint sits and how its network is peered. They are worth reading together precisely because they do not move together. Note also that edge latency is the delay before generation starts; it says nothing about tokens per second once the model is running.

Latency: measured by us, see methodology · Prices: published by providers via OpenRouter · All models: price vs latency