DeepSeek: DeepSeek V3.2 — price and measured latency by host

DeepSeek: DeepSeek V3.2 is served by 14 providers; 6 of them are measured here. Input tokens cost $0.259 per million at siliconflow and $3.000 at sambanova — 11.6× more for the same model. The fastest of them to answer is sambanova at 24 ms p50 from Asia (Tokyo).

What does DeepSeek: DeepSeek V3.2 cost at each provider?

HostInput $/MOutput $/MContextFastest p50FromUptime
siliconflow$0.259$0.420163,84082 msUS (Central)100%
deepinfra$0.260$0.380163,840226 msUS (Central)99.9%
novita$0.269$0.400163,84071 msUS (Central)100%
friendli$0.500$1.500163,840185 msUS (Central)100%
google$0.560$1.680163,84040 msUS (Central)100%
sambanova$3.000$4.50032,76824 msAsia (Tokyo)99.8%

Why does the same model cost different amounts?

The weights are identical; what differs is the hardware it runs on, the quantisation the host chose, how much context they will accept, and what margin they take. A host running the model at reduced precision serves it more cheaply and, usually, faster — the trade is accuracy you cannot see in a price table. Where the host published a quantisation tag, it is in the source data linked below.

Are the price and the latency measured the same way?

No, and the difference matters. The latency figures are ours: probes in 4 regions open a real connection to each host’s API and time the response, every five minutes. The prices are not measured — they are what each provider publishes, collected through OpenRouter’s public catalogue on 2026-07-25. Prices in this market change without notice, so treat the table as a pointer and confirm on the provider’s own page before you commit to one.

Does the cheapest host give you the worst latency?

Not reliably, and that is the useful part. Price per token and time-to-first-byte are set by different things — one by GPU economics and margin, the other by where the endpoint sits and how its network is peered. They are worth reading together precisely because they do not move together. Note also that edge latency is the delay before generation starts; it says nothing about tokens per second once the model is running.

Latency: measured by us, see methodology · Prices: published by providers via OpenRouter · All models: price vs latency