Qwen: Qwen3 VL 30B A3B Instruct — price and measured latency by host

Qwen: Qwen3 VL 30B A3B Instruct is served by 5 providers; 3 of them are measured here. Input tokens cost $0.150 per million at deepinfra and $0.290 at siliconflow — 1.9× more for the same model. The fastest of them to answer is novita at 71 ms p50 from US (Central).

What does Qwen: Qwen3 VL 30B A3B Instruct cost at each provider?

HostInput $/MOutput $/MContextFastest p50FromUptime
deepinfra$0.150$0.600262,144226 msUS (Central)99.9%
novita$0.200$0.700131,07271 msUS (Central)100%
siliconflow$0.290$1.000262,14482 msUS (Central)100%

Why does the same model cost different amounts?

The weights are identical; what differs is the hardware it runs on, the quantisation the host chose, how much context they will accept, and what margin they take. A host running the model at reduced precision serves it more cheaply and, usually, faster — the trade is accuracy you cannot see in a price table. Where the host published a quantisation tag, it is in the source data linked below.

Are the price and the latency measured the same way?

No, and the difference matters. The latency figures are ours: probes in 4 regions open a real connection to each host’s API and time the response, every five minutes. The prices are not measured — they are what each provider publishes, collected through OpenRouter’s public catalogue on 2026-07-25. Prices in this market change without notice, so treat the table as a pointer and confirm on the provider’s own page before you commit to one.

Does the cheapest host give you the worst latency?

Not reliably, and that is the useful part. Price per token and time-to-first-byte are set by different things — one by GPU economics and margin, the other by where the endpoint sits and how its network is peered. They are worth reading together precisely because they do not move together. Note also that edge latency is the delay before generation starts; it says nothing about tokens per second once the model is running.

Latency: measured by us, see methodology · Prices: published by providers via OpenRouter · All models: price vs latency