MoonshotAI: Kimi K2.6 — price and measured latency by host
MoonshotAI: Kimi K2.6 is served by 21 providers; 6 of them are measured here. Input tokens cost $0.750 per million at deepinfra and $1.200 at together — 1.6× more for the same model. The fastest of them to answer is fireworks at 17 ms p50 from Asia (Tokyo).
What does MoonshotAI: Kimi K2.6 cost at each provider?
| Host | Input $/M | Output $/M | Context | Fastest p50 | From | Uptime |
|---|---|---|---|---|---|---|
| deepinfra | $0.750 | $3.500 | 262,144 | 226 ms | US (Central) | 99.9% |
| siliconflow | $0.770 | $3.400 | 262,144 | 82 ms | US (Central) | 100% |
| novita | $0.800 | $3.400 | 262,144 | 71 ms | US (Central) | 100% |
| baseten | $0.950 | $4.000 | 262,000 | 54 ms | US (Central) | 100% |
| fireworks | $0.950 | $4.000 | 262,144 | 17 ms | Asia (Tokyo) | 100% |
| together | $1.200 | $4.500 | 262,144 | 111 ms | US (Central) | 100% |
Why does the same model cost different amounts?
The weights are identical; what differs is the hardware it runs on, the quantisation the host chose, how much context they will accept, and what margin they take. A host running the model at reduced precision serves it more cheaply and, usually, faster — the trade is accuracy you cannot see in a price table. Where the host published a quantisation tag, it is in the source data linked below.
Are the price and the latency measured the same way?
No, and the difference matters. The latency figures are ours: probes in 4 regions open a real connection to each host’s API and time the response, every five minutes. The prices are not measured — they are what each provider publishes, collected through OpenRouter’s public catalogue on 2026-07-25. Prices in this market change without notice, so treat the table as a pointer and confirm on the provider’s own page before you commit to one.
Does the cheapest host give you the worst latency?
Not reliably, and that is the useful part. Price per token and time-to-first-byte are set by different things — one by GPU economics and margin, the other by where the endpoint sits and how its network is peered. They are worth reading together precisely because they do not move together. Note also that edge latency is the delay before generation starts; it says nothing about tokens per second once the model is running.
Latency: measured by us, see methodology · Prices: published by providers via OpenRouter · All models: price vs latency