Fastest AI API from Asia (Tokyo)
· edge + inference latency
As of September 07, 2026, measured from Asia (Tokyo), the fastest AI inference API by edge latency (time-to-first-byte) is fireworks at 20 ms p50 (n=288).
Which AI API is fastest from Asia (Tokyo)?
| # | Provider | p50 TTFB | p95 | Uptime | Samples |
|---|---|---|---|---|---|
| 1 | fireworks | 20 ms | 68 ms | 100% | 288 |
| 2 | sambanova | 22 ms | 57 ms | 100% | 288 |
| 3 | openrouter | 48 ms | 102 ms | 100% | 288 |
| 4 | upstage | 58 ms | 77 ms | 100% | 288 |
| 5 | 63 ms | 105 ms | 100% | 288 | |
| 6 | glm | 93 ms | 135 ms | 100% | 288 |
| 7 | meta-llama | 125 ms | 323 ms | 100% | 288 |
| 8 | kimi | 127 ms | 196 ms | 100% | 288 |
| 9 | ernie | 139 ms | 167 ms | 100% | 288 |
| 10 | groq | 143 ms | 199 ms | 100% | 288 |
| 11 | novita | 150 ms | 197 ms | 100% | 288 |
| 12 | baseten | 157 ms | 172 ms | 100% | 288 |
| 13 | qwen | 158 ms | 217 ms | 100% | 288 |
| 14 | xai | 170 ms | 253 ms | 100% | 288 |
| 15 | cohere | 171 ms | 209 ms | 100% | 288 |
| 16 | replicate | 171 ms | 479 ms | 100% | 288 |
| 17 | deepseek | 174 ms | 236 ms | 100% | 288 |
| 18 | baichuan | 178 ms | 195 ms | 100% | 288 |
| 19 | together | 181 ms | 380 ms | 100% | 288 |
| 20 | yi-01ai | 191 ms | 890 ms | 92.4% | 288 |
| 21 | cerebras | 198 ms | 268 ms | 100% | 288 |
| 22 | openai | 205 ms | 283 ms | 99.7% | 288 |
| 23 | minimax | 214 ms | 250 ms | 100% | 288 |
| 24 | iflytek | 230 ms | 1350 ms | 99.3% | 288 |
| 25 | reka | 234 ms | 281 ms | 100% | 288 |
| 26 | anthropic | 239 ms | 298 ms | 100% | 288 |
| 27 | sarvam | 244 ms | 272 ms | 99.7% | 288 |
| 28 | perplexity | 244 ms | 588 ms | 100% | 288 |
| 29 | writer | 258 ms | 292 ms | 100% | 288 |
| 30 | friendli | 278 ms | 506 ms | 99.7% | 288 |
| 31 | stepfun | 293 ms | 759 ms | 100% | 288 |
| 32 | ai21 | 298 ms | 369 ms | 100% | 288 |
| 33 | mistral | 319 ms | 417 ms | 100% | 288 |
| 34 | targon | 331 ms | 392 ms | 100% | 288 |
| 35 | siliconflow | 377 ms | 400 ms | 100% | 288 |
| 36 | sensenova | 443 ms | 1164 ms | 100% | 288 |
| 37 | featherless | 446 ms | 990 ms | 100% | 288 |
| 38 | hyperbolic | 471 ms | 536 ms | 100% | 288 |
| 39 | doubao | 493 ms | 1792 ms | 98.3% | 288 |
| 40 | nebius | 497 ms | 566 ms | 100% | 288 |
| 41 | nscale | 499 ms | 504 ms | 100% | 288 |
| 42 | deepinfra | 518 ms | 739 ms | 100% | 288 |
| 43 | aleph-alpha | 538 ms | 590 ms | 100% | 288 |
| 44 | inference-net | 747 ms | 1350 ms | 100% | 288 |
| 45 | hunyuan | 1117 ms | 1126 ms | 100% | 288 |
fireworks leads by 3 ms over sambanova here. The median provider measured from Asia (Tokyo) answers in 214 ms, so the spread between the fastest and a typical option is 194 ms on every single request — before the model has produced anything.
Inference latency (time-to-first-token)
Each provider is measured on the model named below — model size affects TTFT far more than the network does, so these numbers are not a like-for-like ranking of providers. Use the edge-latency table above for that. Coverage is limited to providers we hold an API key for.
| # | Provider | Model | p50 TTFT | p95 | Samples |
|---|---|---|---|---|---|
| 1 | groq | openai/gpt-oss-120b | 1366 ms | 1768 ms | 72 |
| 2 | gemini-flash-lite-latest | 1487 ms | 2066 ms | 72 |
Why measure from Asia (Tokyo) separately?
Latency is a property of a route, not of a company. The same provider can lead in one region and trail in another, so a ranking produced from a single location tells you very little about what your users will experience somewhere else. This page reports only what was measured from Asia (Tokyo); the other regions are ranked independently and often disagree — see the cross-region summary for how far apart they get.
How should these numbers be read?
p50 is the typical request and p95 is the slow tail: if p95 is far above p50, that provider is inconsistent from here, which usually hurts more than a slightly higher median. Uptime counts a probe as failed only on a genuine service failure — an authentication error means the endpoint answered correctly and counts as up. All figures cover the last 24 hours and are recomputed continuously; the methodology page states the limits of this data plainly.