Fastest AI API from Asia (Tokyo)

· edge + inference latency

As of September 07, 2026, measured from Asia (Tokyo), the fastest AI inference API by edge latency (time-to-first-byte) is fireworks at 20 ms p50 (n=288).

Which AI API is fastest from Asia (Tokyo)?

#Providerp50 TTFBp95UptimeSamples
1fireworks20 ms68 ms100%288
2sambanova22 ms57 ms100%288
3openrouter48 ms102 ms100%288
4upstage58 ms77 ms100%288
5google63 ms105 ms100%288
6glm93 ms135 ms100%288
7meta-llama125 ms323 ms100%288
8kimi127 ms196 ms100%288
9ernie139 ms167 ms100%288
10groq143 ms199 ms100%288
11novita150 ms197 ms100%288
12baseten157 ms172 ms100%288
13qwen158 ms217 ms100%288
14xai170 ms253 ms100%288
15cohere171 ms209 ms100%288
16replicate171 ms479 ms100%288
17deepseek174 ms236 ms100%288
18baichuan178 ms195 ms100%288
19together181 ms380 ms100%288
20yi-01ai191 ms890 ms92.4%288
21cerebras198 ms268 ms100%288
22openai205 ms283 ms99.7%288
23minimax214 ms250 ms100%288
24iflytek230 ms1350 ms99.3%288
25reka234 ms281 ms100%288
26anthropic239 ms298 ms100%288
27sarvam244 ms272 ms99.7%288
28perplexity244 ms588 ms100%288
29writer258 ms292 ms100%288
30friendli278 ms506 ms99.7%288
31stepfun293 ms759 ms100%288
32ai21298 ms369 ms100%288
33mistral319 ms417 ms100%288
34targon331 ms392 ms100%288
35siliconflow377 ms400 ms100%288
36sensenova443 ms1164 ms100%288
37featherless446 ms990 ms100%288
38hyperbolic471 ms536 ms100%288
39doubao493 ms1792 ms98.3%288
40nebius497 ms566 ms100%288
41nscale499 ms504 ms100%288
42deepinfra518 ms739 ms100%288
43aleph-alpha538 ms590 ms100%288
44inference-net747 ms1350 ms100%288
45hunyuan1117 ms1126 ms100%288

fireworks leads by 3 ms over sambanova here. The median provider measured from Asia (Tokyo) answers in 214 ms, so the spread between the fastest and a typical option is 194 ms on every single request — before the model has produced anything.

Inference latency (time-to-first-token)

Each provider is measured on the model named below — model size affects TTFT far more than the network does, so these numbers are not a like-for-like ranking of providers. Use the edge-latency table above for that. Coverage is limited to providers we hold an API key for.

#ProviderModelp50 TTFTp95Samples
1groqopenai/gpt-oss-120b1366 ms1768 ms72
2googlegemini-flash-lite-latest1487 ms2066 ms72

Why measure from Asia (Tokyo) separately?

Latency is a property of a route, not of a company. The same provider can lead in one region and trail in another, so a ranking produced from a single location tells you very little about what your users will experience somewhere else. This page reports only what was measured from Asia (Tokyo); the other regions are ranked independently and often disagree — see the cross-region summary for how far apart they get.

How should these numbers be read?

p50 is the typical request and p95 is the slow tail: if p95 is far above p50, that provider is inconsistent from here, which usually hurts more than a slightly higher median. Uptime counts a probe as failed only on a genuine service failure — an authentication error means the endpoint answered correctly and counts as up. All figures cover the last 24 hours and are recomputed continuously; the methodology page states the limits of this data plainly.