Fastest AI API from US (Central)

· edge + inference latency

As of September 07, 2026, measured from US (Central), the fastest AI inference API by edge latency (time-to-first-byte) is google at 42 ms p50 (n=288).

Which AI API is fastest from US (Central)?

#Providerp50 TTFBp95UptimeSamples
1google42 ms99 ms100%288
2fireworks48 ms89 ms100%288
3openrouter55 ms107 ms100%288
4meta-llama68 ms146 ms100%288
5baseten69 ms77 ms100%288
6cerebras71 ms140 ms100%288
7novita77 ms148 ms100%288
8cohere78 ms143 ms100%288
9siliconflow83 ms103 ms100%288
10reka85 ms153 ms100%288
11perplexity88 ms139 ms100%288
12targon90 ms99 ms100%288
13anthropic91 ms154 ms100%288
14replicate94 ms155 ms100%288
15xai110 ms171 ms100%288
16together112 ms171 ms100%288
17groq118 ms190 ms100%288
18sambanova133 ms190 ms100%288
19openai133 ms222 ms100%288
20hyperbolic136 ms286 ms100%288
21writer138 ms165 ms100%288
22friendli174 ms501 ms100%288
23inference-net176 ms762 ms100%288
24ai21181 ms266 ms100%288
25mistral188 ms275 ms100%288
26nscale196 ms210 ms100%288
27nebius242 ms304 ms100%288
28yi-01ai246 ms293 ms99.7%288
29baichuan255 ms587 ms100%288
30deepinfra258 ms868 ms100%288
31glm265 ms425 ms100%288
32kimi271 ms367 ms100%288
33ernie274 ms302 ms100%288
34aleph-alpha304 ms342 ms100%288
35featherless305 ms421 ms100%288
36deepseek311 ms371 ms100%288
37upstage350 ms361 ms100%288
38minimax364 ms395 ms100%288
39sarvam405 ms450 ms100%288
40stepfun415 ms486 ms100%288
41sensenova432 ms481 ms100%288
42iflytek445 ms1109 ms100%288
43qwen450 ms520 ms100%288
44doubao493 ms574 ms100%288
45hunyuan1384 ms1429 ms100%288

google leads by 6 ms over fireworks here. The median provider measured from US (Central) answers in 176 ms, so the spread between the fastest and a typical option is 135 ms on every single request — before the model has produced anything.

Inference latency (time-to-first-token)

Each provider is measured on the model named below — model size affects TTFT far more than the network does, so these numbers are not a like-for-like ranking of providers. Use the edge-latency table above for that. Coverage is limited to providers we hold an API key for.

#ProviderModelp50 TTFTp95Samples
1groqopenai/gpt-oss-120b1362 ms1749 ms72
2googlegemini-flash-lite-latest1480 ms1937 ms72

Why measure from US (Central) separately?

Latency is a property of a route, not of a company. The same provider can lead in one region and trail in another, so a ranking produced from a single location tells you very little about what your users will experience somewhere else. This page reports only what was measured from US (Central); the other regions are ranked independently and often disagree — see the cross-region summary for how far apart they get.

How should these numbers be read?

p50 is the typical request and p95 is the slow tail: if p95 is far above p50, that provider is inconsistent from here, which usually hurts more than a slightly higher median. Uptime counts a probe as failed only on a genuine service failure — an authentication error means the endpoint answered correctly and counts as up. All figures cover the last 24 hours and are recomputed continuously; the methodology page states the limits of this data plainly.