Fastest AI API from US (Central)
· edge + inference latency
As of September 07, 2026, measured from US (Central), the fastest AI inference API by edge latency (time-to-first-byte) is google at 42 ms p50 (n=288).
Which AI API is fastest from US (Central)?
| # | Provider | p50 TTFB | p95 | Uptime | Samples |
|---|---|---|---|---|---|
| 1 | 42 ms | 99 ms | 100% | 288 | |
| 2 | fireworks | 48 ms | 89 ms | 100% | 288 |
| 3 | openrouter | 55 ms | 107 ms | 100% | 288 |
| 4 | meta-llama | 68 ms | 146 ms | 100% | 288 |
| 5 | baseten | 69 ms | 77 ms | 100% | 288 |
| 6 | cerebras | 71 ms | 140 ms | 100% | 288 |
| 7 | novita | 77 ms | 148 ms | 100% | 288 |
| 8 | cohere | 78 ms | 143 ms | 100% | 288 |
| 9 | siliconflow | 83 ms | 103 ms | 100% | 288 |
| 10 | reka | 85 ms | 153 ms | 100% | 288 |
| 11 | perplexity | 88 ms | 139 ms | 100% | 288 |
| 12 | targon | 90 ms | 99 ms | 100% | 288 |
| 13 | anthropic | 91 ms | 154 ms | 100% | 288 |
| 14 | replicate | 94 ms | 155 ms | 100% | 288 |
| 15 | xai | 110 ms | 171 ms | 100% | 288 |
| 16 | together | 112 ms | 171 ms | 100% | 288 |
| 17 | groq | 118 ms | 190 ms | 100% | 288 |
| 18 | sambanova | 133 ms | 190 ms | 100% | 288 |
| 19 | openai | 133 ms | 222 ms | 100% | 288 |
| 20 | hyperbolic | 136 ms | 286 ms | 100% | 288 |
| 21 | writer | 138 ms | 165 ms | 100% | 288 |
| 22 | friendli | 174 ms | 501 ms | 100% | 288 |
| 23 | inference-net | 176 ms | 762 ms | 100% | 288 |
| 24 | ai21 | 181 ms | 266 ms | 100% | 288 |
| 25 | mistral | 188 ms | 275 ms | 100% | 288 |
| 26 | nscale | 196 ms | 210 ms | 100% | 288 |
| 27 | nebius | 242 ms | 304 ms | 100% | 288 |
| 28 | yi-01ai | 246 ms | 293 ms | 99.7% | 288 |
| 29 | baichuan | 255 ms | 587 ms | 100% | 288 |
| 30 | deepinfra | 258 ms | 868 ms | 100% | 288 |
| 31 | glm | 265 ms | 425 ms | 100% | 288 |
| 32 | kimi | 271 ms | 367 ms | 100% | 288 |
| 33 | ernie | 274 ms | 302 ms | 100% | 288 |
| 34 | aleph-alpha | 304 ms | 342 ms | 100% | 288 |
| 35 | featherless | 305 ms | 421 ms | 100% | 288 |
| 36 | deepseek | 311 ms | 371 ms | 100% | 288 |
| 37 | upstage | 350 ms | 361 ms | 100% | 288 |
| 38 | minimax | 364 ms | 395 ms | 100% | 288 |
| 39 | sarvam | 405 ms | 450 ms | 100% | 288 |
| 40 | stepfun | 415 ms | 486 ms | 100% | 288 |
| 41 | sensenova | 432 ms | 481 ms | 100% | 288 |
| 42 | iflytek | 445 ms | 1109 ms | 100% | 288 |
| 43 | qwen | 450 ms | 520 ms | 100% | 288 |
| 44 | doubao | 493 ms | 574 ms | 100% | 288 |
| 45 | hunyuan | 1384 ms | 1429 ms | 100% | 288 |
google leads by 6 ms over fireworks here. The median provider measured from US (Central) answers in 176 ms, so the spread between the fastest and a typical option is 135 ms on every single request — before the model has produced anything.
Inference latency (time-to-first-token)
Each provider is measured on the model named below — model size affects TTFT far more than the network does, so these numbers are not a like-for-like ranking of providers. Use the edge-latency table above for that. Coverage is limited to providers we hold an API key for.
| # | Provider | Model | p50 TTFT | p95 | Samples |
|---|---|---|---|---|---|
| 1 | groq | openai/gpt-oss-120b | 1362 ms | 1749 ms | 72 |
| 2 | gemini-flash-lite-latest | 1480 ms | 1937 ms | 72 |
Why measure from US (Central) separately?
Latency is a property of a route, not of a company. The same provider can lead in one region and trail in another, so a ranking produced from a single location tells you very little about what your users will experience somewhere else. This page reports only what was measured from US (Central); the other regions are ranked independently and often disagree — see the cross-region summary for how far apart they get.
How should these numbers be read?
p50 is the typical request and p95 is the slow tail: if p95 is far above p50, that provider is inconsistent from here, which usually hurts more than a slightly higher median. Uptime counts a probe as failed only on a genuine service failure — an authentication error means the endpoint answered correctly and counts as up. All figures cover the last 24 hours and are recomputed continuously; the methodology page states the limits of this data plainly.