AI Latency Tracker
· 45 providers · 4 regions
As of September 06, 2026, the fastest-responding AI inference API by edge latency (time-to-first-byte) is fireworks at 20 ms p50, measured from Asia (Tokyo) (100% uptime, n=288). Rankings differ by region — see the table below.
Independent, provider-neutral latency and uptime of AI inference APIs (OpenAI, Anthropic, Google, Mistral, DeepSeek, xAI, OpenRouter and more), measured directly from multiple regions and updated automatically. How this is measured →
Is any AI API down right now?
1 confirmed outage(s) in the last 24 hours: yi-01ai
45 providers · probed every 5 minutes from 4 regions · last probe Sep 6, 17:10 UTC
None of the 10 providers with a machine-readable status page reports an open incident (checked Sep 6, 16:17 UTC).
all probes succeededsome probes failedhalf or more failedno data
Hour by hour for the last 72 hours, all regions combined; hover a cell for the count. Each provider page shows the same view per region. Full incident log →
Which AI API is fastest right now?
| Region | Fastest AI API | p50 TTFB | Uptime |
|---|---|---|---|
| Asia (Tokyo) | fireworks | 20 ms | 100% |
| Europe (Germany) | fireworks | 98 ms | 100% |
| South America (São Paulo) | openrouter | 56 ms | 100% |
| US (Central) | 40 ms | 100% |
Which AI APIs are fastest overall?
Composite score 0–100 (100 = as fast as the region leader), averaged across all 4 measured region(s). How it’s computed →
| # | Provider | Speed Index |
|---|---|---|
| 1 | openrouter | 79 |
| 2 | fireworks | 77 |
| 3 | 66 | |
| 4 | meta-llama | 49 |
| 5 | sambanova | 41 |
| 6 | baseten | 40 |
| 7 | mistral | 37 |
| 8 | cerebras | 37 |
| 9 | cohere | 37 |
| 10 | inference-net | 37 |
| 11 | nscale | 34 |
| 12 | replicate | 34 |
| 13 | reka | 34 |
| 14 | perplexity | 33 |
| 15 | anthropic | 33 |
Full ranking — Europe (Germany)
Edge latency (time-to-first-byte), p50/p95, last 24h. Other regions: see the per-region pages below.
| # | Provider | p50 TTFB | p95 | Uptime | Samples |
|---|---|---|---|---|---|
| 1 | fireworks | 98 ms | 200 ms | 100% | 288 |
| 2 | nscale | 99 ms | 202 ms | 100% | 288 |
| 3 | openrouter | 99 ms | 201 ms | 100% | 288 |
| 4 | aleph-alpha | 99 ms | 201 ms | 100% | 288 |
| 5 | 99 ms | 204 ms | 100% | 288 | |
| 6 | nebius | 100 ms | 205 ms | 100% | 288 |
| 7 | mistral | 100 ms | 296 ms | 100% | 288 |
| 8 | meta-llama | 101 ms | 299 ms | 100% | 288 |
| 9 | inference-net | 107 ms | 801 ms | 100% | 288 |
| 10 | baichuan | 181 ms | 605 ms | 100% | 288 |
| 11 | baseten | 198 ms | 300 ms | 100% | 288 |
| 12 | cerebras | 199 ms | 382 ms | 100% | 288 |
| 13 | glm | 199 ms | 302 ms | 100% | 288 |
| 14 | perplexity | 199 ms | 396 ms | 100% | 288 |
| 15 | anthropic | 200 ms | 302 ms | 100% | 288 |
How fast is each region?
- Fastest AI API from Asia (Tokyo)
- Fastest AI API from Europe (Germany)
- Fastest AI API from South America (São Paulo)
- Fastest AI API from US (Central)
How fast is a specific provider?
- ai21 latency by region
- aleph-alpha latency by region
- anthropic latency by region
- baichuan latency by region
- baseten latency by region
- cerebras latency by region
- cohere latency by region
- deepinfra latency by region
- deepseek latency by region
- doubao latency by region
- ernie latency by region
- featherless latency by region
- fireworks latency by region
- friendli latency by region
- glm latency by region
- google latency by region
- groq latency by region
- hunyuan latency by region
- hyperbolic latency by region
- iflytek latency by region
- inference-net latency by region
- kimi latency by region
- meta-llama latency by region
- minimax latency by region
- mistral latency by region
- nebius latency by region
- novita latency by region
- nscale latency by region
- openai latency by region
- openrouter latency by region
- perplexity latency by region
- qwen latency by region
- reka latency by region
- replicate latency by region
- sambanova latency by region
- sarvam latency by region
- sensenova latency by region
- siliconflow latency by region
- stepfun latency by region
- targon latency by region
- together latency by region
- upstage latency by region
- writer latency by region
- xai latency by region
- yi-01ai latency by region