Which AI provider has the best uptime?
As of July 28, 2026, 27 of 45 providers measured in all 4 regions answered every single probe over the last 24 hours (51,981 checks total). Uptime is not what separates AI APIs today — latency is. Where availability does differ, it usually differs in one region rather than everywhere: kimi sits at 99.9% on average but 99.7% from Asia (Tokyo).
AI API uptime ranking
| Provider | Uptime | Worst region | where | p50 TTFB | Probes |
|---|---|---|---|---|---|
| anthropic | 100% | 100% | Asia (Tokyo) | 194 ms | 1156 |
| 100% | 100% | Asia (Tokyo) | 90 ms | 1156 | |
| together | 100% | 100% | Asia (Tokyo) | 213 ms | 1156 |
| xai | 100% | 100% | Asia (Tokyo) | 206 ms | 1156 |
| ai21 | 100% | 100% | Asia (Tokyo) | 289 ms | 1155 |
| aleph-alpha | 100% | 100% | Asia (Tokyo) | 364 ms | 1155 |
| baseten | 100% | 100% | Asia (Tokyo) | 139 ms | 1155 |
| cerebras | 100% | 100% | Asia (Tokyo) | 165 ms | 1155 |
| cohere | 100% | 100% | Asia (Tokyo) | 148 ms | 1155 |
| doubao | 100% | 100% | Asia (Tokyo) | 451 ms | 1155 |
| featherless | 100% | 100% | Asia (Tokyo) | 568 ms | 1155 |
| fireworks | 100% | 100% | Asia (Tokyo) | 101 ms | 1155 |
| friendli | 100% | 100% | Asia (Tokyo) | 262 ms | 1155 |
| glm | 100% | 100% | Asia (Tokyo) | 547 ms | 1155 |
| hyperbolic | 100% | 100% | Asia (Tokyo) | 456 ms | 1155 |
| inference-net | 100% | 100% | Asia (Tokyo) | 212 ms | 1155 |
| meta-llama | 100% | 100% | Asia (Tokyo) | 120 ms | 1155 |
| minimax | 100% | 100% | Asia (Tokyo) | 318 ms | 1155 |
| mistral | 100% | 100% | Asia (Tokyo) | 215 ms | 1155 |
| novita | 100% | 100% | Asia (Tokyo) | 188 ms | 1155 |
| nscale | 100% | 100% | Asia (Tokyo) | 288 ms | 1155 |
| perplexity | 100% | 100% | Asia (Tokyo) | 192 ms | 1155 |
| qwen | 100% | 100% | Asia (Tokyo) | 433 ms | 1155 |
| reka | 100% | 100% | Asia (Tokyo) | 341 ms | 1155 |
| replicate | 100% | 100% | Asia (Tokyo) | 164 ms | 1155 |
Last 24 hours, providers measured in all 4 regions. Ties broken by sample size: 100% over 200 probes is a stronger claim than 100% over 20. Detected outages →
Which AI providers dropped requests?
| Provider | Uptime | Worst region | where | p50 TTFB | Probes |
|---|---|---|---|---|---|
| kimi | 99.9% | 99.7% | Asia (Tokyo) | 259 ms | 1156 |
| openai | 99.9% | 99.7% | Europe (Germany) | 212 ms | 1156 |
| baichuan | 99.9% | 99.7% | South America (São Paulo) | 235 ms | 1155 |
| deepinfra | 99.9% | 99.7% | South America (São Paulo) | 484 ms | 1155 |
| deepseek | 99.9% | 99.7% | US (Central) | 306 ms | 1155 |
| ernie | 99.9% | 99.7% | South America (São Paulo) | 274 ms | 1155 |
| groq | 99.9% | 99.7% | US (Central) | 198 ms | 1155 |
| hunyuan | 99.9% | 99.7% | South America (São Paulo) | 1422 ms | 1155 |
| nebius | 99.9% | 99.7% | South America (São Paulo) | 329 ms | 1155 |
| openrouter | 99.9% | 99.7% | US (Central) | 66 ms | 1155 |
| sambanova | 99.9% | 99.7% | Europe (Germany) | 234 ms | 1155 |
| sarvam | 99.9% | 99.7% | Europe (Germany) | 398 ms | 1155 |
| targon | 99.9% | 99.7% | Asia (Tokyo) | 258 ms | 1155 |
| writer | 99.9% | 99.7% | US (Central) | 228 ms | 1155 |
| yi-01ai | 99.9% | 99.7% | Asia (Tokyo) | 383 ms | 1155 |
A provider can be perfect on average and still be unreliable for you: what matters is the region your traffic comes from, which is why the worst region is shown next to the average.
Why uptime is the wrong question for AI APIs
Endpoint availability among major inference providers has converged. Over the last 24 hours we recorded 51,981 probes across 45 providers and 4 regions, and almost all of them answered every time. Choosing a provider on uptime alone therefore tells you almost nothing.
What still varies by a wide margin is how fast the same provider answers depending on where the request starts — often by several times for identical infrastructure. That difference is measurable, persistent, and it affects every request you make. See the cross-region latency summary.
Where does this data come from?
Independent probes in 4 regions open a real connection to each provider’s official API host every five minutes and record whether it answered and how quickly. Requests go directly to the provider, never through a gateway or aggregator, so nothing is attributed to the wrong party. The method and its limits are on the methodology page; raw figures are available as JSON under CC BY 4.0.