Which AI provider has the best uptime?
Independent, third-party measured availability for 45 AI inference APIs — including OpenAI, Anthropic, Google, Mistral, Groq and DeepSeek — probed directly from 4 regions every five minutes, not read off a status page and not routed through a gateway. As of September 12, 2026, 37 of 45 providers answered every probe over the last 24 hours (51,899 checks); over the full observation window since 2026-07-23 that is 2,609,603 checks. Uptime is not what separates AI APIs today — latency is. Where availability does differ, it usually differs in one region rather than everywhere: OpenAI sits at 99.9% on average but 99.7% from Asia (Tokyo).
AI API uptime comparison — OpenAI, Anthropic, Google, Mistral, Groq, DeepSeek and 39 more
| Provider | Uptime, 24h | Observed availability, 51d | Worst region | where | p50 TTFB | Probes |
|---|---|---|---|---|---|---|
| anthropic | 100% | 100% | 100% | Asia (Tokyo) | 182 ms | 1154 |
| cohere | 100% | 99.9% | 100% | Asia (Tokyo) | 150 ms | 1154 |
| deepseek | 100% | 100% | 100% | Asia (Tokyo) | 308 ms | 1154 |
| 100% | 100% | 100% | Asia (Tokyo) | 91 ms | 1154 | |
| groq | 100% | 100% | 100% | Asia (Tokyo) | 201 ms | 1154 |
| hyperbolic | 100% | 100% | 100% | Asia (Tokyo) | 456 ms | 1154 |
| meta-llama | 100% | 99.9% | 100% | Asia (Tokyo) | 124 ms | 1154 |
| mistral | 100% | 100% | 100% | Asia (Tokyo) | 210 ms | 1154 |
| openrouter | 100% | 100% | 100% | Asia (Tokyo) | 65 ms | 1154 |
| reka | 100% | 100% | 100% | Asia (Tokyo) | 181 ms | 1154 |
| replicate | 100% | 100% | 100% | Asia (Tokyo) | 170 ms | 1154 |
| xai | 100% | 100% | 100% | Asia (Tokyo) | 201 ms | 1154 |
| ai21 | 100% | 100% | 100% | Asia (Tokyo) | 311 ms | 1153 |
| aleph-alpha | 100% | 99.9% | 100% | Asia (Tokyo) | 340 ms | 1153 |
| baseten | 100% | 99.9% | 100% | Asia (Tokyo) | 138 ms | 1153 |
| cerebras | 100% | 100% | 100% | Asia (Tokyo) | 163 ms | 1153 |
| deepinfra | 100% | 99.9% | 100% | Asia (Tokyo) | 496 ms | 1153 |
| ernie | 100% | 100% | 100% | Asia (Tokyo) | 277 ms | 1153 |
| featherless | 100% | 99.9% | 100% | Asia (Tokyo) | 546 ms | 1153 |
| fireworks | 100% | 100% | 100% | Asia (Tokyo) | 101 ms | 1153 |
| friendli | 100% | 98.6% | 100% | Asia (Tokyo) | 262 ms | 1153 |
| glm | 100% | 99.9% | 100% | Asia (Tokyo) | 275 ms | 1153 |
| hunyuan | 100% | 99.9% | 100% | Asia (Tokyo) | 1435 ms | 1153 |
| inference-net | 100% | 100% | 100% | Asia (Tokyo) | 249 ms | 1153 |
| kimi | 100% | 100% | 100% | Asia (Tokyo) | 260 ms | 1153 |
Two windows on purpose: 24 hours answers “is it working now”, the 51-day column answers “is it reliable”. Observation started 2026-07-23; the raw series is retained for 90 days, so this window grows until it reaches that. Providers measured in all 4 regions only. Ties broken by sample size: 100% over 200 probes is a stronger claim than 100% over 20. Detected outages and incident history →
Which AI providers dropped requests?
| Provider | Uptime, 24h | Observed availability, 51d | Worst region | where | p50 TTFB | Probes |
|---|---|---|---|---|---|---|
| openai | 99.9% | 97.6% | 99.7% | Asia (Tokyo) | 191 ms | 1154 |
| baichuan | 99.9% | 99.8% | 99.7% | Europe (Germany) | 230 ms | 1153 |
| perplexity | 99.8% | 100% | 99.3% | South America (São Paulo) | 184 ms | 1154 |
| writer | 99.8% | 99.9% | 99.3% | Europe (Germany) | 231 ms | 1153 |
| doubao | 99.7% | 99.6% | 99.3% | Europe (Germany) | 498 ms | 1153 |
| iflytek | 99.7% | 99.5% | 99.3% | Asia (Tokyo) | 396 ms | 1153 |
| sambanova | 99.5% | 99.7% | 99.0% | South America (São Paulo) | 227 ms | 1153 |
| yi-01ai | 97.4% | 97.1% | 89.9% | Asia (Tokyo) | 249 ms | 1153 |
A provider can be perfect on average and still be unreliable for you: what matters is the region your traffic comes from, which is why the worst region is shown next to the average.
Does AI API availability differ by region?
Yes, and by more than the headline numbers suggest. The same providers are probed from every region on the same schedule, so a gap between rows is a property of the path to the provider, not of different sampling. Over the last 7 days:
| Region | Probes | Failed | Failure rate | Observed availability |
|---|---|---|---|---|
| US (Central) | 91,730 | 48 | 0.052% | 99.948% |
| Europe (Germany) | 91,881 | 69 | 0.075% | 99.925% |
| South America (São Paulo) | 91,839 | 219 | 0.238% | 99.762% |
| Asia (Tokyo) | 91,730 | 312 | 0.340% | 99.660% |
Asia (Tokyo) failed 6.5× more often than US (Central) (0.340% against 0.052%) across identical requests to identical endpoints. A monitor that probes from a single location — which is what most uptime dashboards are — cannot see this difference at all, and will report whichever number its own vantage point happens to produce.
How many AI providers had at least one failure this week?
24 of 45 tracked providers dropped at least one probe in the last 7 days — against 8 of 45 in the last 24 hours. That gap is the whole point of looking at more than a day: at 24-hour resolution almost everything reads as 100%, and the failures that do matter are rare enough to hide inside a short window. Availability figures quoted without a window length are not comparable to anything.
Which providers publish a machine-readable status feed?
Only 10 of the 45 providers measured here expose a status feed a program can read. For the rest there is nothing to check a failure against: no feed, no timestamps, no components — which is worth knowing before you plan a fallback around someone’s status page. Where a feed does exist, we compare our observations against it on the incident log, and both outcomes are informative: a matching entry means two independent records agree, a missing one means we saw a failure their page never mentioned.
Why uptime is the wrong question for AI APIs
Here is the same fleet measured on speed instead. Over the last 24 hours the fastest endpoint answered in 65 ms (openrouter) and the slowest in 1435 ms (hunyuan) — a 21.9× spread between providers whose availability is indistinguishable. Uptime separates nobody; latency separates everybody.
Endpoint availability among major inference providers has converged. Over the last 24 hours we recorded 51,899 probes across 45 providers and 4 regions, and almost all of them answered every time. Choosing a provider on uptime alone therefore tells you almost nothing.
What still varies by a wide margin is how fast the same provider answers depending on where the request starts — often by several times for identical infrastructure. That difference is measurable, persistent, and it affects every request you make. See the cross-region latency summary.
Observed availability vs published SLA
The figures above are observed availability, not an SLA. They report the share of probes that received a valid HTTP response from the provider’s official API host, measured from outside their network. A published SLA is a contractual promise — usually enterprise-only, measured by the provider, and often excluding exactly the failure modes a client notices first. Neither number substitutes for the other, and this page deliberately does not restate vendor SLA percentages: we publish what we measured, and the contractual terms are the vendor’s to state.
Is there an independent third-party monitor for LLM API uptime?
This is one. Two properties are what make third-party monitoring worth anything, and both are checkable here: probes run directly against the provider’s own endpoint, never through an aggregator or gateway that would attribute its own failures to the provider; and they run from 4 regions at once, which is the only way to separate a provider outage from a broken network path. Monitors built on status-page polling inherit the provider’s blind spots by construction — they can only report what was already admitted.
Compare the uptime of two specific providers
Each pair below has its own page: observed availability for both, the worst region for each, and latency per region side by side.
- anthropic vs openai uptime and latency
- anthropic vs google uptime and latency
- google vs openai uptime and latency
- cerebras vs groq uptime and latency
- deepseek vs openai uptime and latency
- groq vs sambanova uptime and latency
- fireworks vs together uptime and latency
- mistral vs openai uptime and latency
Where does this data come from?
Independent probes in 4 regions open a real connection to each provider’s official API host every five minutes and record whether it answered and how quickly. Requests go directly to the provider, never through a gateway or aggregator, so nothing is attributed to the wrong party. The method and its limits are on the methodology page; raw figures are available as JSON under CC BY 4.0.