Fastest AI API from South America (São Paulo)

· edge + inference latency

As of September 07, 2026, measured from South America (São Paulo), the fastest AI inference API by edge latency (time-to-first-byte) is openrouter at 56 ms p50 (n=288).

Which AI API is fastest from South America (São Paulo)?

#Providerp50 TTFBp95UptimeSamples
1openrouter56 ms80 ms100%288
2baseten149 ms158 ms100%288
3cohere160 ms188 ms100%288
4google160 ms297 ms100%288
5replicate184 ms468 ms100%288
6cerebras184 ms250 ms100%288
7anthropic198 ms248 ms100%288
8reka198 ms266 ms100%288
9perplexity199 ms446 ms100%288
10meta-llama214 ms300 ms100%288
11inference-net218 ms639 ms100%288
12xai228 ms378 ms100%288
13groq230 ms292 ms100%288
14mistral236 ms299 ms100%288
15openai239 ms462 ms100%288
16novita244 ms311 ms100%288
17writer244 ms279 ms100%288
18fireworks250 ms274 ms100%288
19together261 ms649 ms100%288
20friendli272 ms393 ms100%288
21siliconflow272 ms275 ms100%288
22ai21287 ms355 ms100%288
23targon319 ms329 ms100%288
24kimi330 ms374 ms100%288
25baichuan343 ms1204 ms100%288
26yi-01ai351 ms1246 ms99.7%288
27sambanova371 ms394 ms99.3%288
28nscale408 ms412 ms100%288
29ernie412 ms476 ms100%288
30deepseek436 ms498 ms100%288
31aleph-alpha448 ms504 ms100%288
32nebius473 ms529 ms100%288
33minimax481 ms516 ms100%288
34glm548 ms703 ms100%288
35sarvam578 ms617 ms100%288
36upstage600 ms615 ms100%288
37deepinfra610 ms1426 ms99.7%288
38sensenova629 ms1546 ms100%288
39iflytek651 ms1539 ms99.7%288
40stepfun653 ms1668 ms100%288
41doubao695 ms1565 ms100%288
42qwen701 ms749 ms100%288
43featherless703 ms843 ms100%288
44hyperbolic834 ms908 ms100%288
45hunyuan1657 ms1769 ms100%288

openrouter leads by 93 ms over baseten here. The median provider measured from South America (São Paulo) answers in 319 ms, so the spread between the fastest and a typical option is 263 ms on every single request — before the model has produced anything.

Inference latency (time-to-first-token)

Each provider is measured on the model named below — model size affects TTFT far more than the network does, so these numbers are not a like-for-like ranking of providers. Use the edge-latency table above for that. Coverage is limited to providers we hold an API key for.

#ProviderModelp50 TTFTp95Samples
1groqopenai/gpt-oss-120b739 ms1048 ms72
2googlegemini-flash-lite-latest1290 ms9414 ms72

Why measure from South America (São Paulo) separately?

Latency is a property of a route, not of a company. The same provider can lead in one region and trail in another, so a ranking produced from a single location tells you very little about what your users will experience somewhere else. This page reports only what was measured from South America (São Paulo); the other regions are ranked independently and often disagree — see the cross-region summary for how far apart they get.

How should these numbers be read?

p50 is the typical request and p95 is the slow tail: if p95 is far above p50, that provider is inconsistent from here, which usually hurts more than a slightly higher median. Uptime counts a probe as failed only on a genuine service failure — an authentication error means the endpoint answered correctly and counts as up. All figures cover the last 24 hours and are recomputed continuously; the methodology page states the limits of this data plainly.