Fastest AI API from South America (São Paulo)
· edge + inference latency
As of September 07, 2026, measured from South America (São Paulo), the fastest AI inference API by edge latency (time-to-first-byte) is openrouter at 56 ms p50 (n=288).
Which AI API is fastest from South America (São Paulo)?
| # | Provider | p50 TTFB | p95 | Uptime | Samples |
|---|---|---|---|---|---|
| 1 | openrouter | 56 ms | 80 ms | 100% | 288 |
| 2 | baseten | 149 ms | 158 ms | 100% | 288 |
| 3 | cohere | 160 ms | 188 ms | 100% | 288 |
| 4 | 160 ms | 297 ms | 100% | 288 | |
| 5 | replicate | 184 ms | 468 ms | 100% | 288 |
| 6 | cerebras | 184 ms | 250 ms | 100% | 288 |
| 7 | anthropic | 198 ms | 248 ms | 100% | 288 |
| 8 | reka | 198 ms | 266 ms | 100% | 288 |
| 9 | perplexity | 199 ms | 446 ms | 100% | 288 |
| 10 | meta-llama | 214 ms | 300 ms | 100% | 288 |
| 11 | inference-net | 218 ms | 639 ms | 100% | 288 |
| 12 | xai | 228 ms | 378 ms | 100% | 288 |
| 13 | groq | 230 ms | 292 ms | 100% | 288 |
| 14 | mistral | 236 ms | 299 ms | 100% | 288 |
| 15 | openai | 239 ms | 462 ms | 100% | 288 |
| 16 | novita | 244 ms | 311 ms | 100% | 288 |
| 17 | writer | 244 ms | 279 ms | 100% | 288 |
| 18 | fireworks | 250 ms | 274 ms | 100% | 288 |
| 19 | together | 261 ms | 649 ms | 100% | 288 |
| 20 | friendli | 272 ms | 393 ms | 100% | 288 |
| 21 | siliconflow | 272 ms | 275 ms | 100% | 288 |
| 22 | ai21 | 287 ms | 355 ms | 100% | 288 |
| 23 | targon | 319 ms | 329 ms | 100% | 288 |
| 24 | kimi | 330 ms | 374 ms | 100% | 288 |
| 25 | baichuan | 343 ms | 1204 ms | 100% | 288 |
| 26 | yi-01ai | 351 ms | 1246 ms | 99.7% | 288 |
| 27 | sambanova | 371 ms | 394 ms | 99.3% | 288 |
| 28 | nscale | 408 ms | 412 ms | 100% | 288 |
| 29 | ernie | 412 ms | 476 ms | 100% | 288 |
| 30 | deepseek | 436 ms | 498 ms | 100% | 288 |
| 31 | aleph-alpha | 448 ms | 504 ms | 100% | 288 |
| 32 | nebius | 473 ms | 529 ms | 100% | 288 |
| 33 | minimax | 481 ms | 516 ms | 100% | 288 |
| 34 | glm | 548 ms | 703 ms | 100% | 288 |
| 35 | sarvam | 578 ms | 617 ms | 100% | 288 |
| 36 | upstage | 600 ms | 615 ms | 100% | 288 |
| 37 | deepinfra | 610 ms | 1426 ms | 99.7% | 288 |
| 38 | sensenova | 629 ms | 1546 ms | 100% | 288 |
| 39 | iflytek | 651 ms | 1539 ms | 99.7% | 288 |
| 40 | stepfun | 653 ms | 1668 ms | 100% | 288 |
| 41 | doubao | 695 ms | 1565 ms | 100% | 288 |
| 42 | qwen | 701 ms | 749 ms | 100% | 288 |
| 43 | featherless | 703 ms | 843 ms | 100% | 288 |
| 44 | hyperbolic | 834 ms | 908 ms | 100% | 288 |
| 45 | hunyuan | 1657 ms | 1769 ms | 100% | 288 |
openrouter leads by 93 ms over baseten here. The median provider measured from South America (São Paulo) answers in 319 ms, so the spread between the fastest and a typical option is 263 ms on every single request — before the model has produced anything.
Inference latency (time-to-first-token)
Each provider is measured on the model named below — model size affects TTFT far more than the network does, so these numbers are not a like-for-like ranking of providers. Use the edge-latency table above for that. Coverage is limited to providers we hold an API key for.
| # | Provider | Model | p50 TTFT | p95 | Samples |
|---|---|---|---|---|---|
| 1 | groq | openai/gpt-oss-120b | 739 ms | 1048 ms | 72 |
| 2 | gemini-flash-lite-latest | 1290 ms | 9414 ms | 72 |
Why measure from South America (São Paulo) separately?
Latency is a property of a route, not of a company. The same provider can lead in one region and trail in another, so a ranking produced from a single location tells you very little about what your users will experience somewhere else. This page reports only what was measured from South America (São Paulo); the other regions are ranked independently and often disagree — see the cross-region summary for how far apart they get.
How should these numbers be read?
p50 is the typical request and p95 is the slow tail: if p95 is far above p50, that provider is inconsistent from here, which usually hurts more than a slightly higher median. Uptime counts a probe as failed only on a genuine service failure — an authentication error means the endpoint answered correctly and counts as up. All figures cover the last 24 hours and are recomputed continuously; the methodology page states the limits of this data plainly.