Fastest AI API from Europe (Germany)
· edge + inference latency
As of September 07, 2026, measured from Europe (Germany), the fastest AI inference API by edge latency (time-to-first-byte) is fireworks at 98 ms p50 (n=290).
Which AI API is fastest from Europe (Germany)?
| # | Provider | p50 TTFB | p95 | Uptime | Samples |
|---|---|---|---|---|---|
| 1 | fireworks | 98 ms | 200 ms | 100% | 290 |
| 2 | nscale | 98 ms | 202 ms | 100% | 290 |
| 3 | openrouter | 99 ms | 201 ms | 100% | 290 |
| 4 | 99 ms | 202 ms | 100% | 290 | |
| 5 | aleph-alpha | 99 ms | 202 ms | 100% | 290 |
| 6 | nebius | 100 ms | 203 ms | 100% | 290 |
| 7 | mistral | 100 ms | 294 ms | 100% | 290 |
| 8 | meta-llama | 102 ms | 299 ms | 100% | 290 |
| 9 | baichuan | 181 ms | 250 ms | 100% | 290 |
| 10 | inference-net | 196 ms | 802 ms | 100% | 290 |
| 11 | baseten | 198 ms | 300 ms | 100% | 290 |
| 12 | glm | 199 ms | 392 ms | 100% | 290 |
| 13 | cerebras | 199 ms | 394 ms | 100% | 290 |
| 14 | perplexity | 199 ms | 305 ms | 100% | 290 |
| 15 | anthropic | 200 ms | 305 ms | 100% | 290 |
| 16 | reka | 200 ms | 400 ms | 100% | 290 |
| 17 | replicate | 200 ms | 401 ms | 100% | 290 |
| 18 | cohere | 200 ms | 388 ms | 100% | 290 |
| 19 | openai | 204 ms | 402 ms | 100% | 290 |
| 20 | yi-01ai | 223 ms | 695 ms | 100% | 290 |
| 21 | siliconflow | 293 ms | 395 ms | 100% | 290 |
| 22 | sarvam | 294 ms | 313 ms | 100% | 290 |
| 23 | writer | 296 ms | 403 ms | 100% | 290 |
| 24 | friendli | 296 ms | 452 ms | 100% | 290 |
| 25 | xai | 296 ms | 399 ms | 100% | 290 |
| 26 | groq | 297 ms | 496 ms | 100% | 290 |
| 27 | novita | 297 ms | 401 ms | 100% | 290 |
| 28 | targon | 297 ms | 368 ms | 100% | 290 |
| 29 | together | 298 ms | 596 ms | 100% | 290 |
| 30 | deepseek | 299 ms | 401 ms | 100% | 290 |
| 31 | kimi | 300 ms | 504 ms | 100% | 290 |
| 32 | iflytek | 301 ms | 751 ms | 100% | 290 |
| 33 | hyperbolic | 305 ms | 804 ms | 100% | 290 |
| 34 | ernie | 328 ms | 399 ms | 100% | 290 |
| 35 | doubao | 350 ms | 796 ms | 100% | 290 |
| 36 | sensenova | 391 ms | 1020 ms | 99.7% | 290 |
| 37 | ai21 | 395 ms | 596 ms | 100% | 290 |
| 38 | sambanova | 396 ms | 500 ms | 99.3% | 290 |
| 39 | qwen | 399 ms | 502 ms | 100% | 290 |
| 40 | minimax | 399 ms | 504 ms | 100% | 290 |
| 41 | stepfun | 413 ms | 1088 ms | 100% | 290 |
| 42 | upstage | 565 ms | 601 ms | 100% | 290 |
| 43 | deepinfra | 600 ms | 1094 ms | 100% | 290 |
| 44 | featherless | 699 ms | 1997 ms | 100% | 290 |
| 45 | hunyuan | 1577 ms | 2392 ms | 100% | 290 |
fireworks leads by 1 ms over nscale here. The median provider measured from Europe (Germany) answers in 296 ms, so the spread between the fastest and a typical option is 198 ms on every single request — before the model has produced anything.
Inference latency (time-to-first-token)
Each provider is measured on the model named below — model size affects TTFT far more than the network does, so these numbers are not a like-for-like ranking of providers. Use the edge-latency table above for that. Coverage is limited to providers we hold an API key for.
| # | Provider | Model | p50 TTFT | p95 | Samples |
|---|---|---|---|---|---|
| 1 | groq | openai/gpt-oss-120b | 899 ms | 1393 ms | 72 |
| 2 | gemini-flash-lite-latest | 1005 ms | 1615 ms | 72 |
Why measure from Europe (Germany) separately?
Latency is a property of a route, not of a company. The same provider can lead in one region and trail in another, so a ranking produced from a single location tells you very little about what your users will experience somewhere else. This page reports only what was measured from Europe (Germany); the other regions are ranked independently and often disagree — see the cross-region summary for how far apart they get.
How should these numbers be read?
p50 is the typical request and p95 is the slow tail: if p95 is far above p50, that provider is inconsistent from here, which usually hurts more than a slightly higher median. Uptime counts a probe as failed only on a genuine service failure — an authentication error means the endpoint answered correctly and counts as up. All figures cover the last 24 hours and are recomputed continuously; the methodology page states the limits of this data plainly.