Fastest AI API from Europe (Germany)

· edge + inference latency

As of September 07, 2026, measured from Europe (Germany), the fastest AI inference API by edge latency (time-to-first-byte) is fireworks at 98 ms p50 (n=290).

Which AI API is fastest from Europe (Germany)?

#Providerp50 TTFBp95UptimeSamples
1fireworks98 ms200 ms100%290
2nscale98 ms202 ms100%290
3openrouter99 ms201 ms100%290
4google99 ms202 ms100%290
5aleph-alpha99 ms202 ms100%290
6nebius100 ms203 ms100%290
7mistral100 ms294 ms100%290
8meta-llama102 ms299 ms100%290
9baichuan181 ms250 ms100%290
10inference-net196 ms802 ms100%290
11baseten198 ms300 ms100%290
12glm199 ms392 ms100%290
13cerebras199 ms394 ms100%290
14perplexity199 ms305 ms100%290
15anthropic200 ms305 ms100%290
16reka200 ms400 ms100%290
17replicate200 ms401 ms100%290
18cohere200 ms388 ms100%290
19openai204 ms402 ms100%290
20yi-01ai223 ms695 ms100%290
21siliconflow293 ms395 ms100%290
22sarvam294 ms313 ms100%290
23writer296 ms403 ms100%290
24friendli296 ms452 ms100%290
25xai296 ms399 ms100%290
26groq297 ms496 ms100%290
27novita297 ms401 ms100%290
28targon297 ms368 ms100%290
29together298 ms596 ms100%290
30deepseek299 ms401 ms100%290
31kimi300 ms504 ms100%290
32iflytek301 ms751 ms100%290
33hyperbolic305 ms804 ms100%290
34ernie328 ms399 ms100%290
35doubao350 ms796 ms100%290
36sensenova391 ms1020 ms99.7%290
37ai21395 ms596 ms100%290
38sambanova396 ms500 ms99.3%290
39qwen399 ms502 ms100%290
40minimax399 ms504 ms100%290
41stepfun413 ms1088 ms100%290
42upstage565 ms601 ms100%290
43deepinfra600 ms1094 ms100%290
44featherless699 ms1997 ms100%290
45hunyuan1577 ms2392 ms100%290

fireworks leads by 1 ms over nscale here. The median provider measured from Europe (Germany) answers in 296 ms, so the spread between the fastest and a typical option is 198 ms on every single request — before the model has produced anything.

Inference latency (time-to-first-token)

Each provider is measured on the model named below — model size affects TTFT far more than the network does, so these numbers are not a like-for-like ranking of providers. Use the edge-latency table above for that. Coverage is limited to providers we hold an API key for.

#ProviderModelp50 TTFTp95Samples
1groqopenai/gpt-oss-120b899 ms1393 ms72
2googlegemini-flash-lite-latest1005 ms1615 ms72

Why measure from Europe (Germany) separately?

Latency is a property of a route, not of a company. The same provider can lead in one region and trail in another, so a ranking produced from a single location tells you very little about what your users will experience somewhere else. This page reports only what was measured from Europe (Germany); the other regions are ranked independently and often disagree — see the cross-region summary for how far apart they get.

How should these numbers be read?

p50 is the typical request and p95 is the slow tail: if p95 is far above p50, that provider is inconsistent from here, which usually hurts more than a slightly higher median. Uptime counts a probe as failed only on a genuine service failure — an authentication error means the endpoint answered correctly and counts as up. All figures cover the last 24 hours and are recomputed continuously; the methodology page states the limits of this data plainly.