# Which LLM API has the lowest time to first token?

Updated August 17, 2026. Lowest measured time to first token (TTFT): 757 ms — cerebras on gpt-oss-120b, requested from South America (São Paulo). Not a provider ranking: each provider answers on a different model, and model size drives TTFT more than the provider does.

| Requested from | Provider | Model answering | TTFT p50 | TTFT p95 | Samples |
| --- | --- | --- | --- | --- | --- |
| South America (São Paulo) | cerebras | gpt-oss-120b | 757 ms | 1215 ms | 58 |
| South America (São Paulo) | groq | llama-3.3-70b-versatile | 764 ms | 1126 ms | 58 |
| Europe (Germany) | cerebras | gpt-oss-120b | 801 ms | 1398 ms | 58 |
| Europe (Germany) | groq | llama-3.3-70b-versatile | 897 ms | 1370 ms | 58 |
| South America (São Paulo) | openrouter | nvidia/nemotron-3-nano-30b-a3b:free | 1037 ms | 1343 ms | 58 |
| South America (São Paulo) | google | gemini-flash-lite-latest | 1119 ms | 1724 ms | 58 |
| Europe (Germany) | openrouter | nvidia/nemotron-3-nano-30b-a3b:free | 1264 ms | 2009 ms | 58 |
| Asia (Tokyo) | groq | llama-3.3-70b-versatile | 1280 ms | 1622 ms | 58 |
| Asia (Tokyo) | openrouter | nvidia/nemotron-3-nano-30b-a3b:free | 1292 ms | 1749 ms | 58 |
| Europe (Germany) | google | gemini-flash-lite-latest | 1323 ms | 2384 ms | 58 |
| Asia (Tokyo) | cerebras | gpt-oss-120b | 1324 ms | 1659 ms | 58 |
| US (Central) | groq | llama-3.3-70b-versatile | 1372 ms | 1965 ms | 58 |
| US (Central) | google | gemini-flash-lite-latest | 1486 ms | 2140 ms | 58 |
| Asia (Tokyo) | google | gemini-flash-lite-latest | 1488 ms | 1827 ms | 58 |
| US (Central) | cerebras | gpt-oss-120b | 1491 ms | 1841 ms | 58 |
| US (Central) | openrouter | nvidia/nemotron-3-nano-30b-a3b:free | 1570 ms | 1981 ms | 58 |


## TTFT vs TTFB

TTFB (time to first byte) is DNS + TCP + TLS + first response byte: the network floor, no model involved. TTFT (time to first token) is a real streamed completion's first token: that network time plus queueing plus model prefill.

## How much of TTFT is network latency?

Between 3% and 33% in these measurements; the rest is queueing and model prefill.

| Requested from | Provider | TTFT p50 | TTFB p50 | Network share |
| --- | --- | --- | --- | --- |
| South America (São Paulo) | cerebras | 757 ms | 183 ms | 24% |
| South America (São Paulo) | groq | 764 ms | 233 ms | 31% |
| Europe (Germany) | cerebras | 801 ms | 200 ms | 25% |
| Europe (Germany) | groq | 897 ms | 296 ms | 33% |
| South America (São Paulo) | openrouter | 1037 ms | 58 ms | 6% |
| South America (São Paulo) | google | 1119 ms | 160 ms | 14% |
| Europe (Germany) | openrouter | 1264 ms | 99 ms | 8% |
| Asia (Tokyo) | groq | 1280 ms | 146 ms | 11% |
| Asia (Tokyo) | openrouter | 1292 ms | 55 ms | 4% |
| Europe (Germany) | google | 1323 ms | 101 ms | 8% |
| Asia (Tokyo) | cerebras | 1324 ms | 196 ms | 15% |
| US (Central) | groq | 1372 ms | 108 ms | 8% |
| US (Central) | google | 1486 ms | 42 ms | 3% |
| Asia (Tokyo) | google | 1488 ms | 57 ms | 4% |
| US (Central) | cerebras | 1491 ms | 68 ms | 5% |
| US (Central) | openrouter | 1570 ms | 58 ms | 4% |


Inference is measured for 4 of 45 providers (a paid API key per provider is required); edge latency is measured for all 45.

Method: https://llmlatency.dev/methodology · Data (CC BY 4.0): https://llmlatency.dev/api/rankings.json
Source: https://llmlatency.dev/time-to-first-token