Which LLM API has the lowest time to first token?
Measured over the last 24 hours, the lowest median time to first token (TTFT) is 491 ms — typesafe running jev-1.13.0, requested from South America (São Paulo). Read that as one measurement, not as a provider ranking: every provider here answers on a different model, and model size moves TTFT far more than the provider does. What is comparable like-for-like is the same provider on the same model from different regions — and that alone varies 1.8× for typesafe.
Measured time to first token, by region
| Requested from | Provider | Model answering | TTFT p50 | TTFT p95 | Samples |
|---|---|---|---|---|---|
| South America (São Paulo) | typesafe | jev-1.13.0 | 491 ms | 696 ms | 72 |
| US (Central) | typesafe | jev-1.13.0 | 659 ms | 844 ms | 72 |
| South America (São Paulo) | groq | openai/gpt-oss-120b | 779 ms | 1362 ms | 72 |
| Asia (Tokyo) | typesafe | jev-1.13.0 | 880 ms | 1104 ms | 72 |
| Europe (Germany) | groq | openai/gpt-oss-120b | 899 ms | 1302 ms | 72 |
| Europe (Germany) | typesafe | jev-1.13.0 | 904 ms | 1399 ms | 72 |
| Europe (Germany) | gemini-flash-lite-latest | 1100 ms | 5502 ms | 72 | |
| South America (São Paulo) | gemini-flash-lite-latest | 1144 ms | 6527 ms | 72 | |
| US (Central) | groq | openai/gpt-oss-120b | 1371 ms | 1836 ms | 72 |
| Asia (Tokyo) | groq | openai/gpt-oss-120b | 1515 ms | 2509 ms | 72 |
| US (Central) | gemini-flash-lite-latest | 1592 ms | 6123 ms | 72 | |
| Asia (Tokyo) | gemini-flash-lite-latest | 1626 ms | 4915 ms | 72 |
Real streamed completions, last 24 hours, requested from each probe region. The model is named on every row because it is the dominant factor in the number next to it.
What is the difference between time to first token and time to first byte?
Time to first byte (TTFB) is how long it takes the network and the provider’s edge to begin responding at all: DNS resolution, TCP connect, TLS handshake and the first byte back. It does not involve the model. Time to first token (TTFT) is how long a real streamed completion takes to emit its first token: that same network time, plus queueing at the provider, plus the model’s prefill over your prompt.
TTFB is the floor — you pay it on every request no matter which model you call, and no prompt engineering removes it. TTFT is what a human actually waits for before words appear on screen. They are different questions and they have different answers, which is why this site measures both and never averages them together.
How much of time to first token is network latency?
Between 3% and 48% of time to first token is network time — the rest is the provider queuing your request and the model generating its first token. That is why a TTFT number quoted without the model name tells you almost nothing about the API: most of what it measures is the model.
| Requested from | Provider | TTFT p50 | TTFB p50 | Network share of TTFT |
|---|---|---|---|---|
| South America (São Paulo) | typesafe | 491 ms | 238 ms | 48% |
| US (Central) | typesafe | 659 ms | 115 ms | 17% |
| South America (São Paulo) | groq | 779 ms | 231 ms | 30% |
| Asia (Tokyo) | typesafe | 880 ms | 149 ms | 17% |
| Europe (Germany) | groq | 899 ms | 297 ms | 33% |
| Europe (Germany) | typesafe | 904 ms | 232 ms | 26% |
| Europe (Germany) | 1100 ms | 100 ms | 9% | |
| South America (São Paulo) | 1144 ms | 161 ms | 14% | |
| US (Central) | groq | 1371 ms | 121 ms | 9% |
| Asia (Tokyo) | groq | 1515 ms | 147 ms | 10% |
| US (Central) | 1592 ms | 43 ms | 3% | |
| Asia (Tokyo) | 1626 ms | 66 ms | 4% |
Can providers be ranked by time to first token?
Not across different models, and that is the honest answer rather than a hedge. Each provider here is measured on whichever model that account can call — the models are named in the table above — and model size drives first-token time far more than the network or the provider’s hardware does. A ranking built from those numbers would be a ranking of models wearing provider names.
Measured from South America (São Paulo), google has the lowest edge latency (161 ms TTFB) while typesafe has the lowest time to first token (491 ms TTFT, on jev-1.13.0). Same region, same probes, same schedule — two different winners. “Fastest AI API” is not a question that has one answer until you say which of the two you are paying for.
Which providers have measured TTFT here, and why not all of them?
Inference is measured for 3 of 46 providers (google, groq, typesafe), because a streamed completion needs a paid API key for that provider and every probe spends tokens. Edge latency needs no key, so TTFB is published for all 46. We would rather publish 3 measured TTFT figures than estimate 46: an estimate presented as a measurement is the one thing this site does not do.
Method in full, including limits: methodology · Edge latency for every provider: fastest AI API by region · Availability: uptime comparison · Raw numbers under CC BY 4.0: JSON API.