Which LLM API has the lowest time to first token?

Measured over the last 24 hours, the lowest median time to first token (TTFT) is 491 ms — typesafe running jev-1.13.0, requested from South America (São Paulo). Read that as one measurement, not as a provider ranking: every provider here answers on a different model, and model size moves TTFT far more than the provider does. What is comparable like-for-like is the same provider on the same model from different regions — and that alone varies 1.8× for typesafe.

Measured time to first token, by region

Requested fromProviderModel answeringTTFT p50TTFT p95Samples
South America (São Paulo)typesafejev-1.13.0491 ms696 ms72
US (Central)typesafejev-1.13.0659 ms844 ms72
South America (São Paulo)groqopenai/gpt-oss-120b779 ms1362 ms72
Asia (Tokyo)typesafejev-1.13.0880 ms1104 ms72
Europe (Germany)groqopenai/gpt-oss-120b899 ms1302 ms72
Europe (Germany)typesafejev-1.13.0904 ms1399 ms72
Europe (Germany)googlegemini-flash-lite-latest1100 ms5502 ms72
South America (São Paulo)googlegemini-flash-lite-latest1144 ms6527 ms72
US (Central)groqopenai/gpt-oss-120b1371 ms1836 ms72
Asia (Tokyo)groqopenai/gpt-oss-120b1515 ms2509 ms72
US (Central)googlegemini-flash-lite-latest1592 ms6123 ms72
Asia (Tokyo)googlegemini-flash-lite-latest1626 ms4915 ms72

Real streamed completions, last 24 hours, requested from each probe region. The model is named on every row because it is the dominant factor in the number next to it.

What is the difference between time to first token and time to first byte?

Time to first byte (TTFB) is how long it takes the network and the provider’s edge to begin responding at all: DNS resolution, TCP connect, TLS handshake and the first byte back. It does not involve the model. Time to first token (TTFT) is how long a real streamed completion takes to emit its first token: that same network time, plus queueing at the provider, plus the model’s prefill over your prompt.

TTFB is the floor — you pay it on every request no matter which model you call, and no prompt engineering removes it. TTFT is what a human actually waits for before words appear on screen. They are different questions and they have different answers, which is why this site measures both and never averages them together.

How much of time to first token is network latency?

Between 3% and 48% of time to first token is network time — the rest is the provider queuing your request and the model generating its first token. That is why a TTFT number quoted without the model name tells you almost nothing about the API: most of what it measures is the model.

Requested fromProviderTTFT p50TTFB p50Network share of TTFT
South America (São Paulo)typesafe491 ms238 ms48%
US (Central)typesafe659 ms115 ms17%
South America (São Paulo)groq779 ms231 ms30%
Asia (Tokyo)typesafe880 ms149 ms17%
Europe (Germany)groq899 ms297 ms33%
Europe (Germany)typesafe904 ms232 ms26%
Europe (Germany)google1100 ms100 ms9%
South America (São Paulo)google1144 ms161 ms14%
US (Central)groq1371 ms121 ms9%
Asia (Tokyo)groq1515 ms147 ms10%
US (Central)google1592 ms43 ms3%
Asia (Tokyo)google1626 ms66 ms4%

Can providers be ranked by time to first token?

Not across different models, and that is the honest answer rather than a hedge. Each provider here is measured on whichever model that account can call — the models are named in the table above — and model size drives first-token time far more than the network or the provider’s hardware does. A ranking built from those numbers would be a ranking of models wearing provider names.

Measured from South America (São Paulo), google has the lowest edge latency (161 ms TTFB) while typesafe has the lowest time to first token (491 ms TTFT, on jev-1.13.0). Same region, same probes, same schedule — two different winners. “Fastest AI API” is not a question that has one answer until you say which of the two you are paying for.

Which providers have measured TTFT here, and why not all of them?

Inference is measured for 3 of 46 providers (google, groq, typesafe), because a streamed completion needs a paid API key for that provider and every probe spends tokens. Edge latency needs no key, so TTFB is published for all 46. We would rather publish 3 measured TTFT figures than estimate 46: an estimate presented as a measurement is the one thing this site does not do.

Method in full, including limits: methodology · Edge latency for every provider: fastest AI API by region · Availability: uptime comparison · Raw numbers under CC BY 4.0: JSON API.