AI glossary · Models and context
What is time to first token (TTFT)?
Also called: TTFT, first-token latency
Definition
Time to first token (TTFT) is how long a model takes from receiving a request to returning the first token of its reply, the wait a user feels before streamed text starts to appear.
Explained
How it works
NVIDIA’s benchmarking docs define TTFT as the time from submitting a query to receiving the first token, and note it generally includes queueing, prefill (processing the whole prompt) and network time, so longer prompts raise it. End-to-end latency is TTFT plus the time to generate every remaining token.
Three related numbers get mixed up. Inter-token latency is the average gap between tokens after the first. Tokens per second per user is the reply length divided by end-to-end latency. Total tokens per second for a system adds up every request being served at once.
On reasoning models the thinking happens before the first visible token, so TTFT can stretch to many seconds. Streaming doesn’t make the model faster, but it shows text as soon as TTFT is over; OpenAI’s latency guide calls it the single most effective way to make users wait less.
Example
Two 600-token replies that feel different
Take two made-up replies of 600 tokens. Reply A starts after 0.4 s and finishes at 8.0 s. Reply B thinks first, starts at 6.0 s and finishes at 9.0 s. Applying NVIDIA’s formulas:
Once it starts, B streams about 2.5 times as fast, yet A feels far quicker in a chat because text appears almost at once. For a background job that only uses the finished result, A’s total of 8.0 s against B’s 9.0 s is what counts. Our Gemini playground shows the first-token and total time under every reply.
| Metric | Reply A | Reply B |
|---|---|---|
| Time to first token | 0.4 s | 6.0 s |
| End-to-end latency | 8.0 s | 9.0 s |
| Inter-token latency | 12.7 ms | 5.0 ms |
| Speed after the first token (1 ÷ inter-token latency) | 79 tokens/s | 200 tokens/s |
| Tokens per second per user | 75 tokens/s | 67 tokens/s |
The timings are invented for the illustration; the formulas are from NVIDIA’s benchmarking metrics.
Cost and quality
Why it matters
TTFT decides how responsive a chat or coding assistant feels. End-to-end latency decides batch jobs and agents that chain many calls. Streaming cuts the wait users feel. To cut TTFT itself, choose lower reasoning effort when the task allows, and keep long prompts short or cache a repeated prefix; Anthropic says caching reduces processing time as well as cost.
To cut total time, generate fewer tokens. OpenAI’s latency guide estimates that halving output tokens may cut about half the latency, while halving the prompt may only gain 1–5%.
Don’t mix up
Common confusions
- TTFT vs end-to-end latency
- TTFT ends when the first token arrives; end-to-end latency ends with the last. A fast start followed by a long reply can still take a long time to finish.
- Per-user vs system tokens per second
- Benchmarks often quote total throughput across many simultaneous requests. NVIDIA notes that as concurrency rises, system throughput goes up while each user’s throughput goes down.
Go deeper
Try it and read more
- Free toolGemini API PlaygroundTry the Gemini API free with your own key.
- Free toolPrompt Caching CalculatorEstimate savings from prompt caching.
- Guide · 11 min readPrompt caching explained: how it works on OpenAI, Anthropic and Gemini, and when it paysHow prompt caching works on OpenAI, Anthropic and Gemini: what gets cached, minimum lengths, how long caches last, write and read prices, and when it pays off.
Related
Related terms
- Prompt cachingPrompt caching is an API feature that stores the processed start of a prompt, so later requests that begin with exactly the same tokens are billed at a much lower cached-input price and start answering sooner.
- Reasoning tokensReasoning tokens are the tokens a reasoning model generates while it works through a problem before writing its answer, billed as output tokens even though the API hides them or returns only a summary.
- KV cacheThe KV cache is the memory where a language model keeps the attention keys and values of every token it has already processed, so each new token is computed without reprocessing the whole sequence.
- Context windowA context window is the maximum number of tokens a language model can work with in one request, counting the system prompt, tool definitions, conversation history, documents and the reply it writes.