Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model’s context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Prompt Caching CalculatorEstimate savings from prompt caching.Tool

AI glossary · Models and context

What is time to first token (TTFT)?

Also called: TTFT, first-token latency

Definition

Time to first token (TTFT) is how long a model takes from receiving a request to returning the first token of its reply, the wait a user feels before streamed text starts to appear.

Explained

How it works

NVIDIA’s benchmarking docs define TTFT as the time from submitting a query to receiving the first token, and note it generally includes queueing, prefill (processing the whole prompt) and network time, so longer prompts raise it. End-to-end latency is TTFT plus the time to generate every remaining token.

Three related numbers get mixed up. Inter-token latency is the average gap between tokens after the first. Tokens per second per user is the reply length divided by end-to-end latency. Total tokens per second for a system adds up every request being served at once.

On reasoning models the thinking happens before the first visible token, so TTFT can stretch to many seconds. Streaming doesn’t make the model faster, but it shows text as soon as TTFT is over; OpenAI’s latency guide calls it the single most effective way to make users wait less.

Example

Two 600-token replies that feel different

Take two made-up replies of 600 tokens. Reply A starts after 0.4 s and finishes at 8.0 s. Reply B thinks first, starts at 6.0 s and finishes at 9.0 s. Applying NVIDIA’s formulas:

Once it starts, B streams about 2.5 times as fast, yet A feels far quicker in a chat because text appears almost at once. For a background job that only uses the finished result, A’s total of 8.0 s against B’s 9.0 s is what counts. Our Gemini playground shows the first-token and total time under every reply.

Same length, different shape
MetricReply AReply B
Time to first token0.4 s6.0 s
End-to-end latency8.0 s9.0 s
Inter-token latency12.7 ms5.0 ms
Speed after the first token (1 ÷ inter-token latency)79 tokens/s200 tokens/s
Tokens per second per user75 tokens/s67 tokens/s

The timings are invented for the illustration; the formulas are from NVIDIA’s benchmarking metrics.

Cost and quality

Why it matters

TTFT decides how responsive a chat or coding assistant feels. End-to-end latency decides batch jobs and agents that chain many calls. Streaming cuts the wait users feel. To cut TTFT itself, choose lower reasoning effort when the task allows, and keep long prompts short or cache a repeated prefix; Anthropic says caching reduces processing time as well as cost.

To cut total time, generate fewer tokens. OpenAI’s latency guide estimates that halving output tokens may cut about half the latency, while halving the prompt may only gain 1–5%.

Don’t mix up

Common confusions

TTFT vs end-to-end latency
TTFT ends when the first token arrives; end-to-end latency ends with the last. A fast start followed by a long reply can still take a long time to finish.
Per-user vs system tokens per second
Benchmarks often quote total throughput across many simultaneous requests. NVIDIA notes that as concurrency rises, system throughput goes up while each user’s throughput goes down.

Go deeper

Try it and read more

Related

All 40 terms in the AI glossary

Written by Tahir Nazir. Checked .

How this was checked: Metric definitions and formulas checked against NVIDIA’s NIM benchmarking docs (version 2.0.0), streaming and output-length guidance against OpenAI’s latency guide, and caching latency against Anthropic’s prompt caching docs on 2026-10-11. The example timings are invented; the derived figures are computed on this page. The playground readout was checked in our own code.