Why Your AI Takes 1.5 Seconds to Start Talking and Then Speeds Up
Quick summary
- The first token waits on prompt processing, scheduling, and delivery. Later tokens reuse work already done.
- Long prompts can slow the start. Under load, queueing can be the bigger delay.
- Measure the request path, find the slow stage, and optimize that stage first.
Your AI takes 1.5 seconds to start talking. Then each new token arrives in about 30 milliseconds. That gap can look strange. It makes sense once you see that the model does two different jobs.
Those figures are an example, not a universal benchmark. Model size, prompt length, hardware, load, caching, and network distance all change the result.
How AI starts with two different jobs
Prefill: the model processes your prompt and builds its key-value (KV) attention state. This is one large pass across the input. More uncached context usually means more prefill work.
Decode: the model generates the answer one token at a time, reusing that state. Each step can be much faster. At scale, decode often runs into memory bandwidth limits, but the exact bottleneck depends on the model and serving setup.
So the first token pays for the prompt and the wait to be scheduled. The next tokens reuse the setup. The Sarathi-Serve paper explains why prefill and decode need different scheduling decisions.
And the model is only part of the wait. Time to first token (TTFT) can include:
- Application work, such as retrieval, tool calls, or authentication.
- Network travel between your app and the model server.
- Queueing while the server waits for capacity.
- Prompt processing, then generating and delivering the first token.
A long prompt can slow prefill. A busy server can make you wait before prefill even begins. Measure both.
How to reduce AI’s first-token wait
- Send less, more useful context.
- Remove stale chat history or summarize it when that keeps the needed details.
- In RAG, retrieve the passages that best answer the question. Do not assume top five is always better than top twenty; measure latency and answer quality together.
- Keep tool descriptions and instructions focused. Every unnecessary input token still has to be handled.
- Reuse a stable prompt prefix.
- If many requests share the same opening instructions, prompt caching may let the provider reuse their computed KV state.
- Keep stable content together at the start and changing content later. Cache rules vary by provider, so confirm hits in the usage metrics.
- A system prompt alone does not guarantee a cache hit. See OpenAI’s prompt caching guide for one provider’s rules.
- Fix the queue.
- If queue time dominates, trimming the prompt will not solve the main problem.
- Review capacity, admission limits, and request scheduling under real peak load.
- Continuous batching can keep the hardware busy, while chunked prefill can stop one long prompt from disrupting active generations. These policies trade off TTFT, throughput, and time between tokens.
- Cut avoidable network and application hops.
- Reuse connections and avoid unnecessary proxy layers.
- Serve users from a nearby region where practical.
- Run independent retrieval or checks in parallel only when their dependencies and authorization rules allow it.
- Stream as soon as the first token exists.
- Streaming does not make the model produce its first token sooner.
- It does let users see output immediately instead of waiting for the whole answer to finish.
- Show a useful progress state while retrieval or other prerequisite work runs.
For serving-side options, vLLM’s performance guide covers chunked prefill. Its disaggregated prefill guide explains how separate workers can tune prompt processing and token generation independently, with extra coordination costs.
Find the bottleneck before you tune
Track client-observed TTFT and, where available, split it into app time, network time, queue wait, and prefill. Compare p50 with p95. Separate cache hits from misses. Keep prompt length and answer quality beside the latency numbers.
Then fix the largest slice. If prefill dominates, trim or cache input. If queueing dominates, tune capacity and scheduling. If the network or app dominates, remove those delays.
That is the practical answer to the 1.5-second pause: do not guess at the model. Find out what the request is waiting for.