During a recent interview, I was asked a question that sounded straightforward:
What happens from the moment you send a prompt to an LLM until you get the response?
I answered with the concepts I knew: tokenization, embeddings, transformer processing, and token generation. But as the discussion continued, the interviewer connected the question to something more practical: latency in a retrieval-augmented generation application.
If a RAG application takes several seconds to answer, where exactly are those seconds going?
I couldn't explain the complete picture as clearly as I wanted to. After the interview, I explored how the individual pieces fit together. This article is the explanation I wish I had been able to give: the journey from a user's question to retrieval, model processing, the first visible token, and the completed answer.
The central lesson is simple: before optimizing a slow RAG application, identify which part of the request is slow.
1. A RAG request is a pipeline
At a high level, an ordinary LLM interaction looks like this:
Prompt → LLM → ResponseA typical RAG interaction has more stages:
User question
↓
Request handling and optional query preparation
↓
Query embedding
↓
Retrieval and document fetching
↓
Optional reranking
↓
Context selection and prompt construction
↓
LLM request, queueing, and prefill
↓
Token generation
↓
Streaming, optional validation, and displayThis is a reference architecture, not a requirement. Keyword retrieval may run without an embedding model. Some applications skip reranking. Others add query rewriting, multiple retrievers, tool calls, or repeated retrieval rounds.
Every additional stage can add computation, waiting, or a network round trip. The important question is which stages sit on the path the user must wait for.
2. Separate ingestion from answering
Before diagnosing request latency, distinguish the work done when documents enter the system from the work done when a user asks a question.
An ingestion pipeline usually reads documents, extracts text, splits it into chunks, creates document embeddings, and stores those embeddings alongside text or references to it.
Documents → Parsing → Chunking → Document embeddings → IndexThat work normally happens ahead of time. A user request should not need to embed the entire document collection again.
At query time, dense retrieval typically embeds the user's question and searches the existing index:
Question → Query embedding → Search existing document vectorsIf a system parses uploaded files or builds an index inside the same request, that work belongs in its latency budget too. But it is a different architecture from querying an already prepared knowledge base.
There is also an important distinction between retrieval embeddings and the LLM's internal token embeddings. A retrieval model produces vectors used to find relevant documents. Inside a text-generating transformer, token IDs are mapped to learned representations for model processing. These are different operations; ordinary text generation does not require a separate vector-database embedding API call.
3. Define what “slow” means
There are several useful latency measurements, and they answer different questions.
| Metric | What it measures | What it helps diagnose |
|---|---|---|
| Application time to first token | Time from submitting the question to seeing the first answer text | The initial user wait across the whole pipeline |
| LLM endpoint time to first token | Time from sending the generation request to receiving its first output token | Provider transport, queueing, and initial model processing |
| Inter-token latency | Time between successive generated tokens | How smoothly the answer arrives |
| Time to final answer | Time from submitting the question to receiving the completed answer | Overall completion time |
| Time to useful answer | Time until the user has enough information to act | Whether the response is useful early enough |
Define measurement boundaries explicitly. A first stream event containing only metadata is not the first visible answer token. Also distinguish when your backend receives text from when the browser renders it.
For a simple sequential pipeline, an approximate accounting model is:
Application TTFT ≈
request preparation
+ query embedding
+ retrieval and document fetching
+ reranking
+ prompt construction
+ LLM endpoint TTFT
+ delivery and rendering delay
Time to final answer ≈
application TTFT
+ generation after the first visible token
+ final processing and deliveryEach stage measurement should include its own waiting and transport costs. Avoid adding network time again if it is already included in a measured API duration.
These equations describe a sequential request. When stages overlap, the elapsed time follows their dependencies rather than the sum of every span.
4. Query preparation can hide expensive LLM calls
Consider a follow-up question:
Does it support offline use?
Without conversation history, a retriever may not know what “it” refers to. Rewriting the question into a standalone query can improve retrieval.
But rewriting may require another model call. So can intent classification, choosing a retriever, or deciding whether retrieval is needed.
Imagine this sequence:
LLM classification
→ LLM query rewriting
→ LLM retrieval routing
→ Retrieval
→ LLM answer generationThe user waits for all four model calls because each stage depends on the previous one.
For each intermediate call, ask whether the application needs it for every question. Could a simple rule handle an obvious case? Could related decisions share one call? Could a smaller model perform the task adequately? Would the original query retrieve useful evidence without rewriting?
These are experimental choices. Removing a rewrite step that resolves conversational ambiguity can make the system faster while making its answers worse.
5. Query embeddings have their own latency
An embedding request can include tokenization, inference, transport, service queueing, and potentially a cold start. A short question does not guarantee a short API round trip.
Measure the complete embedding call. If it becomes a bottleneck, investigate service placement, connection reuse, deployment readiness, and whether repeated exact queries can reuse embeddings.
An embedding cache key should identify the text and the embedding configuration, including the model version. Reusing a vector after changing models can produce incompatible retrieval results.
Changing embedding models is also an index decision. Query and document vectors need compatible representations; switching only the query model is generally insufficient.
6. Retrieval includes more than vector search
A database search may be fast while the retrieval stage is slow.
For example, an application might search an index for chunk IDs, fetch text from another database, load source metadata, check access permissions, and merge keyword and dense search results before it has usable context.
Measure these operations separately. Otherwise, “retrieval took 800 milliseconds” says little about what needs fixing.
There are also two different limits to tune:
- Candidate count: how many chunks the retriever returns for consideration.
- Final context count or token budget: how much evidence reaches the answering model.
Retrieving 30 candidates and passing 5 selected chunks is different from putting all 30 into the prompt.
Metadata filters can restrict results to the relevant product, document version, or tenant. That can improve relevance, but filtering performance depends on the index and filter strategy. Do not assume every filter makes a query faster. Access restrictions must remain correct regardless of their performance effect.
Chunking creates another tradeoff. Large chunks may contain useful surrounding explanations alongside irrelevant material. Small chunks may retrieve precisely but split an answer across several pieces. Tune chunk boundaries against representative questions, not just a preferred chunk size.
7. Reranking adds work but may reduce work later
A common reranking setup scores query-document pairs after initial retrieval and reorders the candidates by relevance. This second stage adds latency, and its cost depends partly on candidate count and document length. Pinecone's reranking guide describes these tradeoffs.
The useful question is whether that extra work earns its place in the complete pipeline.
Suppose initial retrieval finds many plausible chunks. A reranker might let you send a smaller, more relevant set to the LLM. Its added time could be offset by less model input processing, or justified by better answers. That is a hypothesis to test, not an automatic speedup.
Keep the reranker's input count separate from its output count. Asking it to return only three results does not necessarily mean it scored only three candidates.
Compare retrieval quality, answer quality, application TTFT, and completion time with and without reranking. A comparison of the reranking span alone misses its downstream effects.
8. Context construction connects retrieval to generation
After retrieval, the application builds the actual model input. This may include system instructions, conversation history, the question, selected evidence, source identifiers, and formatting requirements.
The relevant quantity is the final prompt's token count, not the number of retrieved chunks.
Five chunks of 150 tokens are very different from five chunks of 1,500 tokens. Repeated headers, overlapping passages, HTML markup, and unnecessary metadata can also consume space without adding useful evidence.
A context budget gives selection a concrete constraint:
Select relevant evidence within a token budget
→ retain source identifiers
→ remove duplicate passages
→ preserve the information needed to answerCompression is not free. An extra LLM summarization call adds latency and can omit decisive details. Try deterministic cleanup and better evidence selection first, then evaluate whether model-based compression improves the overall result.
9. Inside the LLM: tokenization, prefill, and decoding
For a typical autoregressive text-generating transformer, the input is tokenized into IDs and processed by the model. It helps to separate two inference phases.
Prefill: processing the input
During prefill, the model processes the prompt and establishes the state used to generate the continuation. Prompt positions can be processed in parallel within this phase, subject to hardware and serving constraints.
Longer uncached prompts generally require more input processing. However, input length alone does not predict TTFT: queueing, hardware, batching, cache reuse, and serving implementation also matter.
Decoding: producing the continuation
In ordinary autoregressive decoding, the model predicts a next-token distribution, selects a token, and repeats using the expanded sequence. Generation continues until a stopping condition or output limit is reached.
The KV cache retains earlier attention keys and values so the model can reuse them during subsequent steps. It avoids recomputing that state, but consumes memory that grows with the cached sequence. Hugging Face's cache documentation explains the mechanism and memory tradeoffs.
For a simplified illustration, if 299 tokens remain after the first token and the effective rate is 50 tokens per second, the remaining generation takes about 6 seconds. Real rates vary throughout a request and across workloads.
This explains why a response can start quickly but finish slowly. Long prompts and long answers affect different parts of the wait, and both need measurement.
Some reasoning models also perform internal generation before exposing answer text. Visible answer length may therefore be an incomplete explanation of their latency.
10. More context does not imply a proportional slowdown
The chain from the LinkedIn post is useful:
More retrieved text
→ more input tokens
→ potentially more prefill work
→ potentially higher TTFTThe word “potentially” matters. A cached prefix, a lightly loaded endpoint, or efficient prompt processing can change the observed effect. For moderate prompts, output generation or extra serial requests may dominate instead.
OpenAI's latency optimization guide discusses reducing output, avoiding unnecessary requests, parallelizing independent work, and trimming large inputs. Those are useful categories for investigation, rather than promises of a particular speedup.
Test different context budgets on the same evaluation questions. Record input tokens, output tokens, cache usage where available, latency, and answer quality. Do not infer the benefit from token reduction alone.
11. Cache the right layer
“Add caching” is incomplete advice because each cache avoids different work.
| Cache | Reuses | Work it may avoid | Important invalidation boundary |
|---|---|---|---|
| Query embedding cache | A vector for an identical query and configuration | Query embedding computation | Embedding model and preprocessing configuration |
| Retrieval cache | Retrieved candidates or selected evidence | Search and some document fetching | Index version, filters, and authorization scope |
| Answer cache | A previously generated response | Much of the answering pipeline | Evidence freshness, user context, and authorization |
| Prompt prefix cache | Previously processed matching prompt prefix | Some model input processing | Provider rules and prefix identity |
For query embeddings, the invalidation boundary is the embedding model and preprocessing configuration. An answer cache also needs to distinguish conversation state and personalized responses.
Semantic answer caching can reuse an answer for a similar question, but similarity is not equivalence. “Can I cancel?” and “Can I cancel after renewal?” may require different evidence and answers.
Prompt caching is a different mechanism. Matching prefixes can reuse prior input processing; they do not replace retrieval or reuse the complete answer. Stable instructions before dynamic evidence can help prefix reuse when supported. Eligibility, retention, and cache behavior vary by provider. OpenAI's prompt caching documentation describes its implementation.
Measure cold requests and cache hits separately, then report their real production mixture. An all-hit benchmark does not describe first-time users or newly updated documents.
12. Parallelize independent work
Hybrid retrieval may combine dense and keyword searches. If keyword search needs only the prepared query while dense search needs its embedding, a possible dependency graph is:
Prepared query
├── Keyword search ──────────────────┐
└── Query embedding → Dense search ─┤
↓
Merge results
↓
Rerank
↓
AnswerIgnoring overhead, the retrieval branches complete in approximately:
max(keyword search time, embedding time + dense search time)Running them sequentially would add their times instead.
Parallelism helps only where dependencies permit it. Dense search must wait for its vector, and reranking must wait for its candidates. It can also increase pressure on shared services. Benchmark under realistic concurrency, not just one isolated request.
Multi-query retrieval has the same tradeoff: searching several query variants in parallel may improve recall, but generating the variants still adds work, and the expanded candidate pool may increase reranking cost.
13. Streaming changes the user's wait
Streaming lets the user read the answer while generation continues. It can substantially improve perceived responsiveness without reducing the model's total generation work.
But it does not remove the retrieval, reranking, and initial model processing that must happen before answer text is available.
The full delivery path matters too. A backend can receive tokens promptly while an intermediary buffers them, or the frontend delays rendering. Measure both backend receipt and client display.
If validation must finish before any answer text is shown, account for that gate in user-visible latency. A status message such as “Searching documents” is helpful feedback, but it should not be counted as the first answer token.
14. A worked latency budget
The following numbers are invented for illustration, not benchmarks or results from my application.
Assume one sequential request with no overlapping stages:
| Stage | Duration |
|---|---|
| Request preparation | 40 ms |
| LLM query rewrite | 450 ms |
| Query embedding | 100 ms |
| Search and document fetching | 160 ms |
| Reranking | 250 ms |
| Context construction | 20 ms |
| LLM endpoint TTFT | 900 ms |
| First-token delivery and rendering | 30 ms |
| Generation after the first token | 5,000 ms |
| Final processing and delivery | 50 ms |
Application TTFT is 1,950 ms. Completion time is 7,000 ms.
Now consider several independent hypothetical improvements:
- Halving search and document-fetch time saves 80 ms from both measurements.
- Removing the rewrite on questions that do not need it saves 450 ms.
- Reducing the endpoint TTFT by 300 ms saves 300 ms.
- Reducing remaining generation from 5 seconds to 3 seconds saves 2 seconds from completion time, while leaving TTFT unchanged.
These calculations expose which changes address which symptom. They do not establish that a rewrite is dispensable or that a shorter answer is equally useful; evaluation must establish that separately.
15. Instrument the request before changing it
Give each request a trace ID and record spans for query preparation, embedding, search, document fetching, reranking, context construction, the generation request, first output text, final output, and client rendering.
Record enough context to interpret the timing: model and configuration, candidate count, selected context tokens, output tokens, cache hits, retries, and concurrent load. Use monotonic clocks for elapsed durations and distributed tracing to connect services.
An external model API may expose only request-level timing. In that case, endpoint TTFT is measurable, but the split between provider queueing and prefill may be unknown. Do not label the entire interval “prefill” without supporting telemetry.
Track p50, p95, and p99 latency. A median can look healthy while a smaller group of users experiences long waits from rate limits, retries, cold starts, or overloaded services.
Percentiles cannot simply be added across stages: the request at one stage's p95 may not be the request at another stage's p95. Inspect traces of slow end-to-end requests to understand their actual paths.
16. Optimize against quality as well as latency
A fast unsupported answer does not solve the engineering problem.
Build an evaluation set with factual questions, questions requiring several documents, follow-ups, exact identifiers, unanswerable questions, and document-version differences.
Compare retrieval recall, evidence relevance, grounded answer correctness, citation accuracy, and appropriate abstention alongside latency and cost. For multi-document questions, a low context limit may remove one of the passages needed to answer correctly.
A practical optimization loop is:
- Establish baseline quality and latency under representative load.
- Inspect the stages of slow requests and identify the dominant wait.
- Change one factor, such as rewrite routing, candidate count, context budget, or model choice.
- Repeat the same quality evaluation and load measurement.
- Keep the change only if its quality, latency, and cost tradeoff meets the application's requirements.
Also distinguish isolated speed from capacity. A deployment may process more requests per second while making individual requests wait longer. Choose the balance based on the user experience you need.
17. The answer I would give in the interview now
In a RAG application, the question first passes through request handling and any query preparation. For dense retrieval, an embedding model converts the query into a vector used to search an existing document index. The application fetches relevant text, optionally reranks it, and builds a prompt from the question, instructions, history, and selected evidence.
The generation service tokenizes and processes that input, then generates the continuation. Prefill handles the prompt; autoregressive decoding produces the answer while reusing attention state through a KV cache. The application receives and displays the output, potentially as a stream.
To diagnose latency, I would measure the full request and its stages, distinguish application TTFT from generation endpoint TTFT and completion time, and inspect slow requests under realistic load. Then I would optimize the dominant stage while checking answer quality.
That interview pushed me to connect ideas I previously understood separately. Retrieval decisions affect prompt size. Prompt size affects model processing. Serial calls affect the wait before generation begins. Output length affects completion time. Caching helps different parts of the system depending on what is reused.
The question I would start with now is: where is the user actually waiting, and what evidence tells me why?