← All posts
RAGLLMGenerative AIAI EngineeringLatency
Created
Updated

RAG Application Latency: What Happens Between a Question and the Final Answer?

Follow a RAG request from query embedding and retrieval to LLM prefill, token generation, and streaming. Learn how to measure latency and optimize it without sacrificing answer quality.

During a recent interview, I was asked a question that sounded straightforward:

What happens from the moment you send a prompt to an LLM until you get the response?

I answered with the concepts I knew: tokenization, embeddings, transformer processing, and token generation. But as the discussion continued, the interviewer connected the question to something more practical: latency in a retrieval-augmented generation application.

If a RAG application takes several seconds to answer, where exactly are those seconds going?

I couldn't explain the complete picture as clearly as I wanted to. After the interview, I explored how the individual pieces fit together. This article is the explanation I wish I had been able to give: the journey from a user's question to retrieval, model processing, the first visible token, and the completed answer.

The central lesson is simple: before optimizing a slow RAG application, identify which part of the request is slow.

1. A RAG request is a pipeline

At a high level, an ordinary LLM interaction looks like this:

text
Prompt → LLM → Response

A typical RAG interaction has more stages:

text
User question
    ↓
Request handling and optional query preparation
    ↓
Query embedding
    ↓
Retrieval and document fetching
    ↓
Optional reranking
    ↓
Context selection and prompt construction
    ↓
LLM request, queueing, and prefill
    ↓
Token generation
    ↓
Streaming, optional validation, and display

This is a reference architecture, not a requirement. Keyword retrieval may run without an embedding model. Some applications skip reranking. Others add query rewriting, multiple retrievers, tool calls, or repeated retrieval rounds.

Every additional stage can add computation, waiting, or a network round trip. The important question is which stages sit on the path the user must wait for.

2. Separate ingestion from answering

Before diagnosing request latency, distinguish the work done when documents enter the system from the work done when a user asks a question.

An ingestion pipeline usually reads documents, extracts text, splits it into chunks, creates document embeddings, and stores those embeddings alongside text or references to it.

text
Documents → Parsing → Chunking → Document embeddings → Index

That work normally happens ahead of time. A user request should not need to embed the entire document collection again.

At query time, dense retrieval typically embeds the user's question and searches the existing index:

text
Question → Query embedding → Search existing document vectors

If a system parses uploaded files or builds an index inside the same request, that work belongs in its latency budget too. But it is a different architecture from querying an already prepared knowledge base.

There is also an important distinction between retrieval embeddings and the LLM's internal token embeddings. A retrieval model produces vectors used to find relevant documents. Inside a text-generating transformer, token IDs are mapped to learned representations for model processing. These are different operations; ordinary text generation does not require a separate vector-database embedding API call.

3. Define what “slow” means

There are several useful latency measurements, and they answer different questions.

MetricWhat it measuresWhat it helps diagnose
Application time to first tokenTime from submitting the question to seeing the first answer textThe initial user wait across the whole pipeline
LLM endpoint time to first tokenTime from sending the generation request to receiving its first output tokenProvider transport, queueing, and initial model processing
Inter-token latencyTime between successive generated tokensHow smoothly the answer arrives
Time to final answerTime from submitting the question to receiving the completed answerOverall completion time
Time to useful answerTime until the user has enough information to actWhether the response is useful early enough

Define measurement boundaries explicitly. A first stream event containing only metadata is not the first visible answer token. Also distinguish when your backend receives text from when the browser renders it.

For a simple sequential pipeline, an approximate accounting model is:

text
Application TTFT ≈
    request preparation
  + query embedding
  + retrieval and document fetching
  + reranking
  + prompt construction
  + LLM endpoint TTFT
  + delivery and rendering delay
 
Time to final answer ≈
    application TTFT
  + generation after the first visible token
  + final processing and delivery

Each stage measurement should include its own waiting and transport costs. Avoid adding network time again if it is already included in a measured API duration.

These equations describe a sequential request. When stages overlap, the elapsed time follows their dependencies rather than the sum of every span.

4. Query preparation can hide expensive LLM calls

Consider a follow-up question:

Does it support offline use?

Without conversation history, a retriever may not know what “it” refers to. Rewriting the question into a standalone query can improve retrieval.

But rewriting may require another model call. So can intent classification, choosing a retriever, or deciding whether retrieval is needed.

Imagine this sequence:

text
LLM classification
    → LLM query rewriting
    → LLM retrieval routing
    → Retrieval
    → LLM answer generation

The user waits for all four model calls because each stage depends on the previous one.

For each intermediate call, ask whether the application needs it for every question. Could a simple rule handle an obvious case? Could related decisions share one call? Could a smaller model perform the task adequately? Would the original query retrieve useful evidence without rewriting?

These are experimental choices. Removing a rewrite step that resolves conversational ambiguity can make the system faster while making its answers worse.

5. Query embeddings have their own latency

An embedding request can include tokenization, inference, transport, service queueing, and potentially a cold start. A short question does not guarantee a short API round trip.

Measure the complete embedding call. If it becomes a bottleneck, investigate service placement, connection reuse, deployment readiness, and whether repeated exact queries can reuse embeddings.

An embedding cache key should identify the text and the embedding configuration, including the model version. Reusing a vector after changing models can produce incompatible retrieval results.

Changing embedding models is also an index decision. Query and document vectors need compatible representations; switching only the query model is generally insufficient.

6. Retrieval includes more than vector search

A database search may be fast while the retrieval stage is slow.

For example, an application might search an index for chunk IDs, fetch text from another database, load source metadata, check access permissions, and merge keyword and dense search results before it has usable context.

Measure these operations separately. Otherwise, “retrieval took 800 milliseconds” says little about what needs fixing.

There are also two different limits to tune:

Retrieving 30 candidates and passing 5 selected chunks is different from putting all 30 into the prompt.

Metadata filters can restrict results to the relevant product, document version, or tenant. That can improve relevance, but filtering performance depends on the index and filter strategy. Do not assume every filter makes a query faster. Access restrictions must remain correct regardless of their performance effect.

Chunking creates another tradeoff. Large chunks may contain useful surrounding explanations alongside irrelevant material. Small chunks may retrieve precisely but split an answer across several pieces. Tune chunk boundaries against representative questions, not just a preferred chunk size.

7. Reranking adds work but may reduce work later

A common reranking setup scores query-document pairs after initial retrieval and reorders the candidates by relevance. This second stage adds latency, and its cost depends partly on candidate count and document length. Pinecone's reranking guide describes these tradeoffs.

The useful question is whether that extra work earns its place in the complete pipeline.

Suppose initial retrieval finds many plausible chunks. A reranker might let you send a smaller, more relevant set to the LLM. Its added time could be offset by less model input processing, or justified by better answers. That is a hypothesis to test, not an automatic speedup.

Keep the reranker's input count separate from its output count. Asking it to return only three results does not necessarily mean it scored only three candidates.

Compare retrieval quality, answer quality, application TTFT, and completion time with and without reranking. A comparison of the reranking span alone misses its downstream effects.

8. Context construction connects retrieval to generation

After retrieval, the application builds the actual model input. This may include system instructions, conversation history, the question, selected evidence, source identifiers, and formatting requirements.

The relevant quantity is the final prompt's token count, not the number of retrieved chunks.

Five chunks of 150 tokens are very different from five chunks of 1,500 tokens. Repeated headers, overlapping passages, HTML markup, and unnecessary metadata can also consume space without adding useful evidence.

A context budget gives selection a concrete constraint:

text
Select relevant evidence within a token budget
    → retain source identifiers
    → remove duplicate passages
    → preserve the information needed to answer

Compression is not free. An extra LLM summarization call adds latency and can omit decisive details. Try deterministic cleanup and better evidence selection first, then evaluate whether model-based compression improves the overall result.

9. Inside the LLM: tokenization, prefill, and decoding

For a typical autoregressive text-generating transformer, the input is tokenized into IDs and processed by the model. It helps to separate two inference phases.

Prefill: processing the input

During prefill, the model processes the prompt and establishes the state used to generate the continuation. Prompt positions can be processed in parallel within this phase, subject to hardware and serving constraints.

Longer uncached prompts generally require more input processing. However, input length alone does not predict TTFT: queueing, hardware, batching, cache reuse, and serving implementation also matter.

Decoding: producing the continuation

In ordinary autoregressive decoding, the model predicts a next-token distribution, selects a token, and repeats using the expanded sequence. Generation continues until a stopping condition or output limit is reached.

The KV cache retains earlier attention keys and values so the model can reuse them during subsequent steps. It avoids recomputing that state, but consumes memory that grows with the cached sequence. Hugging Face's cache documentation explains the mechanism and memory tradeoffs.

For a simplified illustration, if 299 tokens remain after the first token and the effective rate is 50 tokens per second, the remaining generation takes about 6 seconds. Real rates vary throughout a request and across workloads.

This explains why a response can start quickly but finish slowly. Long prompts and long answers affect different parts of the wait, and both need measurement.

Some reasoning models also perform internal generation before exposing answer text. Visible answer length may therefore be an incomplete explanation of their latency.

10. More context does not imply a proportional slowdown

The chain from the LinkedIn post is useful:

text
More retrieved text
    → more input tokens
    → potentially more prefill work
    → potentially higher TTFT

The word “potentially” matters. A cached prefix, a lightly loaded endpoint, or efficient prompt processing can change the observed effect. For moderate prompts, output generation or extra serial requests may dominate instead.

OpenAI's latency optimization guide discusses reducing output, avoiding unnecessary requests, parallelizing independent work, and trimming large inputs. Those are useful categories for investigation, rather than promises of a particular speedup.

Test different context budgets on the same evaluation questions. Record input tokens, output tokens, cache usage where available, latency, and answer quality. Do not infer the benefit from token reduction alone.

11. Cache the right layer

“Add caching” is incomplete advice because each cache avoids different work.

CacheReusesWork it may avoidImportant invalidation boundary
Query embedding cacheA vector for an identical query and configurationQuery embedding computationEmbedding model and preprocessing configuration
Retrieval cacheRetrieved candidates or selected evidenceSearch and some document fetchingIndex version, filters, and authorization scope
Answer cacheA previously generated responseMuch of the answering pipelineEvidence freshness, user context, and authorization
Prompt prefix cachePreviously processed matching prompt prefixSome model input processingProvider rules and prefix identity

For query embeddings, the invalidation boundary is the embedding model and preprocessing configuration. An answer cache also needs to distinguish conversation state and personalized responses.

Semantic answer caching can reuse an answer for a similar question, but similarity is not equivalence. “Can I cancel?” and “Can I cancel after renewal?” may require different evidence and answers.

Prompt caching is a different mechanism. Matching prefixes can reuse prior input processing; they do not replace retrieval or reuse the complete answer. Stable instructions before dynamic evidence can help prefix reuse when supported. Eligibility, retention, and cache behavior vary by provider. OpenAI's prompt caching documentation describes its implementation.

Measure cold requests and cache hits separately, then report their real production mixture. An all-hit benchmark does not describe first-time users or newly updated documents.

12. Parallelize independent work

Hybrid retrieval may combine dense and keyword searches. If keyword search needs only the prepared query while dense search needs its embedding, a possible dependency graph is:

text
Prepared query
    ├── Keyword search ──────────────────┐
    └── Query embedding → Dense search ─┤
                                        ↓
                                  Merge results
                                        ↓
                                     Rerank
                                        ↓
                                     Answer

Ignoring overhead, the retrieval branches complete in approximately:

text
max(keyword search time, embedding time + dense search time)

Running them sequentially would add their times instead.

Parallelism helps only where dependencies permit it. Dense search must wait for its vector, and reranking must wait for its candidates. It can also increase pressure on shared services. Benchmark under realistic concurrency, not just one isolated request.

Multi-query retrieval has the same tradeoff: searching several query variants in parallel may improve recall, but generating the variants still adds work, and the expanded candidate pool may increase reranking cost.

13. Streaming changes the user's wait

Streaming lets the user read the answer while generation continues. It can substantially improve perceived responsiveness without reducing the model's total generation work.

But it does not remove the retrieval, reranking, and initial model processing that must happen before answer text is available.

The full delivery path matters too. A backend can receive tokens promptly while an intermediary buffers them, or the frontend delays rendering. Measure both backend receipt and client display.

If validation must finish before any answer text is shown, account for that gate in user-visible latency. A status message such as “Searching documents” is helpful feedback, but it should not be counted as the first answer token.

14. A worked latency budget

The following numbers are invented for illustration, not benchmarks or results from my application.

Assume one sequential request with no overlapping stages:

StageDuration
Request preparation40 ms
LLM query rewrite450 ms
Query embedding100 ms
Search and document fetching160 ms
Reranking250 ms
Context construction20 ms
LLM endpoint TTFT900 ms
First-token delivery and rendering30 ms
Generation after the first token5,000 ms
Final processing and delivery50 ms

Application TTFT is 1,950 ms. Completion time is 7,000 ms.

Now consider several independent hypothetical improvements:

These calculations expose which changes address which symptom. They do not establish that a rewrite is dispensable or that a shorter answer is equally useful; evaluation must establish that separately.

15. Instrument the request before changing it

Give each request a trace ID and record spans for query preparation, embedding, search, document fetching, reranking, context construction, the generation request, first output text, final output, and client rendering.

Record enough context to interpret the timing: model and configuration, candidate count, selected context tokens, output tokens, cache hits, retries, and concurrent load. Use monotonic clocks for elapsed durations and distributed tracing to connect services.

An external model API may expose only request-level timing. In that case, endpoint TTFT is measurable, but the split between provider queueing and prefill may be unknown. Do not label the entire interval “prefill” without supporting telemetry.

Track p50, p95, and p99 latency. A median can look healthy while a smaller group of users experiences long waits from rate limits, retries, cold starts, or overloaded services.

Percentiles cannot simply be added across stages: the request at one stage's p95 may not be the request at another stage's p95. Inspect traces of slow end-to-end requests to understand their actual paths.

16. Optimize against quality as well as latency

A fast unsupported answer does not solve the engineering problem.

Build an evaluation set with factual questions, questions requiring several documents, follow-ups, exact identifiers, unanswerable questions, and document-version differences.

Compare retrieval recall, evidence relevance, grounded answer correctness, citation accuracy, and appropriate abstention alongside latency and cost. For multi-document questions, a low context limit may remove one of the passages needed to answer correctly.

A practical optimization loop is:

  1. Establish baseline quality and latency under representative load.
  2. Inspect the stages of slow requests and identify the dominant wait.
  3. Change one factor, such as rewrite routing, candidate count, context budget, or model choice.
  4. Repeat the same quality evaluation and load measurement.
  5. Keep the change only if its quality, latency, and cost tradeoff meets the application's requirements.

Also distinguish isolated speed from capacity. A deployment may process more requests per second while making individual requests wait longer. Choose the balance based on the user experience you need.

17. The answer I would give in the interview now

In a RAG application, the question first passes through request handling and any query preparation. For dense retrieval, an embedding model converts the query into a vector used to search an existing document index. The application fetches relevant text, optionally reranks it, and builds a prompt from the question, instructions, history, and selected evidence.

The generation service tokenizes and processes that input, then generates the continuation. Prefill handles the prompt; autoregressive decoding produces the answer while reusing attention state through a KV cache. The application receives and displays the output, potentially as a stream.

To diagnose latency, I would measure the full request and its stages, distinguish application TTFT from generation endpoint TTFT and completion time, and inspect slow requests under realistic load. Then I would optimize the dominant stage while checking answer quality.

That interview pushed me to connect ideas I previously understood separately. Retrieval decisions affect prompt size. Prompt size affects model processing. Serial calls affect the wait before generation begins. Output length affects completion time. Caching helps different parts of the system depending on what is reused.

The question I would start with now is: where is the user actually waiting, and what evidence tells me why?