Skip to main content

How to Optimize LLM Inference: TTFT, Throughput, Continuous Batching, KV Cache and GPU Memory Explained

Calculating read time…

LLM inference optimization is the engineering process of making a deployed language model serve useful responses within explicit latency, throughput, memory, reliability, quality and cost constraints. It is not simply “making the GPU faster.” A production engineer must understand how requests move through queueing, prompt processing, token generation, batching and KV-cache allocation—and then determine which stage is actually limiting the service. ⚡

This matters because production traffic rarely looks like a clean benchmark. One user may send a short question while another submits a long RAG context and hundreds more arrive simultaneously. Optimizing only average latency, GPU utilization or requests per second can therefore improve one dashboard while making the user experience worse. The production goal is useful throughput while meeting latency and quality objectives. 🏗️

Original LLM inference optimization diagram showing requests entering a continuous batching scheduler, consuming GPU model and KV-cache memory, followed by TTFT, throughput and cache measurements in a feedback optimization loop

🔀 Quick Comparison

Metric What it tells you Typical pressure Why users care
TTFT Time from request arrival until the first generated token. Queueing and prompt/prefill processing. How quickly an interactive response appears to begin.
TPOT Request-level time per generated output token after the first token. Decode scheduling and GPU contention. How quickly the answer continues after it starts.
ITL Observed gap between streamed output events. Decode contention and scheduling behavior. Perceived smoothness of streaming.
Throughput Useful work completed over time. Batching, GPU compute, memory and workload mix. Determines how much demand the deployment can sustain.
KV usage How much serving cache capacity is occupied. Context lengths and active sequences. Can constrain concurrency before raw GPU compute is exhausted.

1. Foundations: What Are We Actually Optimizing?

🧒 Sticky-note analogy: Imagine a supermarket. Making one customer's checkout extremely fast is useful, but opening the store to 500 customers changes the problem. Now queue length, checkout capacity and how efficiently cashiers share work matter too.

LLM inference has the same distinction. Latency asks how long an individual request waits. Throughput asks how much useful inference work the system completes over time. Optimizing one can affect the other.

For an interactive LLM, one end-to-end latency number hides several different experiences. vLLM's current production metrics expose separate measurements for time to first token, request-level time per output token, inter-token latency, queue time, prefill time and decode time. That separation is operationally valuable because each metric points toward a different class of bottleneck.

TTFT—Time to First Token is the delay before the first generated token reaches the user. A large prompt can increase prompt-processing work; a saturated server can add queue time before that work even begins.

TPOT—Time Per Output Token describes request-level generation speed after the first token. vLLM currently calculates its request-level TPOT metric from the post-TTFT latency divided across subsequent output tokens.

ITL—Inter-Token Latency measures the wall-clock gap between streamed output events. vLLM explicitly notes that ITL and request-level TPOT are related but not always identical, particularly when an output event contains multiple tokens.

Throughput must also be defined carefully. Requests per second alone can be misleading when one request contains a tiny prompt and another contains thousands of tokens. Token-oriented measurements and workload distributions usually provide more useful context for LLM capacity engineering.

💡 Production rule: Never say “the optimized configuration is faster” without naming the workload and metric. A configuration can improve aggregate throughput while increasing TTFT, or reduce TTFT while lowering the number of simultaneous requests the service can sustain.

🎯 Use this when... you need to define a measurable performance objective before changing inference configuration.

2. Mechanics: Prefill, Decode and Why One Request Has Two Personalities

🧒 Sticky-note analogy: Reading a question and writing the answer are different jobs. You may read an entire page before writing the first word, but once you start answering, words appear one after another.

Autoregressive LLM inference can be understood as two important phases: prefill and decode. NVIDIA's current TensorRT-LLM documentation likewise separates inference into context/prefill and generation/decode phases because they have different compute characteristics.

Prefill: the model processes the input prompt. Attention state needed for later generation is computed and stored in the KV cache. Longer prompts generally mean more prompt-processing work and more state associated with the sequence.

Decode: after the prompt has been processed, generation proceeds autoregressively. Newly generated tokens extend the sequence and the cached attention state is reused rather than reconstructing the entire history from scratch for every token.

This distinction explains why an application can feel slow in two very different ways:

  1. Slow start: the user waits too long before seeing the first token. Investigate queueing, prompt lengths and prefill pressure.
  2. Slow generation: the answer begins quickly but streams slowly. Investigate decode performance, scheduling pressure and competing workloads.

✅ Practical example: Suppose Employee A sends “Summarize this policy” with a very large retrieved document, while Employee B sends “What is our support number?” with a tiny prompt. The first request may place much more pressure on prefill and cache resources even if both eventually generate a similarly short answer. Grouping them together as simply “two requests” hides the important difference.

This is why mature inference benchmarking records prompt-token and generation-token distributions instead of relying only on request counts. Current vLLM metrics include histograms for request prompt tokens and generated tokens alongside queue, prefill, decode and latency measurements.

🎯 Use this when... you need to determine whether a latency problem occurs before generation starts or while tokens are being generated.

3. Continuous Batching: Turning Concurrent Users into Efficient GPU Work

🧒 Sticky-note analogy: A bus does not need every passenger to travel to the same destination. People can board and leave at different stops while the bus continues moving. Continuous batching applies a similar idea to active inference sequences.

Traditional static batching is easy to imagine: collect a group of requests, execute the batch, wait for the entire batch to finish, then start another. That is awkward for generative workloads because requests can have dramatically different prompt and generation lengths.

A serving scheduler can instead continually decide which active sequences should participate in upcoming model-execution steps. Completed requests leave while eligible work can be admitted. This is the key mental model behind continuous or in-flight batching.

Why does this improve throughput? GPU execution becomes less dependent on one request occupying the entire serving path. The scheduler can combine useful work from concurrent sequences and attempt to keep expensive accelerator resources productive.

But batching introduces a fundamental trade-off. More concurrent work can increase aggregate throughput while also increasing competition for scheduling slots, compute and KV-cache capacity. Eventually requests begin waiting.

Current vLLM telemetry makes this visible through metrics such as vllm:num_requests_running, vllm:num_requests_waiting and request queue-time measurements.

Practical experiment:

Run Traffic pattern Observe
AOne request at a time.Baseline TTFT, TPOT and token throughput.
BModerate concurrent arrivals.Whether aggregate throughput improves and how latency distributions move.
CIncrease offered load until requests queue.Queue time, waiting requests, KV pressure and tail latency.

The transition from Run B to Run C is especially valuable. It helps reveal the system's saturation region: additional demand no longer produces proportional useful throughput and instead increasingly appears as queueing or latency.

🎯 Use this when... your GPU serves concurrent users and you need to find the useful throughput region before queueing dominates user experience.

4. KV Cache and GPU Memory: Why “The Model Fits” Is Not Enough

🧒 Sticky-note analogy: Imagine studying with an open notebook. Remembering earlier calculations saves you from solving them again, but every active problem needs notebook space. Eventually the desk fills even though the textbook itself never became larger.

The KV cache is one of the most important concepts in production LLM serving. Transformer attention repeatedly needs information derived from tokens that have already been processed. During autoregressive inference, retaining key/value state allows later decode steps to reuse that work.

A simplified GPU-memory mental model is:

GPU memory
|
+-- model weights
|
+-- KV-cache blocks
|    +-- active request A
|    +-- active request B
|    +-- active request C
|
+-- runtime / temporary workspace

The model weights may remain constant while active request state grows and shrinks. This is why a model successfully loading into GPU memory proves only one thing: the model can load under that configuration. It does not prove that the server can support the required context lengths and concurrency.

Current vLLM exposes vllm:kv_cache_usage_perc, where a value of 1 represents full cache usage. It also exposes preemption and waiting-request metrics that can help correlate memory pressure with scheduling behavior.

Beginner exercise — the GPU-memory detective:

  1. Start the server with one representative model.
  2. Send several short-prompt requests and observe KV-cache usage.
  3. Repeat with longer prompts while keeping concurrency similar.
  4. Increase concurrent active requests.
  5. Watch KV-cache usage, waiting requests, queue time and preemptions together.
  6. Record the point where latency begins deteriorating materially.

The important lesson is not a particular percentage. The lesson is learning to correlate workload shape → memory pressure → scheduling behavior → user-visible latency.

Prefix caching adds another dimension. When requests share previously computed prefixes, cached KV state may be reusable. vLLM currently exposes both prefix-cache query and prefix-cache hit counters, allowing teams to verify whether the intended workload actually benefits rather than assuming it does.

✅ Corporate-style example: Consider an internal assistant where every request begins with the same long corporate policy instructions and stable reference context. Prefix caching may make repeated prompt computation reusable. In contrast, if every request has a largely unique prefix, the expected reuse can be much smaller. Measure cache queries and hits against representative traffic before treating caching as an optimization win.

🎯 Use this when... GPU memory appears sufficient for model loading but concurrency, long contexts or queueing cause production instability.

5. Practical Optimization Lab: Baseline → Load → Diagnose → Improve

🧒 Sticky-note analogy: A doctor measures temperature before giving medicine. If you change five medicines before measuring again, you cannot tell which treatment helped. Performance engineering follows the same discipline.

Example type: hypothetical enterprise workload using documented vLLM metrics and benchmarking capabilities. Imagine an internal AI assistant serving four workloads:

Workload Typical shape Likely engineering concern
Short Q&AShort input and output.Interactive latency.
RAG questionRetrieved context creates a larger prompt.Prefill and KV-cache pressure.
Code generationPotentially longer output.Decode duration and active-sequence lifetime.
Agent stepRepeated instructions plus changing state.Prefix reuse, latency and repeated inference cost.

Stage 1 — build a representative workload. Do not optimize using only “Hello, model.” Preserve realistic prompt-length and output-length distributions, important request types, streaming behavior and arrival patterns.

Stage 2 — establish the baseline. Record the exact model revision, inference image/build, GPU type, replica topology and serving configuration. Then capture latency distributions, throughput, request queueing and KV-cache behavior.

vLLM provides a current benchmarking suite and CLI tooling for controlled performance and regression experiments. Its documentation also describes parameter sweeps, allowing multiple configurations to be evaluated systematically.

Illustrative command pattern:

# Example pattern only: validate flags against your installed release.
vllm bench serve \
  --backend vllm \
  --model <model-id> \
  --request-rate <test-rate> \
  --num-prompts <sample-count>

Stage 3 — raise load gradually. Do not jump directly from one request to an overload test. Increase offered load in controlled steps. For every step, record:

  • TTFT distribution;
  • TPOT or appropriate streaming latency;
  • end-to-end latency;
  • prompt and generated-token throughput;
  • running and waiting requests;
  • queue time;
  • KV-cache utilization;
  • preemptions and failures;
  • GPU telemetry from the surrounding infrastructure.

Stage 4 — classify the bottleneck.

Observation Possible interpretation to investigate Next experiment
TTFT rises with queueOffered load may be exceeding current serving capacity.Reduce arrival rate or add controlled serving capacity and compare.
Long prompts hurt TTFTPrefill workload may be dominating.Separate results by prompt-length bucket.
KV usage stays highActive sequence state may be constraining concurrency.Vary context/output length and concurrency independently.
Good TTFT, poor generationDecode performance or interference may dominate.Compare decode-heavy workloads and concurrency levels.

Stage 5 — change exactly one important variable. Examples include request concurrency, scheduler limits, model precision, allowed context, cache-related settings, parallelism topology or replica count. Repeat the same workload and compare distributions—not just averages.

Stage 6 — verify model quality. An infrastructure optimization can change numerical behavior or model configuration. If quantization or another model-level optimization is introduced, rerun the application's quality regression suite as well as the performance benchmark.

Stage 7 — repeat under failure conditions. Restart a replica, remove capacity, exercise long requests, create a traffic burst and verify that overload is observable and recoverable.

✅ What makes this special: The experiment teaches you to stop asking “Which vLLM flag makes inference fastest?” and start asking “Which resource is limiting this workload under this latency objective?” That second question transfers across models, GPU generations and serving frameworks.

🎯 Use this when... you want a repeatable optimization method rather than copying performance flags from someone else's hardware and workload.

6. Advanced Optimization: Prefix Caching, Quantization and Disaggregated Serving

🧒 Sticky-note analogy: Once a restaurant runs efficiently, the next improvements are specialized: prepare common ingredients in advance, use different equipment for different jobs, or split preparation and cooking into separate stations.

Prefix caching can reduce repeated prompt computation when requests share reusable prefixes. The production question is not “is prefix caching enabled?” but “does my workload generate enough cache hits to matter?” Current vLLM exposes prefix-cache queries and hits so this can be measured.

Quantization reduces numerical precision for some model representations and can lower memory requirements or change execution characteristics. The trade-off must be evaluated across hardware support, serving performance and model-quality regression tests. Do not infer quality from memory savings alone.

Speculative decoding is another advanced technique available in modern inference systems. Its purpose is to accelerate generation by proposing token candidates and validating them with the target model. Whether it helps depends on the model pair or speculative method, acceptance behavior, workload and serving configuration. Benchmark it rather than assuming universal improvement.

Disaggregated prefill and decode is a particularly important architecture to understand. Current NVIDIA TensorRT-LLM documentation describes aggregated serving, where context and generation share GPU resources, and disaggregated serving, where they execute on separate GPU pools.

Why separate them? Prefill and decode have different compute characteristics. NVIDIA documents that sharing resources can create interference in which context processing delays token generation. Disaggregation allows the two phases to use different resources and parallelism strategies, although it introduces the additional problem of transferring KV-cache state between them.

Aggregated serving

Request
   |
   +-- Prefill --+
                 |--> same GPU pool
   +-- Decode ---+


Disaggregated serving

Request
   |
   +--> Prefill GPU pool
           |
        KV state
           |
           v
        Decode GPU pool
           |
        Response

TensorRT-LLM's current design includes mechanisms for KV-cache exchange and overlapping cache transmission with computation from independent requests. This is a good example of why sophisticated optimization creates new costs: separating prefill and decode can reduce interference but now cache movement, routing and topology become part of the performance problem.

💡 Architecture warning: Advanced does not mean automatically better. A single well-sized GPU or conventional multi-GPU replica can be operationally superior when it already satisfies the workload. Add architectural complexity only when measurements identify a constraint that the complexity addresses.

🎯 Use this when... basic batching and capacity tuning no longer meet the required latency-throughput envelope and measurements justify a more specialized architecture.

7. Enterprise Rollout: Treat Performance as a Release Property

🧒 Sticky-note analogy: A new engine is not approved for an airline because it ran fast once. It must repeatedly satisfy defined operating conditions, monitoring and safety checks. Inference performance deserves similar engineering discipline.

Production optimization becomes sustainable when performance is treated as part of the release contract rather than as a one-time tuning exercise.

  1. Assign ownership. Define who owns model quality, inference configuration, GPU capacity, application SLOs, security and incident response.
  2. Version the benchmark workload. Preserve representative prompt/output distributions and important application scenarios alongside the release process.
  3. Version model and serving configuration. Record model revision, tokenizer/configuration, inference image/build, GPU topology and important serving parameters.
  4. Create CI performance gates. Run bounded regression benchmarks when inference-critical components change. vLLM itself currently maintains a performance dashboard driven by automated benchmark runs, illustrating the broader engineering principle of detecting performance regressions continuously.
  5. Gate model quality separately. Performance improvements must not silently bypass accuracy, safety, RAG-quality or application-behavior checks.
  6. Protect production-derived test data. If real prompts are sampled for workload modeling, establish minimization, redaction, access, retention and privacy controls.
  7. Establish capacity budgets. Track accelerator allocation and useful workload throughput together rather than rewarding GPU utilization by itself.
  8. Build operational dashboards. Include TTFT, TPOT or ITL, end-to-end latency, queue time, running/waiting requests, token distributions, KV-cache pressure, failures and GPU telemetry.
  9. Canary performance changes. Route a bounded eligible workload to the candidate and compare it with the approved baseline.
  10. Define rollback criteria before release. Roll back when latency, failures, quality, memory stability or other agreed SLOs materially regress.

Performance testing should be segmented. A single global percentile can hide an important regression. Compare short prompts against long prompts, RAG traffic against ordinary chat, streaming against non-streaming, and interactive requests against batch-oriented jobs when those categories matter to the application.

Cost belongs beside latency and throughput. A useful corporate capacity question is not “How busy is the GPU?” but “How much approved workload do we complete within our latency and quality objectives for the resources allocated?”

🎯 Use this when... inference performance affects business-critical applications and optimization changes need reproducible testing, approval, observability and rollback.

8. Common Mistakes and Their Production Impact

🧒 Sticky-note analogy: If every warning light on a car dashboard is replaced by one light called “CAR OK,” troubleshooting becomes impossible. LLM systems have the same problem when every performance signal is collapsed into one number.

Mistake 1: optimizing average latency. An average can conceal users stuck behind long queues. Production impact: a dashboard may look healthy while tail latency makes interactive use frustrating. Track distributions and separate queue, prefill and generation behavior.

Mistake 2: maximizing GPU utilization. Utilization is a resource signal, not a user SLO. A highly utilized accelerator with rapidly increasing queues can indicate saturation rather than success.

Mistake 3: assuming model fit equals serving capacity. KV cache and runtime memory also consume GPU memory. Long contexts and concurrency can therefore create pressure after the model loaded successfully.

Mistake 4: benchmarking one tiny prompt repeatedly. This creates a workload unlike many RAG, agent and enterprise-assistant deployments. Optimization decisions then target the benchmark rather than production.

Mistake 5: increasing concurrency indefinitely because throughput initially improved. Once serving capacity is saturated, additional demand can mostly become waiting work and tail latency.

Mistake 6: enabling every optimization simultaneously. If performance changes, attribution becomes difficult; if quality regresses, identifying the responsible change becomes harder. Change one meaningful variable at a time during controlled experiments.

Mistake 7: ignoring prompt and output distributions. Two services with the same request rate can impose radically different GPU and cache demands if their token distributions differ.

Mistake 8: treating prefix caching as free performance. Its benefit depends on actual prefix reuse. Observe cache queries and hits rather than assuming that enabling a feature means the workload benefits.

Mistake 9: scaling GPUs before identifying the bottleneck. More hardware may hide a scheduling, memory, workload-shaping or architecture problem while increasing cost. First determine whether the constraint is compute, cache capacity, queueing, communication or something else.

🎯 Use this when... performance tuning produces confusing results or infrastructure cost rises without a proportional improvement in user-visible service.

9. ❓ FAQ

1. Should I optimize TTFT or throughput first?

Start from the application's service objective rather than choosing a universal winner. Interactive assistants may care strongly about TTFT and streaming behavior, while offline processing may tolerate more latency in exchange for useful aggregate throughput. Measure both because improving one can affect the other.

2. Why can GPU memory run out when the model already fits?

Model weights are only part of serving memory. KV-cache state and runtime allocations also consume accelerator memory. Active sequences, context lengths and generated lengths therefore influence how much concurrent work the server can sustain.

3. Does higher concurrency always increase throughput?

No. Concurrency can initially help the scheduler use accelerator resources more effectively, but eventually capacity is reached. Beyond that region, additional demand can increasingly appear as waiting requests, queue time and worse tail latency.

4. When is prefix caching useful?

It is most relevant when requests contain reusable prefixes whose previously computed state can be reused. Measure prefix-cache queries and hits against representative traffic to determine whether the workload actually benefits.

5. When should an enterprise consider disaggregated prefill and decode?

Consider it when measured workloads show that prefill/decode interference or differing resource requirements justify separating the phases. The architecture adds KV-transfer, routing and operational complexity, so it should solve a demonstrated constraint rather than be adopted merely because it is advanced.

10. 🔗 References & Further Reading

Attribution notice: vLLM names and marks belong to their respective project owners; NVIDIA and TensorRT are trademarks or marks of NVIDIA Corporation. Other product names belong to their respective owners.

11. 📝 Summary

  • Foundations: Optimize a defined latency-throughput objective rather than chasing one generic “speed” number.
  • Mechanics: Prefill determines how the prompt enters the model; decode governs iterative token generation, so they expose different bottlenecks.
  • Batching: Continuous scheduling helps concurrent requests share accelerator execution, but excessive offered load eventually becomes queueing.
  • KV cache: Model weights are only one GPU-memory consumer; active sequence state can constrain real serving capacity.
  • Practical lab: Establish a reproducible baseline, raise load gradually, classify the bottleneck and change one variable at a time.
  • Advanced optimization: Prefix caching, quantization, speculative decoding and disaggregated serving solve different problems and require workload-specific validation.
  • Enterprise rollout: Version performance workloads, configurations and quality tests, then enforce them through CI, dashboards, canaries and rollback criteria.
  • Common mistakes: Avoid optimizing averages, utilization or synthetic micro-workloads while ignoring queues, token distributions and cache pressure.
  • FAQ: The correct optimization depends on workload shape and service objectives—not a universal configuration copied from another deployment.

The most valuable inference skill is not memorizing tuning flags. It is learning to look at a slow LLM service, follow the request through queueing, prefill, KV memory and decode, identify the constrained resource, and prove that your change improved the workload that actually matters. Happy optimizing! 

Comments