Modern AI Engineering

Lesson 13.17 · 31 min

TensorRT-LLM: NVIDIA's Optimized Inference Engine

If a GPU can do trillions of operations per second, why does it often sit idle while serving an LLM, and how does NVIDIA's TensorRT-LLM get that time back?

In short: TensorRT-LLM is NVIDIA's open-source library for running LLMs as fast as possible on NVIDIA GPUs. It prepares the model ahead of time (fusing operations into fewer kernels, picking the fastest kernel for each operation, and using low-precision formats like FP8), then serves it with a paged KV cache, in-flight batching, CUDA graphs, speculative decoding and multi-GPU parallelism. Newer versions add a PyTorch backend that keeps most of the speed without a separate build step.

What is inference

Inference means using a trained model to produce outputs, as opposed to training, which adjusts its weights. For an LLM, inference is: take a prompt, run the prefill phase (process all prompt tokens at once and fill the KV cache, the stored keys and values of past tokens), then run the decode phase (produce one token at a time until done).

Inference is where the money goes in production: a model is trained once, but it may answer billions of requests. A 2× faster inference engine can halve the GPU bill or double the number of users. That is the job TensorRT-LLM was built for.

Think of it like a pit crew A race car (the GPU) is incredibly fast, but races are often lost in the pits. A great pit crew rehearses every move in advance, combines tasks, and never leaves the car waiting. TensorRT-LLM is the pit crew: it plans and rehearses the model's work ahead of time so the GPU spends its time racing, not waiting.

Our running example: a company serves a 70B chat model on a server with 8 NVIDIA H100 GPUs and wants the most tokens per second per dollar while keeping replies snappy.

What is a GPU and what is a kernel

A GPU is a processor with thousands of small cores designed to do the same operation on lots of data at once. It also has its own fast memory (HBM, high-bandwidth memory) where the model's weights and KV cache live. Data must be read from HBM into the cores to be used, and that reading has a limited speed (the memory bandwidth).

A kernel is one function that runs on the GPU, such as "multiply these two matrices" or "add a bias and apply an activation". The CPU tells the GPU to run kernels one after another; each such request is a kernel launch, and each launch costs a few microseconds of overhead. A single forward pass through a transformer can launch hundreds of kernels.

The problem: the GPU spends its time on the wrong things

Run a model naively (for example, plain PyTorch, one operation at a time) and the GPU wastes much of its time:

  • Moving data, not computing: each small operation reads its input from HBM and writes its output back, only for the next operation to read it again. Many operations do very little maths per byte moved.
  • Waiting for launches: during decode each kernel finishes in microseconds, so the CPU's launch overhead and Python code can become the bottleneck, leaving the GPU idle between kernels.
  • Generic kernels: a one-size-fits-all kernel is rarely the fastest for a particular matrix shape and GPU model.
  • Too many bits: 16-bit weights and KV cache take twice the memory and bandwidth of 8-bit ones.
  • Poor batching and memory use: without paging and continuous batching, memory is wasted and batches stay small.

Pause and think: During decode with a small batch, a matrix multiply finishes in 8 microseconds and a kernel launch costs 5 microseconds of CPU time. What happens if the CPU cannot launch ahead of the GPU?

The GPU spends a large share of its time idle, waiting for the next launch: roughly 5 of every 13 microseconds (around 40%) are pure overhead. Reducing the number of launches (fusion) or removing launch overhead (CUDA graphs) recovers that time.

What is TensorRT-LLM, and the big idea

TensorRT-LLM is an open-source library from NVIDIA (first released in late 2023, Apache-2.0 licensed) for high-performance LLM inference on NVIDIA GPUs. It builds on TensorRT, NVIDIA's long-standing deep-learning compiler and runtime, and on techniques from NVIDIA's earlier FasterTransformer library. It provides optimized model definitions for many popular LLM families, a Python LLM API, a C++ runtime with a batch scheduler, and serving integrations.

The big idea is to prepare the model ahead of time for one specific situation: this model, this GPU type, this precision, these maximum batch and sequence sizes. With all of that fixed, the compiler can make aggressive choices (fuse operations, pick the fastest kernels, plan memory) that a general framework, which must handle anything at any moment, cannot.

The build step: from a model to an engine

The classic TensorRT-LLM workflow

  1. Start from a checkpoint: Take a model from Hugging Face (safetensors weights plus config).
  2. Convert and optionally quantize: Convert the weights into TensorRT-LLM's checkpoint format for the target parallelism (for example split for 8 GPUs), and optionally quantize them, for example to FP8 using NVIDIA's Model Optimizer with a small calibration set.
  3. Build the engine: Run trtllm-build with limits such as max batch size, max input and output length, and plugins to enable. TensorRT fuses layers, tries candidate kernels for each operation on the real GPU and keeps the fastest, and plans memory.
  4. Get an engine file: The result is a serialized engine: an optimized execution plan tied to that GPU architecture, TensorRT-LLM version, precision and limits.
  5. Run it: The runtime loads the engine and serves requests with the batch scheduler, paged KV cache and the other features below.

Common mistake Building an engine on one GPU type and copying it to another, or building with a max sequence length of 2,048 and then sending 8,000-token prompts. Engines are specific to the GPU architecture, the library version and the limits they were built with. Plan the limits from real traffic, and rebuild when hardware or versions change.

Kernel fusion

Kernel fusion merges several small operations into one kernel. Instead of "add bias, write to memory; read, apply GELU, write; read, add residual, write", a fused kernel reads the inputs once, does all three steps while the data sits in fast on-chip registers, and writes the result once. Fewer launches and far less memory traffic, with exactly the same maths.

fusion_and_graphs.py

import numpy as np
rng = np.random.default_rng(1)
x = rng.normal(size=(8, 4096)).astype(np.float16)        # activations for 8 tokens
bias = rng.normal(size=4096).astype(np.float16)
resid = rng.normal(size=(8, 4096)).astype(np.float16)
def gelu(v):                                              # tanh approximation of GELU
return 0.5 * v * (1 + np.tanh(0.79788456 * (v + 0.044715 * v ** 3)))
# Unfused: three separate "kernels", each reads inputs from memory and writes a result back.
t1 = x + bias                                             # kernel 1
t2 = gelu(t1)                                             # kernel 2
out_unfused = t2 + resid                                  # kernel 3
# Fused: one kernel reads x, bias, resid once and writes the final result once.
out_fused = gelu(x + bias) + resid
print("same result:", np.allclose(out_unfused, out_fused))
n = x.nbytes                                              # bytes in one 8 x 4096 FP16 tensor
unfused_traffic = (n + bias.nbytes + n) + (n + n) + (n + n + n)  # reads + writes per kernel
fused_traffic = n + bias.nbytes + n + n                   # read x, bias, resid; write out
print(f"memory traffic: unfused {unfused_traffic / 1024:.0f} KiB, fused {fused_traffic / 1024:.0f} KiB "
f"({unfused_traffic / fused_traffic:.1f}x less)")
# Launch overhead: decoding one token in a 32-layer model may launch hundreds of kernels.
layers, kernels_per_layer, launch_us, gpu_work_us = 32, 15, 5.0, 2500.0   # illustrative
launches = layers * kernels_per_layer
cpu_us = launches * launch_us
print(f"{launches} kernel launches -> {cpu_us / 1000:.1f} ms of launch overhead per token")
for name, overhead in [("eager launches", cpu_us), ("one CUDA graph replay", launch_us)]:
total = gpu_work_us + overhead
print(f"{name:22s}: {total / 1000:.2f} ms/token -> {1e6 / total:5.0f} tokens/s")

Output:

same result: True
memory traffic: unfused 456 KiB, fused 200 KiB (2.3x less)
480 kernel launches -> 2.4 ms of launch overhead per token
eager launches        : 4.90 ms/token ->   204 tokens/s
one CUDA graph replay : 2.50 ms/token ->   399 tokens/s

TensorRT performs many fusions automatically, and TensorRT-LLM adds hand-written fused kernels for the hot spots of transformers: fused attention, fused normalization plus quantization, fused gated MLP activations, and fused mixture-of-experts routing among others.

Quantization and custom attention kernels

Quantization stores numbers in fewer bits. TensorRT-LLM supports many recipes and matches each to hardware that can run it fast:

Common TensorRT-LLM precision options (support depends on GPU generation and version)
RecipeWhat is low-precisionNotes
FP8Weights and activations (and optionally KV cache) in 8-bit floating pointRuns on FP8 tensor cores of Ada/Hopper/Blackwell GPUs; a popular near-lossless default
NVFP44-bit floating point with fine-grained scalesBlackwell-generation GPUs
INT8 SmoothQuantWeights and activations in INT8, outliers smoothed into weightsWorks on older GPUs too
INT4 AWQ / GPTQ, INT8 / INT4 weight-onlyWeights only; activations stay 16-bitCuts memory for memory-bound decode
FP8 / INT8 KV cacheThe stored keys and valuesFits more tokens and users in memory

Custom attention kernels matter because attention is where prefill and decode behave very differently. For prefill, TensorRT-LLM uses fused multi-head attention kernels in the spirit of FlashAttention: compute attention in tiles that stay in on-chip memory instead of writing the full attention matrix to HBM. For decode, where each request has one new query token attending over a long KV cache, it uses specialised kernels (such as its XQA kernels) tuned for multi-query and grouped-query attention, so the many query heads that share one KV head read that cached data efficiently.

The paged KV cache, in-flight batching and CUDA graphs

Paged KV cache. Like vLLM's PagedAttention, TensorRT-LLM stores the KV cache in fixed-size blocks allocated on demand, instead of reserving one contiguous max-length region per request. This removes most memory waste and lets blocks be reused across requests that share a prefix, such as a common system prompt (KV cache reuse).

In-flight batching is NVIDIA's name for continuous batching: the batch is re-formed at every iteration, so finished requests leave and new ones join immediately. Requests in their prefill phase and requests in decode can be processed in the same batch, and long prompts can be split into chunks so they do not stall everyone else.

CUDA graphs. A CUDA graph records a whole sequence of kernel launches once (capture) and then replays the entire sequence with a single launch. Decode steps repeat the same kernels every token, so TensorRT-LLM captures graphs for common batch sizes and replays them, removing most per-kernel CPU overhead. In the simplified model above, that alone roughly doubled the token rate.

Speculative decoding and running one model across many GPUs

Speculative decoding uses spare GPU compute during decode: something cheap drafts several tokens and the main model verifies them in one pass, keeping the output unchanged. TensorRT-LLM supports several drafting methods, including a separate draft model, Medusa heads, EAGLE-style drafters, ReDrafter, lookahead decoding, and multi-token-prediction modules for models trained with them. Availability varies by backend and version.

A 70B model in FP16 needs about 140 GB for weights alone, more than one 80 GB GPU. TensorRT-LLM splits models across GPUs in three main ways:

  • Tensor parallelism (TP): each weight matrix is split across GPUs; every GPU computes part of every layer and they exchange results after each layer over fast links (NVLink). Lowers latency, but needs fast interconnect.
  • Pipeline parallelism (PP): different GPUs hold different layers; requests flow through them like an assembly line. Less communication, useful across nodes.
  • Expert parallelism (EP): for mixture-of-experts models, different experts live on different GPUs and tokens are routed to them.

Our 70B model In FP8 the 70B model's weights are about 70 GB. With tensor parallelism of 8 on the H100 server, each GPU holds about 9 GB of weights, leaving most of its memory for the KV cache, so many users fit at once. Alternatively the team could run two copies with TP 4 each, trading per-request latency for more independent capacity.

How we serve the model, and the PyTorch backend

An engine is not a web service by itself. Common ways to put TensorRT-LLM behind an API:

  • trtllm-serve: a built-in OpenAI-compatible HTTP server (chat and completions endpoints).
  • Triton Inference Server with the TensorRT-LLM backend: NVIDIA's general model server, with metrics, multiple models and ensembles (for example tokenizer, model and detokenizer as one pipeline).
  • Higher-level NVIDIA offerings such as NIM microservices and the Dynamo distributed serving framework use TensorRT-LLM (among other engines) under the hood.

The classic engine-build workflow is powerful but heavy: long builds, rebuilds for every change, and model code written in TensorRT-LLM's own graph-building API. So the project added a PyTorch backend: models are written in ordinary PyTorch, then accelerated with TensorRT-LLM's custom kernels (attention, fused MoE, quantized GEMMs), CUDA graphs, the paged KV cache, in-flight batching and speculative decoding, without a separate engine-compilation step. During 2025 this became the recommended default path, accessed through the high-level LLM API.

llm_api.py (needs an NVIDIA GPU and the tensorrt_llm package; no output shown)

from tensorrt_llm import LLM, SamplingParams
# Loads a Hugging Face model and prepares it for fast inference on the local GPU(s).
llm = LLM(model="meta-llama/Llama-3.1-8B-Instruct")
params = SamplingParams(temperature=0.7, max_tokens=64)
for out in llm.generate(["How do I reset my password?"], params):
print(out.outputs[0].text)

TensorRT-LLM vs vLLM, and where it works well or fails

Where it works well: stable, high-volume production on modern NVIDIA GPUs; models with first-class support; latency-sensitive services that benefit from FP8/FP4, CUDA graphs and speculative decoding; large multi-GPU deployments.

Where it struggles: non-NVIDIA hardware (not supported); very new or custom architectures without an optimized implementation; fast-changing experiments where rebuilding engines slows the team down; small teams for whom setup and tuning effort outweighs a modest speed gain.

Further learning Deep-dive: revisit the KV Cache, Paged Attention, and Continuous Batching lessons in this module to see how TensorRT-LLM combines all three for maximum GPU utilization.

Going one level deeper

fusion_and_graphs.py looked at one case: 2.5 ms of GPU work and 480 launches per token. Launch overhead is a fixed cost per pass, so how much it hurts depends on how long the pass is. Let us vary the case by hand. Every number here is illustrative, and we assume, as the code did, that launches are not hidden behind GPU work.

The same launches, five situations

  1. Decode, small batch: 480 launches × 5 µs = 2.4 ms on top of 2.5 ms of GPU work. Overhead share: 2.4 ÷ 4.9 ≈ 49%. Nearly half of every step is waiting.
  2. Add fusion: Suppose fusion cuts 15 kernels per layer to 6. Then 32 × 6 = 192 launches cost 0.96 ms. Share: 0.96 ÷ 3.46 ≈ 28%. (Fusion also cuts memory traffic; we hold the GPU work fixed to look at launches alone.)
  3. Replay a CUDA graph: The whole sequence costs about one launch, 0.005 ms. Share: about 0.2%.
  4. Decode, large batch: With many users in the batch the GPU work per step grows, say to 10 ms, while the launches stay at 2.4 ms. Share: 2.4 ÷ 12.4 ≈ 19%.
  5. Prefill of a long prompt: Say the pass needs 60 ms of GPU work. The same 2.4 ms of launches is now 2.4 ÷ 62.4 ≈ 4%.
Share of a pass lost to launch overhead (illustrative numbers)
SituationGPU workLaunch overheadOverhead share
Decode, small batch, eager launches2.5 ms2.4 ms≈ 49%
Decode, small batch, after fusion2.5 ms0.96 ms≈ 28%
Decode, small batch, CUDA graph2.5 ms0.005 ms≈ 0.2%
Decode, large batch, eager launches10 ms2.4 ms≈ 19%
Long prefill, eager launches60 ms2.4 ms≈ 4%

The pattern: the shorter the pass, the more a fixed cost hurts. Small-batch decode has the shortest passes and repeats them for every token, so that is where fusion and CUDA graphs pay off most. A long prefill barely notices the launches; there the fused attention kernels and lower precision matter more, because they cut the GPU work itself.

This gives us a habit for reading any speedup claim. Ask which part of the time it removes, and how big that part is in our own workload. A trick that doubles the token rate at batch size 1 may add only a few percent on a server that always runs large batches.

Practice: try it yourself

The lesson described tensor parallelism in words: each weight matrix is split across GPUs, every GPU computes part of every layer, and the parts are combined. We will do exactly that with a tiny matrix in numpy and check that the answer does not change. Then we use plain arithmetic to see how the weights of our 70B model spread over 80 GB GPUs at different precisions and TP sizes.

practice_tensor_parallel.py

import numpy as np
rng = np.random.default_rng(0)
d_in, d_out, TP = 8, 16, 4
W = rng.normal(size=(d_out, d_in))           # one layer's weight matrix
x = rng.normal(size=d_in)                    # one token's activation vector
# Tensor parallelism: each "GPU" holds a slice of the rows of W.
shards = np.array_split(W, TP, axis=0)
parts = [s @ x for s in shards]              # every GPU computes its part of the layer
y_tp = np.concatenate(parts)                 # the parts are gathered after the layer
print("rows of W on each GPU:", [s.shape[0] for s in shards])
print("same result as one GPU:", np.allclose(y_tp, W @ x))
# Weights per GPU for a 70B model on 80 GB GPUs: does it fit, and what is left?
GPU_GB, PARAMS = 80, 70e9
for name, bytes_per_weight in [("FP16", 2), ("FP8", 1)]:
for tp in (1, 2, 4, 8):
per_gpu = PARAMS * bytes_per_weight / 1e9 / tp
left = GPU_GB - per_gpu
status = f"{left:5.2f} GB left on each GPU" if left > 0 else "does not fit"
print(f"{name:4s} TP={tp}: {per_gpu:6.2f} GB of weights per GPU -> {status}")

Output:

rows of W on each GPU: [4, 4, 4, 4]
same result as one GPU: True
FP16 TP=1: 140.00 GB of weights per GPU -> does not fit
FP16 TP=2:  70.00 GB of weights per GPU -> 10.00 GB left on each GPU
FP16 TP=4:  35.00 GB of weights per GPU -> 45.00 GB left on each GPU
FP16 TP=8:  17.50 GB of weights per GPU -> 62.50 GB left on each GPU
FP8  TP=1:  70.00 GB of weights per GPU -> 10.00 GB left on each GPU
FP8  TP=2:  35.00 GB of weights per GPU -> 45.00 GB left on each GPU
FP8  TP=4:  17.50 GB of weights per GPU -> 62.50 GB left on each GPU
FP8  TP=8:   8.75 GB of weights per GPU -> 71.25 GB left on each GPU

The last line matches the lesson's example: FP8 with TP 8 puts about 9 GB of weights on each GPU. The line for FP8 with TP 4 is the two-copies option: 17.5 GB of weights per GPU. Now change it:

  • Set TP = 3. The 16 rows no longer divide evenly. Predict the rows per GPU and whether the result still matches.
  • Add a 4-bit option, ("INT4", 0.5), to the list of precisions. Predict the weights per GPU at TP 1 and how much memory is left.
  • Split the columns instead: use axis=1, and give each shard its own slice of x with np.array_split(x, TP). Predict how the parts must now be combined. Is it still np.concatenate?

Pause and think: In the practice output, the FP8 model fits on a single GPU (70 GB, with 10 GB left). Why would the team still spread it over 8 GPUs with tensor parallelism?

Two reasons from the lesson. With only 10 GB left there is very little room for the KV cache, so few users fit at once, while TP 8 leaves about 71 GB free on each GPU. And with tensor parallelism every GPU computes part of every layer, which lowers latency. Fitting the weights is the minimum, not the goal.

Pause and think: In our simplified model, CUDA graphs roughly doubled the token rate for small-batch decode. Would they also double the speed of a long prefill?

No. Launch overhead is a fixed cost per pass. In small-batch decode it is about as large as the GPU work, so removing it nearly halves the step. In a long prefill the GPU work is many times larger than the launches, so removing them saves only a few percent.

Key takeaways

  • Naive inference leaves NVIDIA GPUs idle: too much memory traffic, launch overhead, generic kernels and wasted KV memory.
  • TensorRT-LLM prepares the model ahead of time: fused, auto-tuned kernels for one GPU, precision and size limit.
  • Key runtime features: FP8/FP4 and INT quantization, custom attention kernels, paged KV cache, in-flight batching, CUDA graphs.
  • Speculative decoding plus tensor, pipeline and expert parallelism scale it from one GPU to many.
  • Serve via trtllm-serve or Triton; the PyTorch backend removes the separate build step. It is NVIDIA-only, so benchmark against vLLM and SGLang.

Key terms

  • Kernel: A single function that runs on the GPU, such as a matrix multiply or a fused activation.
  • Engine: TensorRT-LLM's compiled, optimized execution plan for one model, GPU architecture, precision and set of limits.
  • Kernel fusion: Combining several operations into one kernel to cut memory traffic and launch overhead.
  • In-flight batching: NVIDIA's term for continuous batching: re-forming the batch at every iteration.
  • CUDA graph: A recorded sequence of GPU kernel launches that can be replayed with a single launch.
  • Tensor parallelism: Splitting each weight matrix across GPUs so they compute every layer together.
  • FP8: An 8-bit floating-point format that recent NVIDIA GPUs can multiply natively and fast.

← 13.16 SGLang: Structured LLM Programs for Efficient Inference · 14.1 Evaluating LLMs: Metrics, Benchmarks, and Methods →