Modern AI Engineering

Lesson 17.1 · 29 min

GPUs for Deep Learning: Parallelism at the Core

Why does a chip designed to draw video-game pixels train and run almost every modern AI model — and why does the size of its memory matter as much as its speed?

In short: A GPU is a processor with thousands of simple cores that all perform the same kind of arithmetic at once, plus very fast memory to feed them. Deep learning is mostly huge matrix multiplications made of millions of independent multiply-adds, which is exactly the kind of work a GPU parallelises well. Specialised Tensor Cores, low-precision number formats and the CUDA software stack make it faster still, while VRAM size and memory bandwidth decide which models fit and how quickly they generate tokens.

What is a GPU?

A GPU (Graphics Processing Unit) is a processor originally built to draw images on screens. Drawing a frame of a video game means computing the colour of millions of pixels, and each pixel's calculation is almost the same as its neighbours' and does not depend on them. So GPU designers made a chip with thousands of small, simple cores that execute the same instructions on different data at the same time, instead of a few very smart cores.

Around 2007, NVIDIA released CUDA, a way to program GPUs for general maths rather than only graphics. In 2012, the AlexNet image classifier was trained on two consumer NVIDIA GPUs and won the ImageNet competition by a wide margin. Since then, GPUs have been the main engine of deep learning.

Our running example in this module: an appliance shop wants to run a 7-billion-parameter (7B) language model as its support chatbot, and later fine-tune it on its own tickets. Every decision — which GPU, how many, which number format — follows from the ideas in this lesson.

The math professor and the thousands of students

One genius or a stadium of students? Imagine we must mark 100,000 simple arithmetic worksheets. Option A: one brilliant maths professor, very fast and able to solve any hard problem, but only one at a time. Option B: a stadium of 10,000 school students, each slower and only able to do simple sums, but all working at once. For one tricky proof, the professor wins. For 100,000 simple worksheets, the students finish long before the professor. A CPU is the professor; a GPU is the stadium.

The analogy has a second lesson hidden in it: the students only help if we can hand out worksheets fast enough. If one clerk walks sheets to the stadium one at a time, most students sit idle. In a GPU, the 'clerk' is memory bandwidth, and we will see that it is often the real limit.

CPU vs GPU

Why deep learning is mostly matrix multiplication

A dense neural-network layer computes Y = X · W (plus a bias and an activation). X holds a batch of inputs (one row per example or token), W holds the learned weights. In a Transformer, the query/key/value projections, the attention output projection and the big feed-forward layers are all matrix multiplications, and together they account for the large majority of the arithmetic.

Each output cell Y[i, j] is a dot product: multiply row i of X by column j of W element by element, then add up. A 32×4096 input times a 4096×4096 weight matrix produces 131,072 output cells, each needing 4,096 multiplies and adds — about 1.1 billion floating-point operations (FLOPs) for one layer. Crucially, no output cell depends on any other, so all 131,072 can be computed at the same time.

Pause and think: How many FLOPs does multiplying a 2×3 matrix by a 3×4 matrix take, using the 2·M·K·N rule?

2 × 2 × 3 × 4 = 48 FLOPs: 8 output cells, each needing 3 multiplies and 3 adds.

Serial work vs parallel work

Serial work is a chain where each step needs the previous result, like following a recipe. Parallel work is a pile of independent jobs, like washing 100 plates with 100 people. Adding workers speeds up only the parallel part.

  • Very parallel: all output cells of a matrix multiply; applying an activation to every number; processing all tokens of a prompt at once in a Transformer.
  • Serial: the layers of a network (layer 2 needs layer 1's output); generating text one token at a time (token 51 needs token 50); a training loop's steps.

GPUs exploit the parallelism inside each step: each layer is a huge parallel matrix multiply, even though the layers run in order. This is also why the Transformer replaced the RNN: an RNN processes tokens one after another, while a Transformer processes a whole sequence in parallel during training.

Amdahl's law in one line If 5% of a job is strictly serial, then even infinitely many cores can speed it up at most 20× (1 / 0.05). This is why data loading, Python overhead and small sequential operations can leave an expensive GPU waiting.

GPU memory (VRAM) and memory bandwidth

A GPU has its own memory, called VRAM (video RAM) or device memory. Data-centre GPUs use HBM (High Bandwidth Memory): memory chips stacked right next to the processor. Two numbers matter:

  • Capacity (GB): how much fits. Examples: NVIDIA's A100 comes in 40 GB and 80 GB versions; the H100 SXM has 80 GB; consumer cards commonly have 8–32 GB.
  • Bandwidth (GB/s or TB/s): how fast data moves between VRAM and the cores. An H100 SXM is rated at about 3.35 TB/s; a typical desktop CPU's main memory delivers tens of GB/s, and server CPUs a few hundred GB/s.

Whether compute or bandwidth is the bottleneck depends on arithmetic intensity: how many FLOPs we do per byte we read. A GPU like the H100 can do on the order of a thousand trillion low-precision FLOPs per second but read only about 3 trillion bytes per second, so it needs roughly 300 FLOPs per byte to keep its cores busy. A matrix multiply with a large batch reuses each weight many times (high intensity: compute-bound). Generating text for one user with batch size 1 reads every weight to do only about 2 FLOPs with it (low intensity: memory-bound).

gpu_numbers.py

import numpy as np
# 1) A neural-network layer is a matrix multiply: (batch x d_in) @ (d_in x d_out)
B, d_in, d_out = 32, 4096, 4096
flops = 2 * B * d_in * d_out                 # one multiply + one add per term
print(f"one layer, batch {B}: {flops / 1e9:.1f} GFLOPs, "
f"{B * d_out:,} outputs that are all independent")
# 2) Every output cell is its own dot product -> perfectly parallel work
rng = np.random.default_rng(0)
X, W = rng.normal(size=(3, 4)), rng.normal(size=(4, 2))
cell = sum(X[1, k] * W[k, 0] for k in range(4))   # what ONE GPU thread does
print("one cell by hand:", round(cell, 4), "| same cell from X@W:", round((X @ W)[1, 0], 4))
# 3) Will the model fit in VRAM?  weights only = params x bytes per param
params = 7e9
for name, nbytes in [("FP32", 4), ("FP16/BF16", 2), ("INT8", 1), ("INT4", 0.5)]:
print(f"7B model in {name:9s}: {params * nbytes / 1e9:5.1f} GB of weights")
# Training needs far more: a common mixed-precision Adam estimate is
# ~16 bytes/param (weights + grads + fp32 master copy + 2 optimizer states)
print(f"7B model, full training state: ~{params * 16 / 1e9:.0f} GB (+ activations)")
# 4) Arithmetic intensity: FLOPs done per byte moved from memory (FP16)
for B in (1, 256):
bytes_moved = 2 * (B * d_in + d_in * d_out + B * d_out)
print(f"batch {B:3d}: {2 * B * d_in * d_out / bytes_moved:6.1f} FLOPs per byte")

Output:

one layer, batch 32: 1.1 GFLOPs, 131,072 outputs that are all independent
one cell by hand: 0.4751 | same cell from X@W: 0.4751
7B model in FP32     :  28.0 GB of weights
7B model in FP16/BF16:  14.0 GB of weights
7B model in INT8     :   7.0 GB of weights
7B model in INT4     :   3.5 GB of weights
7B model, full training state: ~112 GB (+ activations)
batch   1:    1.0 FLOPs per byte
batch 256:  227.6 FLOPs per byte

Why the model must fit in VRAM

Every token the model processes needs every weight. If the weights do not all fit in VRAM, some must be fetched from CPU memory over the PCIe bus, which is far slower than HBM (PCIe 5.0 x16 is roughly 64 GB/s in each direction). The model then runs many times slower, or simply fails with an out-of-memory (OOM) error.

VRAM must hold more than the weights. For inference: weights + the KV cache (stored keys and values for every token in every active conversation) + temporary activations. For training: weights + gradients + optimizer states (Adam keeps two extra numbers per parameter) + saved activations for backpropagation. That is why a 7B model that runs on one 24 GB card in FP16 needs on the order of 100+ GB to fully fine-tune — or tricks like LoRA, quantization and sharding across GPUs.

Pause and think: Our shop's GPU has 24 GB of VRAM. Can it serve the 7B model in FP16? In FP32?

FP16 weights take 14 GB, leaving about 10 GB for the KV cache and overhead, so yes for moderate context lengths and a few users. FP32 weights alone take 28 GB, more than the card has, so no.

Tensor Cores and lower precision (FP16, BF16, INT8)

Regular GPU cores (NVIDIA calls them CUDA cores) do one multiply-add per clock each. Starting with the Volta generation (V100, 2017), NVIDIA added Tensor Cores: units that multiply small matrix tiles (for example 4×4 blocks) in one operation. They deliver far more FLOPs than the regular cores, but only for matrix maths in reduced-precision formats.

Common number formats in deep learning
FormatBitsTypical useNote
FP3232Classic training; master weightsPrecise but slow and large
TF3219 used (stored in 32)FP32 matmuls on Ampere and laterFP32 range with reduced mantissa
FP1616Mixed-precision training, inferenceSmall range: needs loss scaling to avoid underflow
BF1616Mixed-precision training, inferenceSame range as FP32, less precision; usually no loss scaling needed
FP88Training and inference on Hopper and newerNeeds careful scaling
INT8 / INT48 / 4Quantised inferenceHalves or quarters memory again; small accuracy cost if done well

Lower precision helps in three ways at once: Tensor Cores run faster on it, each number takes fewer bytes (so the model fits and memory traffic drops), and more numbers fit in caches. Mixed precision training does the heavy matrix maths in BF16 or FP16 while keeping a master copy of the weights and sensitive reductions in FP32, giving most of the speed with little loss of accuracy.

Common mistake Training in plain FP16 without loss scaling can make small gradients underflow to zero and silently stall learning; switching to BF16 or enabling automatic loss scaling fixes it. Also, a GPU's headline FLOPs number is usually for its lowest-precision format (sometimes with sparsity): compare like with like.

CUDA and the software stack (cuDNN)

Hardware is useless without software that keeps it busy. A deep-learning call like model(x) in PyTorch passes through several layers before any transistor switches:

From Python to silicon

  1. Framework (PyTorch, JAX, TensorFlow): We describe the model in Python. The framework turns each operation into calls to GPU libraries, and compilers (like torch.compile or XLA) can fuse several operations into one.
  2. Libraries (cuBLAS, cuDNN, NCCL): Highly tuned NVIDIA libraries: cuBLAS for matrix maths, cuDNN for neural-network operations such as convolutions, normalisation and attention, NCCL for communication between GPUs.
  3. CUDA kernels: Each library call launches kernels: small programs that run on thousands of GPU threads at once (next lesson). Custom kernels such as FlashAttention speed up specific operations further.
  4. CUDA runtime and driver: Manage GPU memory, move data between CPU and GPU, schedule kernel launches.
  5. The GPU hardware: Streaming Multiprocessors (SMs) run the threads on CUDA cores and Tensor Cores, reading from caches, shared memory and HBM.

Training vs inference on GPUs

Multiple GPUs working together

Large models need many GPUs, either because the model is too big for one card or because training on one card would take years. GPUs inside a server talk over NVLink (much faster than PCIe), and servers talk over networks like InfiniBand or high-speed Ethernet. The main ways to split the work:

Ways to split work across GPUs
StrategyWhat is splitMain cost
Data parallelismEach GPU has a full model copy and its own slice of the batch; gradients are averaged (all-reduce)Every GPU must hold the whole model
Sharded data parallelism (ZeRO / FSDP)Weights, gradients and optimizer states are split across GPUs and gathered when neededMore communication
Tensor parallelismEach matrix is split across GPUs; each computes part of every layerNeeds very fast links (within a server)
Pipeline parallelismDifferent layers live on different GPUs; micro-batches flow through like an assembly lineIdle 'bubbles' at the start and end

Large training runs combine all of these. For our shop's 7B model, one GPU is enough for inference; fine-tuning would typically use LoRA on one GPU or sharded data parallelism across a few.

Why NVIDIA GPUs power modern AI

  • An early start in software. CUDA (2007) gave researchers a practical way to program GPUs years before deep learning took off, and AlexNet (2012) was built on it.
  • The ecosystem. cuDNN, cuBLAS, NCCL, TensorRT and deep framework integration mean new research code usually runs on NVIDIA first. This software moat is as important as the chips.
  • Hardware aimed at AI. Tensor Cores, HBM, and low-precision formats (BF16, FP8 and newer) added generation by generation.
  • Scale-out. NVLink, NVSwitch and networking let thousands of GPUs act as one training system.

It is not the only option AMD GPUs (with the ROCm software stack), Google TPUs, and specialised inference chips (such as LPUs, covered later in this module) compete on price, efficiency or speed for specific workloads. The best choice depends on the model, the software we rely on, availability and cost.

When a GPU is the wrong tool Small models with tiny batches, heavy branching logic, classic machine learning on small tables (gradient-boosted trees often run fine on CPUs), or workloads dominated by data loading. Moving small data to the GPU and back can cost more time than the computation saves.

Worked example, step by step

Let us turn 'decoding is memory-bound' into a number for our shop's chatbot. We want a rough ceiling on how fast one conversation can receive tokens from the 7B model on a GPU with about 3.35 TB/s of memory bandwidth (the H100 SXM figure from earlier). This is back-of-envelope arithmetic, not a benchmark: real speeds are lower.

From bandwidth to tokens per second

  1. What one token costs in bytes: To produce one token for one user, every weight is read once. In FP16 that is 7 × 10⁹ × 2 bytes = 14 GB of memory traffic per token.
  2. What the memory can deliver: About 3.35 TB/s, which is 3,350 GB every second.
  3. Divide: 3,350 / 14 ≈ 239 tokens per second at the very most. No amount of extra compute can beat this, because the weights cannot arrive faster.
  4. Check the compute side: Each weight is used for about 2 FLOPs, so one token needs about 14 billion FLOPs. At 239 tokens/s that is roughly 3.3 trillion FLOPs per second. The GPU can do on the order of a thousand trillion. The cores are idle well over 99% of the time.
  5. Shrink the weights: In INT4 the weights are 3.5 GB, so the ceiling becomes 3,350 / 3.5 ≈ 957 tokens/s. Quantisation does not only make the model fit; it also raises the speed limit.
  6. Share the read: With 8 users batched together, the weights are still read once per step, but that one pass now yields 8 tokens. Total throughput can rise close to 8× while each user sees about the same speed. This keeps working until compute, not bandwidth, becomes the limit.
Decode ceiling for one stream = bandwidth / bytes of weights (3,350 GB/s, 7B parameters)
FormatWeightsCeiling per stream
FP16 / BF1614 GB≈ 239 tokens/s
INT87 GB≈ 479 tokens/s
INT43.5 GB≈ 957 tokens/s

Why is the real number lower? The KV cache must be read too, kernels have launch overhead, sampling the next token takes time, and no kernel uses 100% of the rated bandwidth. But the ceiling is still useful. If a vendor or a teammate quotes a single-stream speed above it, something else is going on, such as a smaller model or a different number format.

Practice: try it yourself

We will put four ideas from this lesson into one short script: serial versus vectorised matrix multiplication, how many rounds a pile of independent cells needs with few or many workers, Amdahl's law, and the decode ceiling we just worked out by hand.

practice_parallel_numbers.py

import math, time
import numpy as np
rng = np.random.default_rng(0)
# 1) Same matrix product two ways: one cell at a time, or all cells at once.
n = 60
X, W = rng.normal(size=(n, n)), rng.normal(size=(n, n))
X @ W                                            # warm-up call, not timed
t0 = time.perf_counter()
Y_loop = [[sum(X[i, k] * W[k, j] for k in range(n)) for j in range(n)]
for i in range(n)]                     # serial: 3,600 cells in turn
t1 = time.perf_counter()
Y_vec = X @ W                                    # vectorised: one library call
t2 = time.perf_counter()
print("same answer:", np.allclose(Y_loop, Y_vec))
print("vectorised call at least 10x faster:", (t1 - t0) > 10 * (t2 - t1))
# 2) Rounds needed if each worker finishes one output cell per round.
cells = 32 * 4096                                # the lesson's layer, batch 32
for name, workers in [("8 fast cores", 8), ("4,096 simple cores", 4096)]:
print(f"{name:18s}: {math.ceil(cells / workers):6,} rounds for {cells:,} cells")
# 3) Amdahl's law: speed-up = 1 / (serial + parallel / workers)
for serial in (0.0, 0.05, 0.5):
s = 1 / (serial + (1 - serial) / 4096)
print(f"serial share {serial:4.0%}: speed-up with 4,096 workers = {s:6.1f}x")
# 4) Decode ceiling: each new token reads every weight once (batch size 1).
bandwidth_gb_s = 3350                            # about 3.35 TB/s, as in the lesson
for name, gb in [("FP16", 14.0), ("INT8", 7.0), ("INT4", 3.5)]:
print(f"7B model in {name}: at most ~{bandwidth_gb_s / gb:.0f} tokens/s per stream")

Output:

same answer: True
vectorised call at least 10x faster: True
8 fast cores      : 16,384 rounds for 131,072 cells
4,096 simple cores:     32 rounds for 131,072 cells
serial share   0%: speed-up with 4,096 workers = 4096.0x
serial share   5%: speed-up with 4,096 workers =   19.9x
serial share  50%: speed-up with 4,096 workers =    2.0x
7B model in FP16: at most ~239 tokens/s per stream
7B model in INT8: at most ~479 tokens/s per stream
7B model in INT4: at most ~957 tokens/s per stream

Now change it:

  • Set cells = 1 * 4096 (batch size 1). Predict the rounds for both machines. What does the answer say about how busy the 4,096 workers are?
  • In part 3, change 4096 to 8 workers. Predict the speed-up for a 5% serial share before running. Is the serial part still the main limit?
  • Set bandwidth_gb_s = 50, a figure in the 'tens of GB/s' range of desktop CPU memory. Predict the FP16 ceiling. Would a faster CPU help?

Pause and think: A training job spends half its time loading and preparing data on the CPU and half on GPU maths. We swap in a GPU that is twice as fast. How much faster is the whole job?

Only about 1.33×. The data half is unchanged (0.5) and the GPU half shrinks to 0.25, so the new time is 0.75 of the old one, and 1 / 0.75 ≈ 1.33. This is Amdahl's law: the part we did not speed up now dominates. Fixing the data loading would pay off more than a third GPU upgrade.

Pause and think: We batch 8 users together on a memory-bound decoder. Does each user now get tokens 8 times more slowly?

No. The expensive part of a decode step is streaming the weights from memory, and that happens once per step however many users share it. Eight users mean roughly eight tokens per pass instead of one, so total throughput rises while each user's speed stays about the same. The extra arithmetic uses cores that were idle anyway. This holds until the batch is large enough for compute to become the limit.

Key takeaways

  • A GPU trades a few smart cores for thousands of simple ones, ideal for parallel matrix maths.
  • Deep learning is dominated by matrix multiplications whose output cells are independent.
  • VRAM capacity decides what fits; memory bandwidth often decides how fast LLMs generate.
  • Tensor Cores plus low precision (BF16, FP8, INT8) multiply speed and cut memory.
  • CUDA, cuBLAS, cuDNN and NCCL turn the hardware into usable speed; multi-GPU setups split data, layers or matrices.

Key terms

  • GPU: A processor with thousands of simple cores designed for parallel, throughput-oriented computation.
  • VRAM / HBM: The GPU's own memory; HBM is high-bandwidth memory stacked beside the chip.
  • Memory bandwidth: How many bytes per second can move between memory and the compute units.
  • Arithmetic intensity: FLOPs performed per byte moved from memory; decides compute-bound vs memory-bound.
  • Tensor Core: A GPU unit that multiplies small matrix tiles in one operation in reduced precision.
  • Mixed precision: Training with low-precision maths while keeping key values (master weights) in FP32.
  • cuDNN: NVIDIA's library of optimised deep-learning operations built on CUDA.

← 16.6 Variational Autoencoders: Learning a Compressed Latent Space · 17.2 CUDA Kernels: Writing Parallel Code for NVIDIA GPUs →