Modern AI Engineering

Lesson 13.14 · 27 min

llama.cpp: Running Large Models on Consumer Hardware

How does a model that was trained on thousands of data-centre GPUs end up answering questions on an ordinary laptop with no internet connection?

In short: llama.cpp is an open-source C/C++ engine that runs LLMs on everyday CPUs and GPUs. It combines four ideas: quantize the weights to about 4–8 bits so they fit, pack everything into one GGUF file, memory-map that file so it loads instantly, and squeeze speed out of the hardware with SIMD instructions, threads and optional GPU offload. Because token generation is limited by memory bandwidth, smaller weights translate almost directly into faster answers.

What is llama.cpp

llama.cpp is an open-source program for running (not training) large language models. It is written in plain C and C++ with no heavy dependencies: no Python, no PyTorch, no CUDA required. It builds on ggml, a small tensor library from the same author that provides the maths (matrix multiplications, attention, quantized formats) and the hardware backends.

It started in March 2023, when Georgi Gerganov showed Meta's newly released LLaMA model running on a MacBook CPU. The project grew quickly into one of the most popular open-source AI projects, supporting a large range of model families (Llama, Mistral, Qwen, Gemma, Phi, DeepSeek and many more), not just LLaMA.

Think of it like a pocket-sized edition of a book A publisher's hardback (the original model) is beautiful but too heavy to carry. A pocket edition uses thinner paper and smaller print (quantization), binds everything in one volume (GGUF) and opens instantly to any page (memory mapping). You lose a little print quality, but you can read it anywhere. llama.cpp is the pocket edition plus the reader.

Our running example: we want a private assistant that summarises our notes, running on a 16 GB laptop with a small 4 GB GPU, using an 8-billion-parameter model.

Why we needed llama.cpp

Before llama.cpp, running an LLM usually meant Python, PyTorch and a large NVIDIA GPU with enough memory to hold the model in 16-bit precision. That excluded most laptops, every Mac without an NVIDIA card, phones, and small servers. Many people also wanted to run models locally for privacy, offline use, cost, or simply to experiment.

  • Privacy: our notes never leave the laptop.
  • No per-token cost and no rate limits.
  • Offline: works on a plane or in a secure network.
  • Control: pick the exact model version and settings, and nothing changes underneath you.

A quick refresher on what an LLM is

An LLM reads text as tokens (word pieces mapped to numbers). Each token becomes a vector through the embedding table, then passes through a stack of identical transformer layers (32 for a typical 8B model). Each layer has an attention part (each token looks at earlier tokens) and a feed-forward part (a per-token transformation). Almost all of the work is multiplying vectors by big weight matrices. At the end, the output head gives a probability for every possible next token, and we pick one.

Generation repeats: append the chosen token and run again. Two phases matter for speed. Prefill processes the whole prompt in one batch, which is heavy on arithmetic. Decode produces one token at a time, and each step must read every weight in the model once. To avoid recomputing old tokens, each layer stores their keys and values in the KV cache.

The real problem: models are too big to fit

Weights dominate memory. An 8B model has about 8 billion weights. In 16-bit floats (2 bytes each) that is about 16 GB, our entire laptop RAM, with nothing left for the operating system or the KV cache. A 70B model would need about 140 GB.

And even when a model fits, decode speed has a hard ceiling: every new token requires reading all the weights from memory. Speed is roughly memory bandwidth ÷ model size. Laptop RAM might move around 50–100 GB/s; a 16 GB model would crawl along at a few tokens per second. Both problems have the same cure: make the weights smaller.

Pause and think: If we quantize our 16 GB model to about 4.9 GB, by roughly how much can decode speed improve on the same laptop?

Roughly 3×, since 16 / 4.9 ≈ 3.3. Decode reads all weights once per token, so a model a third the size can be read about three times faster. It also now fits in RAM with room to spare.

The first big idea: quantization, and names like Q4_K_M

Quantization stores each weight in fewer bits by rounding it to a small grid of values, with a scale per small block of weights to convert back. llama.cpp does this block-wise: in the simple Q8_0 format, every 32 weights share one 16-bit scale and each weight becomes an 8-bit integer. In Q4_0, each weight becomes a 4-bit integer. The newer k-quants use 256-weight super-blocks with their own quantized sub-block scales.

Decoding llama.cpp quantization names
NameMeaningTypical trade-off
Q8_08-bit, simple blocks of 32 with one scale (_0 = scale only)Very close to full quality, about half the size of F16
Q6_K6-bit k-quantNear Q8 quality, smaller
Q5_K_M5-bit k-quant, medium mixHigh quality, moderate size
Q4_K_M4-bit k-quant, medium mix (some sensitive tensors kept at 6-bit)The popular default balance
Q3_K_M / Q2_K3- and 2-bit k-quantsSmall, with clearly visible quality loss
IQ4_XS, IQ3_M …"i-quants", usually guided by an importance matrixBetter quality at very low bits, sometimes slower on CPU

Crucially, llama.cpp computes directly on the quantized blocks. To multiply a 4-bit weight row by an activation vector, it quantizes the activation vector to 8-bit blocks too, does integer multiply-adds inside each block, and applies the two scales once per block. This avoids ever expanding the whole model back to 16-bit in memory.

The GGUF file and memory mapping

llama.cpp stores models in GGUF, a single binary file that contains a header, metadata (architecture, layer count, context length, the tokenizer's vocabulary and the chat template), a table describing each tensor (name, shape, quantization type, offset) and finally the aligned weight data. One file is all the engine needs: everything packed in one box.

Because the weight data is aligned, llama.cpp can load it with memory mapping (mmap): the operating system makes the file appear to be in memory and reads pages from disk only when they are first used. Start-up takes moments instead of copying gigabytes, repeated runs hit the OS page cache, and several processes can share one copy of the weights. (On some setups, an option can also lock the pages in RAM so the OS never swaps them out.)

GPU layers are copied Memory mapping helps most for weights that stay on the CPU. Layers offloaded to a GPU must be copied into GPU memory, so they cost VRAM no matter how the file was opened.

Squeezing speed out of the CPU

A CPU has few cores compared with a GPU, but each core is powerful. llama.cpp uses every trick available:

  • SIMD instructions: one instruction processes many numbers at once, using AVX2 / AVX-512 on x86 chips and NEON on ARM (phones, Apple silicon). Inner loops are hand-tuned for each quantization format.
  • Integer maths on small blocks: 8-bit multiply-adds are cheap and pack more values per instruction than 32-bit floats.
  • Multithreading: the rows of each matrix are split across CPU threads (the -t / --threads option).
  • Cache-friendly layout: small blocks keep data in fast CPU caches while it is used.
  • Fewer bytes to read: the biggest win, because decode is memory-bound. Quantization directly raises tokens per second.

llamacpp_math.py

import numpy as np
# 1) Q8_0-style dot product: integer multiply-adds per 32-value block, then one float multiply.
def q8_0(x):
blocks = x.reshape(-1, 32)
scale = np.abs(blocks).max(axis=1) / 127
return np.round(blocks / scale[:, None]).astype(np.int8), scale
rng = np.random.default_rng(42)
w, a = rng.normal(size=4096), rng.normal(size=4096)      # one weight row, one activation
qw, sw = q8_0(w)
qa, sa = q8_0(a)
int_dots = (qw.astype(np.int32) * qa.astype(np.int32)).sum(axis=1)   # what SIMD units do
approx = float((int_dots * sw * sa).sum())
print(f"exact dot {w @ a:.3f} | Q8_0 dot {approx:.3f}")
# 2) Speed estimate: each new token reads (almost) all weights once.
model_gb, n_layers = 4.9, 32                              # 8B model in Q4_K_M (approx.)
for name, bw in [("laptop DDR5 RAM", 80), ("Apple M-series Max", 400), ("desktop GPU", 900)]:
print(f"{name:20s} ~{bw:4d} GB/s -> at most ~{bw / model_gb:5.1f} tokens/s")
# 3) Partial GPU offload (-ngl): put as many layers as fit in VRAM on the GPU.
vram_free_gb, cpu_bw, gpu_bw = 3.0, 80, 300               # a small laptop GPU (illustrative)
per_layer = model_gb / n_layers
for ngl in (0, 8, 16, int(vram_free_gb // per_layer), n_layers):
if ngl * per_layer > vram_free_gb:
print(f"-ngl {ngl:2d}: does not fit in {vram_free_gb} GB VRAM"); continue
t = (n_layers - ngl) * per_layer / cpu_bw + ngl * per_layer / gpu_bw
print(f"-ngl {ngl:2d}: {ngl * per_layer:4.2f} GB on GPU -> ~{1 / t:5.1f} tokens/s")

Output:

exact dot -132.269 | Q8_0 dot -131.969
laptop DDR5 RAM      ~  80 GB/s -> at most ~ 16.3 tokens/s
Apple M-series Max   ~ 400 GB/s -> at most ~ 81.6 tokens/s
desktop GPU          ~ 900 GB/s -> at most ~183.7 tokens/s
-ngl  0: 0.00 GB on GPU -> ~ 16.3 tokens/s
-ngl  8: 1.23 GB on GPU -> ~ 20.0 tokens/s
-ngl 16: 2.45 GB on GPU -> ~ 25.8 tokens/s
-ngl 19: 2.91 GB on GPU -> ~ 28.9 tokens/s
-ngl 32: does not fit in 3.0 GB VRAM

Sharing the work with the GPU

Through ggml, llama.cpp has backends for many kinds of hardware: CUDA (NVIDIA), Metal (Apple silicon), Vulkan (most GPUs), HIP (AMD), SYCL (Intel) and others. You can put the whole model on the GPU if it fits, or only part of it. The option -ngl N (--n-gpu-layers) sends the first N transformer layers' weights to the GPU, and the rest stay in system RAM and run on the CPU.

This partial offload is a key reason llama.cpp works on everyday hardware. On our laptop, the 4 GB GPU has about 3 GB free, enough for 19 of 32 layers. In the illustrative numbers above, that raises the speed from about 16 to about 29 tokens per second. On Apple silicon, the CPU and GPU share one pool of fast unified memory, so the whole model can usually go on the GPU via Metal.

Common mistakes Setting -ngl higher than VRAM allows causes out-of-memory errors or heavy slowdowns; leave room for the KV cache, which also lives on the GPU for offloaded layers and grows with context length. Using more CPU threads than physical cores often makes things slower, not faster. And picking the biggest model that technically fits leaves no memory for long contexts: plan for weights plus KV cache plus the OS.

The full journey of running a prompt

Running it yourself

  1. Get a model: Download a GGUF file, for example an 8B instruct model in Q4_K_M (about 5 GB).
  2. Chat in the terminal: Run llama-cli -m model.gguf -ngl 19 and type a message.
  3. Or start a server: Run llama-server -m model.gguf -ngl 19 --port 8080. It offers a simple web UI and an OpenAI-compatible HTTP API, so existing client code can point at it.
  4. Tune: Adjust context size (-c), threads (-t) and layer offload until speed and memory are balanced.

Where llama.cpp is used

llama.cpp and its ggml library sit underneath much of the local-AI world. Ollama was built on llama.cpp's engine, LM Studio and GPT4All use it for GGUF models, Mozilla's llamafile packages it into single executables, and many mobile and desktop apps embed it. Hugging Face hosts a large number of GGUF files made for it.

Our notes assistant, solved We run llama-server with an 8B Q4_K_M model and 19 layers offloaded. Our note-taking app calls its OpenAI-compatible endpoint on localhost. Summaries stream at a comfortable reading speed and no data leaves the laptop.

Worked example, step by step

Our earlier estimate put 19 layers on the GPU, counting only the weights. The warning box said to leave room for the KV cache too. Let us redo the sum with the cache included. We need one assumption about the model's shape: 32 layers, 8 key/value heads of size 128 and a 16-bit cache, which gives about 128 KiB of KV cache per token. Check the real shape of your own model; this one is illustrative.

How many layers fit in 3 GB of free VRAM?

  1. Weights per layer: 4.9 GB ÷ 32 layers ≈ 0.153 GB per layer.
  2. KV cache for the whole context: With a context of 8,192 tokens: 8,192 × 128 KiB ≈ 1.07 GB, spread evenly over 32 layers. That is about 0.034 GB per layer.
  3. Cost of one offloaded layer: The layer's weights and its share of the cache both sit on the GPU: 0.153 + 0.034 ≈ 0.187 GB.
  4. Divide and round down: 3.0 ÷ 0.187 ≈ 16.0, so 16 layers fit, not 19. With 19 layers we would need 19 × 0.187 ≈ 3.55 GB.
  5. Ask for a longer context: At 32,768 tokens the cache is about 4.3 GB, or 0.134 GB per layer. One offloaded layer now costs 0.287 GB, and only 10 layers fit.
  6. Check the RAM side: With 16 layers on the GPU, the other 16 layers stay in RAM: about 2.45 GB of weights plus 0.54 GB of cache. That is comfortable on a 16 GB laptop.
Layers that fit in 3 GB of VRAM as the context grows (illustrative model shape)
ContextKV cache in totalCost per offloaded layerLargest -ngl that fits
Weights onlyignored0.153 GB19
8,192 tokens≈ 1.07 GB≈ 0.187 GB16
32,768 tokens≈ 4.3 GB≈ 0.287 GB10

The model file did not change between the rows. Only the context size did. So -ngl and -c have to be chosen together: a setting that works for short chats can run out of GPU memory when we later raise the context, and the fix is to offload fewer layers or to shorten the context.

Practice: try it yourself

Memory mapping lets a model start even when RAM is tight, because the operating system loads pieces of the file on first use and can drop them again. What does that do to speed? We will simulate it. The weights are 32 chunks of 150 MB. Each token reads every chunk once, in order. A chunk in RAM is fast, and a chunk that has to come from disk is slow. When RAM is full, the least recently used chunk is dropped. The speeds are illustrative.

practice_page_cache.py

# What memory mapping does when the model is bigger than free RAM: a toy page cache.
from collections import OrderedDict
N_LAYERS, LAYER_MB = 32, 150           # about 4.8 GB of weights, in 32 equal chunks
DISK_MB_S, RAM_MB_S = 2000, 80000      # illustrative: SSD read speed, RAM bandwidth
def run(free_ram_mb, tokens=5):
cache = OrderedDict()              # chunks now held in RAM, least recently used first
capacity = free_ram_mb // LAYER_MB # how many chunks fit in free RAM
times = []
for _ in range(tokens):
disk_reads = 0
for layer in range(N_LAYERS):  # each token reads every layer once, in order
if layer in cache:
cache.move_to_end(layer)           # already in RAM: fast
else:
disk_reads += 1                    # not in RAM: read it from disk
cache[layer] = True
if len(cache) > capacity:
cache.popitem(last=False)      # RAM full: drop the oldest chunk
t = disk_reads * LAYER_MB / DISK_MB_S + N_LAYERS * LAYER_MB / RAM_MB_S
times.append(t)
return times
for ram in (8000, 4800, 4000, 2400):
times = run(ram)
print(f"free RAM {ram} MB: first token {times[0]:.2f} s, "
f"later tokens {times[-1]:.2f} s ({1 / times[-1]:.1f} tokens/s)")

Output:

free RAM 8000 MB: first token 2.46 s, later tokens 0.06 s (16.7 tokens/s)
free RAM 4800 MB: first token 2.46 s, later tokens 0.06 s (16.7 tokens/s)
free RAM 4000 MB: first token 2.46 s, later tokens 2.46 s (0.4 tokens/s)
free RAM 2400 MB: first token 2.46 s, later tokens 2.46 s (0.4 tokens/s)

The first token always takes about 2.5 s, because the whole file must come from disk once. With 4,800 MB free or more, later tokens run at RAM speed. With 4,000 MB free, every token is as slow as the first. Now change it:

  • Set LAYER_MB = 100, a smaller quantization of about 3.2 GB, and keep 4,000 MB free. Predict the speed of later tokens before you run it.
  • Add 4,650 to the list of RAM sizes: room for 31 of the 32 chunks. Predict how many disk reads each later token needs when we are just one chunk short.
  • Set DISK_MB_S = 200, a slow external drive. Predict the first-token time when RAM is plentiful, and say which line of the output shows that later tokens do not care.

Pause and think: In the worked example, raising the context from 8,192 to 32,768 tokens cut the layers we can offload from 16 to 10, although the model file is the same. Why?

Each offloaded layer brings its share of the KV cache onto the GPU, and the cache grows with the context size. At 32,768 tokens a layer's cache share (about 0.134 GB) is almost as large as its weights (0.153 GB), so every layer costs nearly twice as much VRAM and fewer fit.

Pause and think: In the practice run, 4,000 MB of free RAM is only about 17% short of the 4,800 MB the model needs. Why is the slowdown about 40 times and not about 17%?

Every token reads all the chunks in the same order. By the time we come back to a chunk, it is the one that was used longest ago, so it has already been dropped to make room. Every read then comes from disk, at disk speed instead of RAM speed. Being a little short of RAM is a cliff, not a gentle slope. The fix is a smaller file or more free memory.

Key takeaways

  • llama.cpp is a dependency-light C/C++ inference engine built on the ggml tensor library.
  • Decode speed ≈ memory bandwidth ÷ model size, so quantization (Q4_K_M, Q8_0 …) makes models both fit and run faster.
  • GGUF packs weights, tokenizer and settings into one file; memory mapping makes loading near-instant.
  • SIMD, integer block maths and threads speed up CPUs; -ngl offloads as many layers as fit onto the GPU.
  • It powers much of local AI (Ollama, LM Studio, llamafile) and offers an OpenAI-compatible server.

Key terms

  • llama.cpp: An open-source C/C++ engine for running LLMs efficiently on CPUs, Apple silicon and consumer GPUs.
  • ggml: The tensor library underneath llama.cpp that provides quantized maths and hardware backends.
  • SIMD: Single Instruction, Multiple Data: CPU instructions (AVX, NEON) that process many numbers at once.
  • Layer offload (-ngl): Placing the weights of some transformer layers on the GPU while the rest run on the CPU.
  • Memory bandwidth: How many bytes per second a processor can read from memory; the main limit on decode speed.
  • Unified memory: A single memory pool shared by CPU and GPU, as on Apple silicon, so the GPU can use most of system RAM.

← 13.13 GGUF: The File Format Powering Local LLM Inference · 13.15 vLLM: High-Throughput Serving with PagedAttention →