Modern AI Engineering

Lesson 13.13 · 23 min

GGUF: The File Format Powering Local LLM Inference

Why can you download one file called something like model-Q4_K_M.gguf, double-click it in a desktop app, and be chatting with an LLM a few seconds later?

In short: GGUF is a single-file format for storing an LLM: a small header, a dictionary of metadata (architecture, tokenizer, chat template), a table describing each tensor, and then the (usually quantized) weights, aligned so they can be memory-mapped. Names like Q4_K_M describe how the weights were quantized. Because everything is in one self-describing file that loads almost instantly, GGUF became the standard for running models locally with llama.cpp, Ollama and LM Studio.

What is a model and what are weights

A language model is a big mathematical function. Text goes in as numbers (token IDs), passes through many layers of matrix multiplications, and comes out as probabilities for the next token. The numbers inside those matrices are the weights (also called parameters). Training is the process of finding good weights; after training they stay fixed.

The weights are grouped into tensors: named multi-dimensional arrays such as blk.0.attn_q.weight (the query projection of layer 0) or token_embd.weight (the embedding table). An 8B model has hundreds of tensors holding about 8 billion numbers in total. To use the model on another computer we must save all those tensors plus everything needed to interpret them: the architecture, the sizes, and the tokenizer.

Think of it like a flat-pack furniture box The weights are the boards and screws. On their own they are useless: you also need the instruction sheet (architecture and hyperparameters) and the label that says which bag is which (tensor names and shapes). A GGUF file is one box with the parts, the instructions and the labels packed together.

What is local inference

Inference means using a trained model to produce outputs. Local inference means running it on your own machine (a laptop, desktop, phone or private server) instead of calling a cloud API. People do it for privacy (data never leaves the device), offline use, no per-token cost, and full control over the model version.

The obstacle is size. An 8B model in 16-bit precision is about 16 GB, more than many laptops can spare. So local inference relies on quantization (storing weights in fewer bits) and on a file format and engine that are efficient on ordinary CPUs and GPUs. Our running example: we want to run an 8B chat model on a 16 GB laptop.

The problem before GGUF

Models on Hugging Face are usually shared as a folder: weight files in PyTorch or safetensors format, plus config.json (architecture), tokenizer.json, tokenizer_config.json (with the chat template) and more. That works well for Python training code on GPUs, but is awkward for a lightweight C/C++ engine and for non-technical users who just want "the model".

The llama.cpp project (next lesson) first used its own formats named after its tensor library, GGML (later variants GGMF and GGJT). Their hyperparameters were stored as a fixed, unnamed list of numbers. Every time a new model architecture or feature appeared, the layout had to change, old files broke, and loaders needed hard-coded guesses. There was no clean place to store the tokenizer or new settings.

  • Breaking changes: new versions of the engine could not load old files, and vice versa.
  • Missing information: tokenizer details and special settings lived in separate files or code.
  • No extensibility: adding a new field meant inventing a new format version.

What is GGUF

GGUF is the binary file format that the ggml / llama.cpp project introduced in August 2023 to replace those older formats. Its name builds on GGML; you will see different expansions of the letters online, but in practice everyone just says "GGUF". Its design goals are simple:

  • Single file: weights, architecture, tokenizer and settings together.
  • Self-describing: all settings are named key–value pairs, so a reader knows what each value means.
  • Extensible: new keys can be added without breaking older readers, which simply ignore keys they do not know.
  • Fast to load: tensor data is aligned so it can be memory-mapped straight from disk.
  • Quantization-friendly: each tensor records its own data type, from 32-bit floats down to 2-bit block formats.

What is stored inside a GGUF file

A GGUF file is laid out in four parts, one after another. All numbers are fixed-size and stored little-endian by default, which makes the file the same on every machine.

The code below writes a tiny file with this exact layout (one 2×4 float tensor and three metadata keys), reads it back by walking the bytes, and memory-maps the tensor. Real GGUF files are the same structure, just with hundreds of tensors and thousands of metadata values.

gguf_demo.py

import struct
import numpy as np
def s(text):                                   # GGUF string = uint64 length + UTF-8 bytes
b = text.encode(); return struct.pack("<Q", len(b)) + b
# ---- write a tiny but valid-layout GGUF v3 file ----
meta = [("general.architecture", 8, s("llama")), ("general.name", 8, s("tiny-demo")),
("llama.context_length", 4, struct.pack("<I", 4096))]   # 8 = STRING, 4 = UINT32
weights = np.arange(8, dtype=np.float32).reshape(2, 4)
out = b"GGUF" + struct.pack("<IQQ", 3, 1, len(meta))            # magic, version, #tensors, #kv
for key, vtype, val in meta:
out += s(key) + struct.pack("<I", vtype) + val
out += s("blk.0.ffn_up.weight") + struct.pack("<I", 2)           # tensor name, n_dims
out += struct.pack("<QQIQ", 4, 2, 0, 0)                          # dims (fastest first), type F32=0, offset
out += b"\0" * (-len(out) % 32)                                  # pad to 32-byte alignment
data_start = len(out)
open("tiny.gguf", "wb").write(out + weights.tobytes())
# ---- read it back ----
buf = open("tiny.gguf", "rb").read()
magic, (ver, n_t, n_kv) = buf[:4], struct.unpack_from("<IQQ", buf, 4)
print("magic:", magic, "version:", ver, "tensors:", n_t, "metadata keys:", n_kv)
pos = 24
def read_str():
global pos
n, = struct.unpack_from("<Q", buf, pos); pos += 8 + n
return buf[pos - n:pos].decode()
for _ in range(n_kv):
key = read_str(); vtype, = struct.unpack_from("<I", buf, pos); pos += 4
if vtype == 8: val = read_str()
else: val, = struct.unpack_from("<I", buf, pos); pos += 4
print(f"  {key} = {val}")
name = read_str(); nd, = struct.unpack_from("<I", buf, pos)
dims = struct.unpack_from(f"<{nd}Q", buf, pos + 4)
print("tensor:", name, "dims:", dims, "data starts at byte", data_start)
w = np.memmap("tiny.gguf", dtype=np.float32, mode="r", offset=data_start, shape=(2, 4))
print("memory-mapped weights:\n", w)
# ---- bits per weight of common block layouts (bytes per block / weights per block) ----
for qname, nbytes, nweights in [("Q8_0", 34, 32), ("Q4_0", 18, 32), ("Q4_K", 144, 256), ("Q6_K", 210, 256)]:
print(f"{qname}: {nbytes * 8 / nweights:.4g} bits/weight")

Output:

magic: b'GGUF' version: 3 tensors: 1 metadata keys: 3
general.architecture = llama
general.name = tiny-demo
llama.context_length = 4096
tensor: blk.0.ffn_up.weight dims: (4, 2) data starts at byte 224
memory-mapped weights:
[[0. 1. 2. 3.]
[4. 5. 6. 7.]]
Q8_0: 8.5 bits/weight
Q4_0: 4.5 bits/weight
Q4_K: 4.5 bits/weight
Q6_K: 6.562 bits/weight

What is quantization (in GGUF)

Quantization stores each weight with fewer bits by rounding it onto a small grid of allowed values, with a scale that converts the grid back to real numbers (the previous lesson covers this in depth). GGUF uses block-wise quantization: weights are cut into small blocks, and each block gets its own scale, so one large weight only affects its own block.

How a Q8_0 block is stored

  1. Take 32 weights: Q8_0 works on blocks of 32 consecutive weights from one row of a tensor.
  2. Find the scale: scale = max |weight| / 127, stored as one 16-bit float (2 bytes).
  3. Round each weight: Store round(weight / scale) as a signed 8-bit integer: 32 bytes.
  4. Count the cost: 2 + 32 = 34 bytes per 32 weights = 8.5 bits per weight. Q4_0 does the same with 4-bit integers: 2 + 16 = 18 bytes, 4.5 bits per weight.

The newer k-quants (the "K" in names like Q4_K) use super-blocks of 256 weights split into smaller sub-blocks. Each sub-block gets its own scale (and minimum), and those scales are themselves quantized to 6 bits, with one 16-bit scale for the whole super-block. This two-level trick gives fine-grained scales at low overhead: Q4_K costs 144 bytes per 256 weights, again 4.5 bits per weight, but with better accuracy than Q4_0.

Understanding quantization names like Q4_K_M

When you browse GGUF files you see a menu of suffixes. They decode like this:

Reading a GGUF quantization name
PartMeaningExample
QQuantized weightsQ4_K_M
NumberApproximate bits per weight for most tensors4 → about 4 bits (+ scale overhead)
_0 / _1Older simple block formats: _0 stores a scale only; _1 stores a scale and a minimumQ4_0, Q4_1, Q8_0
Kk-quant: super-blocks with quantized sub-block scalesQ3_K, Q4_K, Q5_K, Q6_K
_S / _M / _LSmall / Medium / Large mix: how many sensitive tensors get a higher-bit typeQ4_K_M keeps some tensors at Q6_K
IQNewer "i-quants", usually built with an importance matrix from calibration text; good at very low bitsIQ2_XXS, IQ3_M, IQ4_XS
F16 / BF16 / F32Unquantized floating pointUsed for reference or small tensors

The S/M/L mix matters because not all tensors are equally sensitive. In llama.cpp's Q4_K_M recipe, most tensors use Q4_K, but some of the most sensitive ones (for example part of the attention value and feed-forward down-projection tensors) use Q6_K. The result averages a little under 5 bits per weight: an 8B model in Q4_K_M is roughly 4.9 GB.

Pause and think: Our laptop has 16 GB of RAM and we also want the browser open. Which file is the safer choice for an 8B model: Q8_0 or Q4_K_M, and why?

Q4_K_M. At about 4.9 GB it leaves plenty of room for the KV cache, the operating system and other apps, while Q8_0 at about 8.5 GB would squeeze memory. Q4_K_M is widely considered a good quality/size balance; Q8_0 is closer to full quality if memory allows.

How GGUF loads fast with memory mapping

The classic way to load a model is to read the whole file and copy it into a freshly allocated block of memory. For a 5 GB file that is slow and needs the full 5 GB up front. GGUF is designed for memory mapping (mmap) instead.

With mmap, the operating system makes the file appear to be in memory without actually reading it. When the engine first touches a tensor, the OS loads just those pages from disk (or from its page cache, if the file was used recently). Because GGUF aligns every tensor's data to a fixed boundary, the engine can point directly at the bytes in the file and use them as-is, with no parsing or copying.

  • Near-instant start: loading takes milliseconds; data streams in as layers are used.
  • Fast restarts: the second launch hits the OS page cache, so it is often far faster than the first.
  • Shared memory: two processes using the same file share one copy in RAM.
  • Graceful limits: if RAM is tight, the OS can drop pages and re-read them from disk later (slower, but it runs).

Common mistake Judging memory use by what your task manager shows right after loading. Mapped pages may not be counted until they are touched, so usage climbs during the first prompt. Also, storing GGUF files on a slow network or USB drive makes that first pass slow. And if you offload layers to a GPU, those tensors are copied into GPU memory, so mmap savings apply mainly to the part that stays on the CPU.

Why GGUF is cross-platform and extensible

Cross-platform: GGUF uses fixed-size integers and floats with a defined byte order, so the same file works on Windows, macOS, Linux, Android and on x86 or ARM chips. The engine reads the metadata and picks the right compute kernels for the local hardware. You never need a "Mac version" of a model.

Extensible: every setting is a named key with a typed value. When a new architecture or feature needs new information (say, a new type of position encoding), the converter just writes new keys such as <arch>.rope.scaling.type. Older readers skip keys they do not understand instead of crashing. The format version only changes for structural changes; version 3 is current.

Self-contained: because the tokenizer vocabulary, merges, special token IDs and even the chat template (the text pattern that wraps user and assistant messages) are stored in the metadata, one file is enough to chat correctly.

GGUF in the real world

GGUF is the de facto standard for local LLMs. Hugging Face hosts a very large number of GGUF files (often in several quantization levels per model) and can show a file's metadata and tensor list in the browser. Tools that run GGUF models include llama.cpp itself, Ollama, LM Studio, GPT4All, Jan and KoboldCpp. Some other engines can also import GGUF weights.

The typical workflow

  1. Start from a published model: Download the original weights (safetensors + configs) from Hugging Face.
  2. Convert: Run llama.cpp's convert_hf_to_gguf.py to produce an F16 or BF16 GGUF with all metadata and tokenizer data.
  3. Quantize: Run llama-quantize model-f16.gguf model-Q4_K_M.gguf Q4_K_M. Optionally compute an importance matrix first for i-quants.
  4. Run: Load the file in llama.cpp, Ollama or LM Studio. Most users skip steps 1–3 and download a ready-made GGUF.

Our laptop, solved We download an 8B instruct model as Q4_K_M (about 4.9 GB), open it in LM Studio or Ollama, and it starts in seconds thanks to mmap. The chat template inside the file makes sure our messages are formatted the way the model was trained.

When not to use GGUF If you are fine-tuning or training, stay with safetensors in BF16. If you are serving many users on data-centre GPUs, engines like vLLM, SGLang or TensorRT-LLM with their own formats (FP8, AWQ, GPTQ) usually give higher throughput. GGUF shines for local, single-user or small-scale inference.

Worked example, step by step

gguf_demo.py printed that the tensor data starts at byte 224. Where does that number come from? Nothing in the file is hidden, so we can add it up by hand. We need two rules from the code: a string costs 8 bytes for its length plus one byte per character, and a value type or a 32-bit number costs 4 bytes.

Adding up the tiny file

  1. Header: 4 magic bytes + 4 for the version + 8 for the tensor count + 8 for the metadata count = 24 bytes.
  2. First key: general.architecture has 20 characters: 8 + 20 = 28. Then 4 for the value type. The value llama is a string: 8 + 5 = 13. Total 45 bytes.
  3. Second and third keys: general.name = tiny-demo: (8 + 12) + 4 + (8 + 9) = 41 bytes. llama.context_length = 4096: (8 + 20) + 4 + 4 = 36 bytes.
  4. Tensor info: The name blk.0.ffn_up.weight has 19 characters: 8 + 19 = 27. Then 4 for the number of dimensions, 2 × 8 for the two sizes, 4 for the data type and 8 for the offset. Total 59 bytes.
  5. Pad to the boundary: 24 + 45 + 41 + 36 + 59 = 205 bytes. The next multiple of 32 is 224, so the writer adds 19 zero bytes.
  6. Tensor data: Eight 32-bit floats take 32 bytes, from byte 224 to byte 255. The whole file is 256 bytes.
Byte map of tiny.gguf
PartBytesSize
Header0 – 2324
Metadata (3 keys)24 – 145122
Tensor info (1 tensor)146 – 20459
Padding205 – 22319
Tensor data224 – 25532

Two lessons hide in this sum. First, a reader never has to guess. Every string says how long it is and every value says what type it is, so the reader always knows how many bytes come next. That is also how it can step over a key it does not know. Second, in this toy file the description is larger than the data: 205 bytes against 32. In a real model it is the other way round by a huge margin: the weights are gigabytes, and the metadata, even with the whole vocabulary inside, is tiny next to them.

Practice: try it yourself

The lesson described how a Q8_0 block is stored: one 16-bit scale followed by 32 small integers. We will pack blocks like that into real bytes and unpack them again. Then we do the same with a simplified 4-bit block that squeezes two weights into every byte. The 4-bit packing here is a teaching version, not the exact layout of any GGUF type.

practice_block_packing.py

import struct
import numpy as np
rng = np.random.default_rng(5)
BLOCK = 32
weights = rng.normal(0, 0.05, size=2 * BLOCK).astype(np.float32)   # two blocks
def pack8(block):                      # 8-bit block: one FP16 scale + one byte per weight
scale = np.float16(np.abs(block).max() / 127)
q = np.round(block / np.float32(scale)).clip(-127, 127).astype(np.int8)
return struct.pack("<e", scale) + q.tobytes()
def unpack8(raw):
scale, = struct.unpack_from("<e", raw, 0)
return np.frombuffer(raw, dtype=np.int8, offset=2) * np.float32(scale)
def pack4(block):                      # simplified 4-bit block: two weights per byte
scale = np.float16(np.abs(block).max() / 7)
q = (np.round(block / np.float32(scale)).clip(-7, 7) + 8).astype(np.uint8)
return struct.pack("<e", scale) + (q[0::2] | (q[1::2] << 4)).tobytes()
def unpack4(raw):
scale, = struct.unpack_from("<e", raw, 0)
b = np.frombuffer(raw, dtype=np.uint8, offset=2)
q = np.empty(2 * len(b), dtype=np.int16)
q[0::2], q[1::2] = b & 15, b >> 4  # split each byte back into two 4-bit values
return (q - 8) * np.float32(scale)
for name, pack, unpack in [("8-bit", pack8, unpack8), ("4-bit", pack4, unpack4)]:
blocks = [pack(weights[i:i + BLOCK]) for i in range(0, len(weights), BLOCK)]
back = np.concatenate([unpack(b) for b in blocks])
size = sum(len(b) for b in blocks)
print(f"{name}: {len(blocks[0])} bytes per block, {size * 8 / len(weights)} bits/weight, "
f"mean error {np.abs(back - weights).mean():.5f}")
print("as FP32:", weights.nbytes, "bytes | as FP16:", weights.nbytes // 2, "bytes")

Output:

8-bit: 34 bytes per block, 8.5 bits/weight, mean error 0.00020
4-bit: 18 bytes per block, 4.5 bits/weight, mean error 0.00324
as FP32: 256 bytes | as FP16: 128 bytes

The byte counts match the lesson exactly: 34 bytes and 8.5 bits per weight for the 8-bit block, 18 bytes and 4.5 bits per weight for the 4-bit one. The 4-bit error is about 16 times larger, because its grid has 15 levels instead of 255. Now change it:

  • Set BLOCK = 16. Predict the bytes per block and the bits per weight for both formats before you run it.
  • Add weights[5] = 3.0 right after weights is created, and print the mean error of each block separately. Predict which block gets worse and which is untouched.
  • Store the scale as a 32-bit float: use np.float32 for the scale, "<f" in struct, and offset=4. Predict the new bits per weight. Is the extra precision of the scale worth it?

Pause and think: We add one more metadata key to tiny.gguf, and its entry takes 30 bytes. At which byte does the tensor data start now?

At byte 256. The part before the padding grows from 205 to 235 bytes, and the next multiple of 32 after 235 is 256. The padding is not a fixed amount: it is whatever is needed to reach the boundary, here 21 bytes instead of 19.

Pause and think: An older reader meets a metadata key it has never heard of. What in the layout lets it skip that entry and still find the next one?

The key is a string that starts with its own length, and it is followed by a value type. The type tells the reader how big the value is, or, for a string, that a length comes first. So the reader can compute where the entry ends without knowing what the key means.

Key takeaways

  • GGUF is one self-describing file: header, typed key–value metadata, tensor infos, then aligned tensor data.
  • Metadata includes architecture, hyperparameters, tokenizer and chat template, so no extra config files are needed.
  • Weights are block-quantized; names like Q4_K_M mean about 4 bits, k-quant super-blocks, medium mix.
  • Aligned data lets engines memory-map the file for near-instant loading and shared memory.
  • GGUF is the standard for local inference (llama.cpp, Ollama, LM Studio); safetensors remains the training format.

Key terms

  • GGUF: The single-file binary format from the ggml / llama.cpp project for storing models with metadata and quantized tensors.
  • Tensor: A named multi-dimensional array of numbers, such as one weight matrix of the model.
  • Metadata: Named, typed key–value pairs in the file that describe the model, its settings and its tokenizer.
  • Block quantization: Quantizing weights in small groups, each with its own scale, so one large value affects only its block.
  • k-quant: GGUF quantization types (Q2_K … Q6_K) using 256-weight super-blocks with quantized sub-block scales.
  • Memory mapping (mmap): An OS feature that makes a file appear in memory and loads its pages from disk only when they are accessed.

← 13.12 Model Quantization: Shrinking Weights Without Breaking Outputs · 13.14 llama.cpp: Running Large Models on Consumer Hardware →