Modern AI Engineering

Lesson 13.12 · 27 min

Model Quantization: Shrinking Weights Without Breaking Outputs

How can a 7-billion-parameter model that needs 28 GB in full precision squeeze into 4 GB and still give nearly the same answers?

In short: Quantization stores a model's numbers with fewer bits, for example 8-bit or 4-bit integers instead of 32-bit floats, using a scale (and sometimes a zero-point) to map between the two. That cuts memory by 4–8× and speeds up the memory-bound parts of inference. The art is choosing the right granularity, method and bit width so that rare large values (outliers) do not destroy accuracy.

What is model quantization?

A neural network is mostly a huge collection of numbers called weights (or parameters). A 7B model has about 7 billion of them. During training, each weight is usually stored as a 32-bit or 16-bit floating-point number. Quantization means storing those numbers (and sometimes the intermediate values computed from them, called activations) with fewer bits, such as 8 or 4.

Fewer bits means fewer possible values. A 4-bit number can only take 16 different values. So quantization is a form of rounding: each original weight is snapped to the nearest value on a coarse grid. We accept a small rounding error in exchange for a much smaller, faster model.

Think of it like rounding prices A shop could price items to the exact tenth of a cent, but tags would be long and hard to read. Rounding to the nearest dollar makes tags tiny, and for most purchases the total barely changes. Quantization rounds a model's weights to a coarse grid; the trick is choosing the grid so the "total" (the model's answers) barely changes.

Our running example: we want to run a 7B chat assistant on a laptop with a 6 GB GPU. In 16-bit it needs about 14 GB just for weights. Quantization is how we make it fit.

How numbers use bits (FP32, INT8, INT4)

A floating-point number stores a sign, an exponent (which sets the scale, like scientific notation) and a mantissa (the significant digits). This lets it cover tiny and huge values. An integer format stores whole numbers evenly spaced, with no exponent.

Common number formats in LLMs
FormatBitsLayout / rangeBytes per weight
FP32321 sign, 8 exponent, 23 mantissa bits; about ±3.4×10³⁸4
FP16161 sign, 5 exponent, 10 mantissa; max about 65,5042
BF16161 sign, 8 exponent, 7 mantissa; FP32's range with less precision2
FP8 (E4M3)81 sign, 4 exponent, 3 mantissa; used on recent GPUs1
INT88Whole numbers −128 … 127 (256 values)1
INT44Whole numbers −8 … 7 (16 values)0.5

Why fewer bits means less memory and faster speed

Memory is simple arithmetic: parameters × bytes per parameter. A 7B model needs 28 GB in FP32, 14 GB in FP16, 7 GB in INT8 and 3.5 GB in INT4, plus a little extra for scales and for the KV cache.

Speed improves for two reasons. First, generating each token needs the GPU or CPU to read every weight from memory, and that reading, not the arithmetic, is usually the bottleneck. Halve the bytes and you roughly halve the reading time. Second, many chips have special units that multiply INT8 or FP8 numbers faster than FP16 ones, which helps in compute-heavy phases like processing a long prompt.

Pause and think: If token generation on our laptop is limited by memory bandwidth and takes 100 ms per token in FP16, what is a rough estimate in INT4 weights?

Around 25–35 ms per token. INT4 weights are a quarter of the bytes of FP16, so reading them takes about a quarter of the time. In practice dequantization overhead and other costs mean the gain is a bit less than 4×.

The core mechanic: scale and zero-point

To map real numbers onto integers we need two values. The scale s is the size of one integer step in real units. The zero-point z is the integer that represents the real value 0.0.

Quantize four weights to INT8 by hand

  1. Find the range: Weights: 0.5, −1.27, 0.02, 1.0. The largest absolute value is 1.27.
  2. Compute the scale: INT8 symmetric uses −127 … 127. s = 1.27 / 127 = 0.01.
  3. Divide and round: 0.5/0.01 = 50, −1.27/0.01 = −127, 0.02/0.01 = 2, 1.0/0.01 = 100. Store [50, −127, 2, 100] as 1 byte each, plus the one scale.
  4. Dequantize when needed: Multiply back: 50 × 0.01 = 0.50, and so on. Here the error is zero because the numbers fell exactly on the grid; real weights usually land between grid points and lose up to half a step (0.005).

Symmetric vs asymmetric quantization

Symmetric quantization fixes z = 0 and centres the grid on zero: s = max|x| / qmax. It is simple and fast because there is no zero-point to subtract. It fits weights well, since they are usually spread roughly evenly around zero.

Asymmetric quantization uses the actual minimum and maximum: s = (max − min) / (qmax − qmin) and a non-zero z. It fits lopsided data. Example: activations after a ReLU are never negative. A symmetric grid wastes half its values on negatives that never occur; an asymmetric grid spends all 256 values on the range that is actually used, roughly halving the error.

Per-tensor vs per-channel quantization

How many weights share one scale? This choice is called granularity.

  • Per-tensor: one scale for a whole weight matrix. Cheapest to store, but if one row has big values, the scale is large and every small weight in the matrix rounds to almost nothing.
  • Per-channel: one scale per output channel (per row of the matrix). Each row gets a grid that fits its own range.
  • Per-group (block-wise): one scale per small group of, say, 32, 64 or 128 consecutive weights. This is standard for 4-bit LLM formats (GPTQ, AWQ and GGUF all use groups or blocks).

The code below quantizes a 64-row matrix where one row is 20× larger than the rest. Watch the per-tensor error explode, especially at 4 bits.

quantize.py

import numpy as np
def quant_sym(x, bits):
qmax = 2 ** (bits - 1) - 1                     # 127 for INT8, 7 for INT4
scale = np.abs(x).max() / qmax                 # one scale, zero-point = 0
q = np.clip(np.round(x / scale), -qmax, qmax)
return q * scale                               # dequantize to compare
def quant_asym(x, bits):
qmin, qmax = 0, 2 ** bits - 1                  # 0..255 for UINT8
scale = (x.max() - x.min()) / (qmax - qmin)
zero = np.round(qmin - x.min() / scale)        # integer that stands for 0.0
q = np.clip(np.round(x / scale) + zero, qmin, qmax)
return (q - zero) * scale
# Tiny worked example: 4 weights to INT8 (symmetric)
w = np.array([0.5, -1.27, 0.02, 1.0])
scale = np.abs(w).max() / 127
print("scale =", round(scale, 4), " ints =", np.round(w / scale).astype(int))
rng = np.random.default_rng(0)
W = rng.normal(0, 0.02, size=(64, 256))            # 64 output channels
W[3] *= 20                                         # one channel with large weights
relu_out = np.abs(rng.normal(0, 1, 1000))          # all-positive activations
err = lambda a, b: np.abs(a - b).mean()
print("ReLU activations, 8-bit: symmetric err %.4f | asymmetric err %.4f" %
(err(relu_out, quant_sym(relu_out, 8)), err(relu_out, quant_asym(relu_out, 8))))
for bits in (8, 4):
per_tensor = quant_sym(W, bits)
per_channel = np.vstack([quant_sym(row, bits) for row in W])
print(f"INT{bits} weights: per-tensor err {err(W, per_tensor):.5f} | "
f"per-channel err {err(W, per_channel):.5f}")
print("memory for 7B params: FP32 %.0f GB, FP16 %.0f GB, INT8 %.0f GB, INT4 %.1f GB"
% tuple(7e9 * b / 8 / 1e9 for b in (32, 16, 8, 4)))

Output:

scale = 0.01  ints = [  50 -127    2  100]
ReLU activations, 8-bit: symmetric err 0.0075 | asymmetric err 0.0039
INT8 weights: per-tensor err 0.00232 | per-channel err 0.00015
INT4 weights: per-tensor err 0.01639 | per-channel err 0.00281
memory for 7B params: FP32 28 GB, FP16 14 GB, INT8 7 GB, INT4 3.5 GB

Look at the INT4 per-tensor error: 0.016, close to the typical weight size of 0.02. With one shared scale, most normal weights rounded to zero. Per-channel scales bring the error down to 0.003. Granularity matters as much as bit width.

PTQ vs QAT, and weight-only vs weight-and-activation

Post-Training Quantization (PTQ) quantizes an already-trained model. Simple PTQ just rounds the weights. Smarter PTQ uses a small calibration set (a few hundred example texts) to see typical activation ranges and to adjust weights so rounding errors cancel. No retraining is needed, so it takes minutes to hours. Almost all LLM quantization is PTQ.

Quantization-Aware Training (QAT) simulates rounding during training or fine-tuning ("fake quantization"), so the model learns weights that survive rounding. Rounding has no useful gradient, so training passes gradients straight through it (the straight-through estimator). QAT gives the best low-bit accuracy but needs training data and compute, so it is used mainly by model makers, for example to release official low-bit versions.

A second choice is what to quantize. Weight-only quantization (written W4A16 or W8A16: 4- or 8-bit weights, 16-bit activations) stores weights in low bits and converts them back to 16-bit just before multiplying. It saves memory and speeds up memory-bound decoding. Weight-and-activation quantization (W8A8, FP8) also quantizes activations, so the multiply itself runs on fast integer or FP8 units. That helps compute-heavy work like long prompts and large batches, but activations are harder to quantize, as the next section shows.

The outlier problem in LLMs

LLMs have a quirk. In larger models, a few hidden dimensions in the activations take values far larger than the rest, often tens of times larger, and they appear consistently across tokens. The LLM.int8() paper (Dettmers et al., 2022) found that these outlier features emerge systematically once models reach roughly the 6.7B-parameter scale, and that naive 8-bit quantization then breaks down.

Why it hurts: as our code showed, one large value forces a large scale, and the large scale crushes every normal value onto zero. Removing outliers is not an option either, because the model relies on them. Several fixes exist:

  • Mixed precision (LLM.int8()): keep the few outlier dimensions in 16-bit and quantize the rest to 8-bit.
  • Smoothing (SmoothQuant): divide activations by a per-channel factor and multiply the matching weights by it. The maths is unchanged, but the outlier difficulty moves from activations (hard) to weights (easy).
  • Protecting salient weights (AWQ): weights that multiply large activations matter most, so scale them up before rounding to reduce their relative error.
  • Finer groups: small blocks of 32–128 weights stop one outlier from setting the scale for millions of weights.

Common mistake Quantizing an LLM with one scale per tensor because "it worked for my image classifier". LLM outliers make per-tensor INT8 activations lose a lot of accuracy and per-tensor INT4 weights fall apart. Use per-channel or per-group scales and an outlier-aware method.

Popular methods: GPTQ, AWQ, bitsandbytes, GGUF / llama.cpp

The methods you will meet most often
MethodCore ideaTypical use
GPTQ (2022)Quantizes weights one column at a time and adjusts the not-yet-quantized weights to cancel the error, using second-order (Hessian) information from a calibration set.4-bit (also 3- and 8-bit) GPU checkpoints; widely published on Hugging Face
AWQ (2023)Activation-aware: finds the small fraction of weight channels that see large activations and scales them to protect them before rounding. No backpropagation.4-bit GPU inference in vLLM, TensorRT-LLM and others
bitsandbytesA library that quantizes on the fly when you load a model: LLM.int8() with outlier handling, and 4-bit NF4/FP4. NF4 is the format QLoRA fine-tuning uses.Quick experiments and QLoRA fine-tuning in Hugging Face Transformers
GGUF / llama.cppA file format plus block-wise "k-quant" schemes (Q4_K_M, Q5_K_M, Q8_0 …) designed for fast CPU and Apple-silicon inference, with optional GPU offload.Running models locally via llama.cpp, Ollama, LM Studio

They are complementary rather than rivals: GPTQ and AWQ are algorithms for choosing good low-bit weights; bitsandbytes is a convenient runtime library; GGUF is a file format and set of block layouts for one popular engine. The next two lessons dig into GGUF and llama.cpp.

The accuracy trade-off and running LLMs locally

How much quality do we lose? Broadly, from published evaluations and community experience (exact results vary by model, method and task):

  • 8-bit (INT8 or FP8) is close to lossless for most LLMs when done with outlier-aware methods.
  • 4-bit with a good method (GPTQ, AWQ, Q4_K_M) usually costs a small amount of quality and is the popular sweet spot for local use.
  • 3-bit and below degrades noticeably; quality can fall apart on reasoning, maths and code first.
  • Bigger models tolerate it better. A 4-bit 70B model typically beats a 16-bit 8B model while needing less memory than a 16-bit 34B one.

Our laptop, solved A 7B model in a 4-bit format with group scales is about 4 GB of weights. It fits in the 6 GB GPU with room for a few thousand tokens of KV cache. The same idea lets a 70B model run in about 40 GB, within reach of a high-memory desktop or Mac.

When not to quantize (much) Avoid aggressive quantization when you are training or doing full fine-tuning (use BF16), when the task is very sensitive (exact maths, long code generation) and you have the memory to spare, or when you serve huge batches where compute, not memory, is the bottleneck and weight-only quantization brings little speedup. Always evaluate the quantized model on your own task, not just on perplexity.

Worked example, step by step

Earlier we quantized four weights by hand with a symmetric grid. Let us now do the asymmetric case, where the zero-point does real work. We use only 4 bits so the numbers stay small. Our five values are lopsided, as activations often are: −0.6, 0.0, 0.47, 1.33 and 2.4.

Asymmetric 4-bit quantization (integers 0 … 15)

  1. Find the range: min = −0.6 and max = 2.4, so the range is 3.0.
  2. Compute the scale: s = (max − min) / (qmax − qmin) = 3.0 / 15 = 0.2. Each integer step is worth 0.2.
  3. Compute the zero-point: z = round(qmin − min / s) = round(0 − (−0.6 / 0.2)) = 3. The integer 3 now stands for the real value 0.0.
  4. Quantize: q = round(x / s) + z. For 0.47: round(2.35) + 3 = 5. For 1.33: round(6.65) + 3 = 10. The ends map to 0 and 15.
  5. Dequantize: x̂ = s · (q − z). For q = 5: 0.2 × 2 = 0.4. For q = 10: 0.2 × 7 = 1.4. Each is off by 0.07.
  6. Compare with symmetric: A symmetric 4-bit grid uses −7 … 7 with s = 2.4 / 7 ≈ 0.343. The step is 70% larger, because the grid also covers −2.4 to −0.6, where we have no values at all.
The same five values on both 4-bit grids
ValueAsymmetric qAsymmetric x̂ErrorSymmetric qSymmetric x̂Error
−0.60−0.60−2−0.6860.086
0.030.0000.00
0.4750.40.0710.3430.127
1.33101.40.0741.3710.041
2.4152.4072.40

The symmetric grid only ever uses the integers −2 to 7: ten of its fifteen levels. The asymmetric grid uses all sixteen. A single value can still come out better on the coarser grid by luck, as 1.33 does here. What we can rely on is the worst case: the error is at most half a step, so 0.1 for the asymmetric grid against about 0.17 for the symmetric one.

Notice also that 0.0 is stored exactly on both grids. That is the reason the zero-point is rounded to a whole integer: zeros are very common in a network, and we do not want them to pick up an error.

Practice: try it yourself

The lesson named per-group scales as the standard for 4-bit LLM formats, but our earlier code only compared per-tensor with per-channel. Here we sweep the group size. We quantize one small weight matrix to 4 bits with one scale per 256, 64, 32 or 8 weights. For each we report the real cost in bits per weight, the weight error, and the error of the layer's output W @ x.

practice_group_size.py

import numpy as np
rng = np.random.default_rng(0)
W = rng.normal(0, 0.02, size=(64, 256))          # a small weight matrix
idx = rng.integers(0, W.size, size=40)
W.flat[idx] *= 25                                # 40 scattered outlier weights
x = rng.normal(size=256)                         # one input vector
BITS = 4
qmax = 2 ** (BITS - 1) - 1                       # 7 for 4-bit symmetric
def quantize_groups(W, group):
# one symmetric scale per `group` consecutive weights in a row
g = W.reshape(-1, group)
scale = np.abs(g).max(axis=1, keepdims=True) / qmax
return (np.clip(np.round(g / scale), -qmax, qmax) * scale).reshape(W.shape)
y = W @ x                                        # the exact layer output
print("group  bits/weight  weight error  output error")
for group in (256, 64, 32, 8):
Wq = quantize_groups(W, group)
bpw = BITS + 16 / group                      # 4-bit ints + one FP16 scale per group
w_err = np.abs(W - Wq).mean()
y_err = np.linalg.norm(Wq @ x - y) / np.linalg.norm(y)
print(f"{group:5d}  {bpw:11.2f}  {w_err:12.5f}  {y_err:12.3f}")
print(f"size: FP16 {W.size * 2} bytes, group 32 {int(W.size * 4.5 / 8)} bytes")

Output:

group  bits/weight  weight error  output error
256         4.06       0.00665         0.374
64         4.25       0.00305         0.183
32         4.50       0.00227         0.121
8         6.00       0.00127         0.082
size: FP16 32768 bytes, group 32 9216 bytes

Going from groups of 256 to groups of 32 costs less than half a bit per weight and cuts the output error from 0.374 to 0.121. Going on to groups of 8 helps a little more but costs 6 bits per weight. Now change it:

  • Set BITS = 3. Predict first: does 3 bits with groups of 8 (5 bits per weight in total) beat 4 bits with groups of 256 (4.06 in total) on output error?
  • Delete the two lines that create the outlier weights. Predict how much the group size still matters when all weights are of similar size.
  • Add x[:4] *= 20 after x is created, to mimic a few large activations. Predict which of the two error columns changes and which cannot change. What does that say about judging quality by weight error alone?

Pause and think: In the worked example, both grids have 4 bits, yet the asymmetric step is 0.2 and the symmetric step is about 0.343. For what kind of data would the two grids be equally good?

Data spread evenly around zero, with min ≈ −max. Then the symmetric grid wastes nothing, the two scales are almost the same, and the zero-point brings no benefit. This is why symmetric grids are the usual choice for weights and asymmetric ones for lopsided activations.

Pause and think: A file is labelled 4-bit, but its size works out to about 4.5 bits per weight. Using the practice output, explain where the extra half bit goes and why we accept it.

It pays for the scales. With one 16-bit scale per 32 weights, each weight carries 16 / 32 = 0.5 extra bits. We accept it because small groups stop one large weight from setting the scale for many others: in the practice run the output error fell from 0.374 with groups of 256 to 0.121 with groups of 32.

Wrapping up model quantization

  • Quantization stores numbers with fewer bits; weights snap to a grid defined by a scale and zero-point.
  • Memory falls in proportion to bits, and memory-bound decoding speeds up almost as much.
  • Symmetric suits weights; asymmetric suits skewed data. Per-channel and per-group scales are essential for LLMs.
  • PTQ is fast and dominant; QAT is best at very low bits. Weight-only saves memory; weight-and-activation also speeds up compute.
  • LLM activation outliers break naive schemes; LLM.int8(), SmoothQuant, AWQ and small groups handle them.
  • GPTQ, AWQ, bitsandbytes and GGUF are the everyday tools; 4-bit is the usual local sweet spot.

Key takeaways

  • Quantization maps real numbers to a few integer levels with a scale (and zero-point): q = round(x/s) + z.
  • Memory scales with bits: a 7B model is 14 GB in FP16 and about 3.5–4 GB in 4-bit.
  • Granularity matters: use per-channel or per-group scales, never one scale for a whole LLM tensor.
  • LLM activation outliers are real and important; LLM.int8(), SmoothQuant and AWQ are designed around them.
  • 8-bit is near lossless; 4-bit with GPTQ/AWQ/GGUF k-quants is the common sweet spot; below that, test carefully.

Key terms

  • Quantization: Representing numbers with fewer bits by rounding them to a coarse grid of allowed values.
  • Scale: The real-valued size of one integer step; multiply an integer by it to get back an approximate real value.
  • Zero-point: The integer that represents real 0.0 in asymmetric quantization.
  • Per-channel / per-group: Using a separate scale for each row or each small block of weights instead of one per tensor.
  • PTQ: Post-Training Quantization: converting an already-trained model to low precision, often with a small calibration set.
  • QAT: Quantization-Aware Training: simulating rounding during training so the model learns to tolerate it.
  • Outlier features: A few activation dimensions in large LLMs with values far bigger than the rest, which break naive quantization.

← 13.11 EAGLE: Feature-Level Drafting for Faster Inference · 13.13 GGUF: The File Format Powering Local LLM Inference →