Lesson 7.1 · 25 min
Small Language Models: Big Capability in Compact Form
When would a 3-billion-parameter model on a phone beat a frontier model in a data centre?
In short: Small Language Models (SLMs) are language models with roughly a few hundred million to around ten billion parameters, small enough to run cheaply, quickly and often on a single device. They stay capable thanks to high-quality and synthetic training data, distillation from larger models, long training, pruning and quantization. They shine for focused, high-volume, private or offline tasks; large models remain better for broad knowledge and hard multi-step reasoning.
SLM = Small + Language Model
The name says it all: a Small Language Model (SLM) is a language model built to be small. It uses the same kind of architecture as the big chat models (usually a decoder-only Transformer) but with far fewer parameters, the learned numbers inside the network. Fewer parameters means less memory, less compute per token, lower cost and lower latency.
Think of it like vehicles A cargo ship (a frontier LLM) can carry anything anywhere, but you would not use one to deliver a pizza. A scooter (an SLM) cannot carry a container, but for short, frequent, specific trips it is faster, cheaper and goes where the ship cannot, like a narrow street or your pocket.
Running example for this lesson: a bank wants to classify customer messages ("card lost", "dispute a charge", "update address") and draft short replies, millions of times a day, ideally on its own servers for privacy. We will keep asking: SLM or LLM?
What is a language model?
A language model assigns probabilities to text. Given some tokens, it predicts a probability for every possible next token. Generating text means repeatedly sampling a next token and appending it. Everything else (answering questions, summarising, classifying) is built on top of that one skill, through training and prompting.
What counts as "small"?
There is no official cut-off, and the line keeps moving as hardware improves. In 2026 most people use "SLM" for models from about 100 million to about 10 billion parameters, with some including models up to roughly 15B. A more practical definition: a model small enough to run on a single consumer device or a single modest GPU, often after quantization.
Size alone does not decide what a model costs to run. Precision matters too: a 3.8B model needs about 7.6 GB in 16-bit floats but under 2 GB at 4-bit quantization. That is why "small" is often discussed together with quantization.
Popular SLMs we should know
| Family | Maker | Small sizes | Known for |
|---|---|---|---|
| Phi (Phi-3, Phi-4-mini) | Microsoft | 3.8B | Heavy use of curated and synthetic "textbook-quality" data |
| Gemma 2 / Gemma 3 | 1B, 2B, 4B, 9B, 12B | Small models trained with distillation from larger ones | |
| Llama 3.2 | Meta | 1B, 3B | On-device use; built with pruning and distillation from larger Llamas |
| Qwen2.5 / Qwen3 | Alibaba | 0.5B to 8B | Strong multilingual and coding ability for the size |
| SmolLM2 | Hugging Face | 135M, 360M, 1.7B | Fully open, very small models |
| Ministral / Mistral 7B | Mistral AI | 3B, 7B, 8B | Efficient attention (GQA, sliding window in Mistral 7B) |
Phone makers also ship their own on-device models, for example Apple's on-device foundation model of about 3B parameters and Google's Gemini Nano on Android. New versions appear every few months, so check current leaderboards before choosing.
How SLMs stay capable despite being small
A 2026 3B model often beats much larger models from a few years earlier. Five techniques explain most of that progress:
The SLM toolkit
- Better data: Filter the web hard and add synthetic, textbook-style data written by larger models. Microsoft's Phi work showed data quality can substitute for a lot of size.
- Train far longer: Chinchilla suggested ~20 tokens per parameter for compute-optimal training, but SLMs are deliberately "over-trained" on trillions of tokens (Llama 3 8B saw about 15T tokens) because a small model that is cheap to serve is worth the extra training cost.
- Distil from a teacher: Knowledge distillation trains the small student to match a large teacher's full probability distribution over next tokens, which carries much more information than the single correct token.
- Prune: Pruning removes less important layers, heads or neurons from a bigger trained model, then retrains briefly, often with distillation. Llama 3.2 1B and 3B were made this way from larger Llama models.
- Quantize and specialise: Store weights in 8 or 4 bits for deployment, and fine-tune (often with LoRA) on the narrow task, where a small model can match a large one.
Why SLMs matter
- Cost: fewer parameters means fewer FLOPs per token and smaller GPUs; at millions of requests per day, this can cut serving cost by an order of magnitude or more.
- Latency: decoding is usually limited by memory bandwidth, so reading 2 GB of weights per token is far faster than reading 140 GB.
- Privacy and compliance: the model can run on-premises or on the user's device, so sensitive data never leaves it.
- Offline and edge: phones, cars, factory devices and laptops without reliable internet.
- Control: easy to fine-tune, version and audit your own model.
- Energy: less compute per request.
will_it_fit.py
# Will a model fit on a device? Weights only, plus ~20% headroom for KV cache and runtime.
BYTES = {"FP16": 2.0, "INT8": 1.0, "INT4": 0.5}
DEVICES = {"phone (8 GB, ~4 GB free)": 4, "laptop GPU (8 GB)": 8, "server GPU (80 GB)": 80}
models = {"0.5B": 0.5e9, "1.5B": 1.5e9, "3.8B": 3.8e9, "8B": 8e9, "70B": 70e9}
def footprint_gb(params, precision):
return params * BYTES[precision] * 1.2 / 1e9 # 20% headroom
print(f"{'model':>6} {'FP16 GB':>8} {'INT4 GB':>8} smallest device that fits (INT4)")
for name, p in models.items():
fp16, int4 = footprint_gb(p, "FP16"), footprint_gb(p, "INT4")
fits = next((d for d, cap in DEVICES.items() if int4 <= cap), "needs several GPUs")
print(f"{name:>6} {fp16:8.1f} {int4:8.1f} {fits}")
# Rough decode speed when memory-bandwidth bound: tokens/s ~ bandwidth / bytes read per token
bandwidth_gbs = 100 # e.g. a laptop-class memory system
for name in ["1.5B", "8B", "70B"]:
gb_per_token = models[name] * BYTES["INT4"] / 1e9
print(f"{name}: ~{bandwidth_gbs / gb_per_token:5.0f} tokens/s upper bound at {bandwidth_gbs} GB/s")Output:
model FP16 GB INT4 GB smallest device that fits (INT4) 0.5B 1.2 0.3 phone (8 GB, ~4 GB free) 1.5B 3.6 0.9 phone (8 GB, ~4 GB free) 3.8B 9.1 2.3 phone (8 GB, ~4 GB free) 8B 19.2 4.8 laptop GPU (8 GB) 70B 168.0 42.0 server GPU (80 GB) 1.5B: ~ 133 tokens/s upper bound at 100 GB/s 8B: ~ 25 tokens/s upper bound at 100 GB/s 70B: ~ 3 tokens/s upper bound at 100 GB/s
Pause and think: Using the rule in the code, about how much memory does a 3.8B model need at INT8, and would it fit in a phone with 4 GB free?
3.8B × 1 byte × 1.2 ≈ 4.6 GB, which does not fit in 4 GB. At INT4 it needs about 2.3 GB and fits. Precision can decide whether a model runs on a device at all.
SLM vs LLM
Where SLMs shine: use cases
- Classification and routing: intent detection, ticket triage, spam or toxicity filtering. Our bank's message categories fit perfectly.
- Extraction: pulling names, dates, amounts or product codes into JSON.
- On-device assistants: autocomplete, smart replies, summarising notifications, offline translation.
- RAG answerers: when retrieval supplies the facts, a small model only needs to read and rephrase them.
- Agent sub-steps: cheap models handle routine tool calls while a larger model plans.
- Draft models for speculative decoding: a small model proposes tokens that a large model verifies.
The bank, decided Fine-tune a 3–8B model on a few thousand labelled messages for classification and short templated replies, run it on the bank's own GPUs, and route the rare complex complaints (legal disputes, multi-issue letters) to a large model with human review.
Trade-offs of SLMs
- Less knowledge: fewer parameters store fewer facts, so SLMs hallucinate more on knowledge questions unless given the facts (RAG).
- Shallower reasoning: long, multi-step problems degrade faster, though small reasoning-tuned models have narrowed the gap on maths.
- Narrower generalisation: a fine-tuned SLM can be excellent on its task and poor just outside it.
- Context limits in practice: even if long context is supported, small devices may lack memory for a big KV cache.
Common mistake Judging an SLM by public benchmarks or by a few hand-typed prompts. Benchmarks may not resemble your task, and some models are tuned towards them. Build a test set of real examples from your domain and measure accuracy, latency and cost for both an SLM and an LLM before deciding.
When to pick an SLM
- The task is narrow and well defined (classify, extract, rewrite, short answers).
- Volume is high or latency must be low (interactive, real-time).
- Data must stay on the device or premises, or the app must work offline.
- You can supply knowledge through retrieval or fine-tuning instead of relying on the model's memory.
- Your evaluation shows the SLM meets the quality bar. If not, try fine-tuning, then a bigger SLM, then routing hard cases to an LLM.
Pause and think: A startup wants a creative writing partner that discusses any topic in depth and keeps a long novel consistent. Is an SLM the right first choice?
Probably not. The task is broad, open-ended and needs wide knowledge and long, coherent reasoning, which are LLM strengths. An SLM might still help with sub-tasks such as fixing grammar or autocomplete.
Worked example, step by step
The advice “use an SLM and escalate hard cases” only pays off if few cases escalate. Let us put numbers on the bank example. All figures here are illustrative: 1,000,000 messages a day, an SLM call costs 1 unit, an LLM call costs 20 units.
Costing a cascade
- Baseline: Send everything to the LLM: 1,000,000 × 20 = 20,000,000 units a day.
- SLM first: Every message goes through the SLM once: 1,000,000 × 1 = 1,000,000 units. This part is paid whatever happens next.
- Escalate the unsure ones: Say the SLM is unsure about 10% of messages. Those 100,000 go to the LLM: 100,000 × 20 = 2,000,000 units.
- Add up: 1,000,000 + 2,000,000 = 3,000,000 units, which is 15% of the baseline.
- Find the break-even: The cascade costs 1 + 20 × e units per message, where e is the escalation rate. It equals the baseline of 20 when e = 0.95. Above 95% escalation the SLM is pure overhead.
| Escalation rate | SLM cost | LLM cost | Total | Share of all-LLM cost |
|---|---|---|---|---|
| 0% | 1,000,000 | 0 | 1,000,000 | 5% |
| 10% | 1,000,000 | 2,000,000 | 3,000,000 | 15% |
| 30% | 1,000,000 | 6,000,000 | 7,000,000 | 35% |
| 100% | 1,000,000 | 20,000,000 | 21,000,000 | 105% |
Two things to watch in practice. First, the escalation rate is set by our confidence threshold, so a stricter threshold raises cost. Second, cost is only half of the picture: we also need the accuracy of the cases the SLM keeps. If the SLM is confidently wrong on many of them, a low escalation rate is a warning sign, not a success.
Practice: try it yourself
We will try knowledge distillation at the smallest possible scale. A “student” with four logits learns the answer to “The capital of France is ___” twice: once from a hard one-hot label, once from a teacher's soft probabilities. Then we compare what each student knows about the wrong answers.
practice_distillation.py
import math
def softmax(z):
e = [math.exp(x - max(z)) for x in z]
return [x / sum(e) for x in e]
def kl(p, q):
# KL(p || q): how far the student q is from the teacher p (0 = identical)
return sum(pi * math.log(pi / qi) for pi, qi in zip(p, q) if pi > 0)
tokens = ["Paris", "Lyon", "London", "banana"]
teacher = softmax([5.0, 2.0, 1.5, -3.0]) # illustrative teacher logits
hard = [1.0, 0.0, 0.0, 0.0] # one-hot label: only "Paris" counts
def train(target, steps=200, lr=0.5):
z = [0.0] * 4 # the student starts knowing nothing
for _ in range(steps):
q = softmax(z)
# gradient of cross-entropy with respect to the logits is (q - target)
z = [zi - lr * (qi - ti) for zi, qi, ti in zip(z, q, target)]
return softmax(z)
print("teacher ", " ".join(f"{t}={p:.3f}" for t, p in zip(tokens, teacher)))
for name, target in [("hard label", hard), ("soft labels", teacher)]:
q = train(target)
print(f"{name:11s}", " ".join(f"{t}={p:.3f}" for t, p in zip(tokens, q)),
f"| KL to teacher={kl(teacher, q):.4f}")Output:
teacher Paris=0.926 Lyon=0.046 London=0.028 banana=0.000 hard label Paris=0.992 Lyon=0.003 London=0.003 banana=0.003 | KL to teacher=0.1350 soft labels Paris=0.924 Lyon=0.044 London=0.026 banana=0.007 | KL to teacher=0.0055
Now change it:
- Cut
stepsfrom200to20. Predict which student is further from its target after so few updates, and check the KL values. - Change the teacher logits to
[5.0, 4.5, 1.5, -3.0], a teacher that is unsure between Paris and Lyon. Predict the soft-label student's top two probabilities. - Soften the teacher: divide every teacher logit by
2before the softmax (a temperature of 2). Predict whether the probabilities of Lyon and London go up or down.
Pause and think: The hard-label student gives Lyon, London and banana exactly the same probability. What knowledge is it missing, and why could the hard label never teach it?
It does not know that Lyon and London are “less wrong” than banana. A one-hot target carries no ranking among wrong answers: the gradient pushes every wrong logit down by the same rule, so they stay equal. The teacher's soft probabilities carry that ranking, which is the extra signal distillation gives.
Pause and think: The soft-label student ends with KL = 0.0055, not 0, and gives banana 0.007 where the teacher gives almost 0. Why?
It started from equal logits and had only 200 updates. Matching a probability near zero needs a very negative logit, and the gradient for that token gets tiny as its probability shrinks, so the last bit of the gap closes slowly. More steps would push the KL closer to 0.
Quick summary
- SLMs are the same kind of model as LLMs, just with roughly 0.1–10B parameters.
- Good data, long training, distillation, pruning and quantization keep them surprisingly capable.
- They win on cost, latency, privacy and offline use; LLMs win on breadth and hard reasoning.
- Pick the smallest model that passes your own evaluation, and route hard cases upward.
Key takeaways
- An SLM is a normal language model with roughly 0.1–10B parameters; the boundary is a convention.
- Data quality, long training, distillation, pruning and quantization keep small models capable.
- Memory ≈ parameters × bytes per parameter; quantization often decides if a model fits a device.
- SLMs win on cost, latency, privacy and offline use; LLMs win on knowledge and hard reasoning.
- Choose by evaluating on your own data, and route hard cases to a larger model.
Key terms
- Small Language Model (SLM): A language model small enough (about 0.1–10B parameters) to run cheaply on one device or modest GPU.
- Parameter: A learned number inside a neural network.
- Knowledge distillation: Training a small student model to imitate a larger teacher's output distributions.
- Pruning: Removing less important parts of a trained model, then retraining briefly.
- Quantization: Storing weights with fewer bits (e.g. 8 or 4) to save memory and speed up inference.
- On-device inference: Running the model directly on a phone, laptop or edge device instead of a server.
← 6.7 DeepSeek-V4: Anatomy of an Open-Source Frontier Model · 7.2 Large Reasoning Models: Chain-of-Thought at Inference Time →