Modern AI Engineering

Lesson 6.7 · 31 min

DeepSeek-V4: Anatomy of an Open-Source Frontier Model

How does an open model read a million tokens while keeping only about a tenth of the memory its predecessor needed?

In short: DeepSeek-V4 (April 2026) is a pair of open Mixture-of-Experts models, V4-Pro (1.6T total, 49B active parameters) and V4-Flash (284B total, 13B active), both with a native 1M-token context. Its main ideas are a hybrid attention that compresses past tokens (CSA: compress 4× and pick the top-k blocks; HCA: compress 128× and attend to all), stronger residual connections called mHC, the Muon optimizer, FP4 quantization-aware training, and a post-training recipe that trains domain specialists and merges them by on-policy distillation. Users choose among three reasoning modes.

The big picture

DeepSeek is a Chinese AI lab known for open-weight models with unusually efficient designs. Its V3 model (late 2024) introduced a widely copied recipe: a large MoE with Multi-head Latent Attention and Multi-Token Prediction. V3.2 added DeepSeek Sparse Attention (DSA), where a cheap "lightning indexer" picks which past tokens each query should attend to. DeepSeek-V4, released in April 2026 with a technical report titled DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence, pushes that line further, aiming at agents that work over very long contexts.

This lesson assumes we know MoE, GQA, sliding windows and the KV cache from earlier lessons. We will decode each new component, using one running example: an AI coding agent that has loaded a 1-million-token code repository plus a long history of tool calls.

Think of it like a researcher's notes Reading a million pages, a good researcher keeps the last page in full detail, a one-line note for every few pages, and a one-line note per chapter. To answer a question, they skim the chapter notes, pick the few most relevant page notes, and reread only the latest pages closely. V4's attention does something very similar with its KV cache.

Two models: DeepSeek-V4-Pro and DeepSeek-V4-Flash

Published configuration (from the model cards and technical report)
DeepSeek-V4-ProDeepSeek-V4-Flash
Total parameters1.6 trillion284 billion
Active parameters per token49 billion (~3%)13 billion (~4.6%)
Context length1M tokens1M tokens
CSA top-k selected entries1,024512
Single-token FLOPs at 1M context vs V3.227%10%
KV cache at 1M context vs V3.210%7%
LicenseMITMIT

Pro is the quality flagship and, at release, among the largest open-weight models. Flash trades some quality for much lower cost and can be served on far smaller hardware. Both share the same architecture ideas; they differ in size and a few settings. Base and instruct checkpoints were published for each.

Hybrid attention with CSA and HCA

At 1M tokens, even V3.2's sparse attention keeps a KV entry per token. V4 instead compresses the sequence: groups of consecutive tokens are merged into a single KV entry. Two kinds of layers use two strengths of compression, and the network interleaves them (in V4-Pro's 61 layers, the first two are HCA and the rest alternate CSA and HCA).

  • CSA, Compressed Sparse Attention. Every 4 consecutive tokens are pooled into one entry using learned, softmax-gated weights. That still leaves 250,000 entries at 1M tokens, too many to attend to, so a lightning indexer (inherited from V3.2's DSA, and itself run in low precision) scores all compressed entries cheaply and each query attends only to the top-k (1,024 in Pro, 512 in Flash). Fine detail, selectively.
  • HCA, Heavily Compressed Attention. Every 128 tokens become one entry. At 1M tokens that is only about 7,800 entries, few enough that each query attends to all of them densely, with no selection step. A coarse but complete view of everything.
  • Sliding window branch. Both layer types also attend to the most recent 128 tokens uncompressed, so local, word-level detail is never lost to compression.

One query in a CSA layer, step by step

  1. Compress: As tokens arrive, each block of 4 tokens' keys and values is pooled into one compressed entry and stored in the cache.
  2. Index: The lightning indexer computes a cheap relevance score between the current query and every compressed entry.
  3. Select: Keep the top-k compressed entries (e.g. 1,024 for Pro). In our agent example, these might be the blocks containing the function being edited and its callers.
  4. Add the local window: Add the last 128 uncompressed tokens, e.g. the agent's current line of reasoning.
  5. Attend: Run ordinary attention over about 1,152 entries instead of 1,000,000.

csa_hca_budget.py

import numpy as np
np.random.seed(0)
# 1) Compress every m token keys into one entry with learned weights (toy version)
m, d = 4, 6
keys = np.random.randn(16, d)                         # 16 tokens of keys
w_logits = np.random.randn(m)                         # learned position weights
w = np.exp(w_logits) / np.exp(w_logits).sum()
compressed = (keys.reshape(-1, m, d) * w[None, :, None]).sum(1)
print("tokens:", keys.shape[0], "-> compressed entries:", compressed.shape[0])
# 2) Top-k selection by a cheap indexer score (toy: dot product with a query)
q = np.random.randn(d)
idx_scores = compressed @ q
top = np.argsort(idx_scores)[-2:][::-1]
print("indexer picks compressed blocks:", top.tolist())
# 3) Entries per layer at a 1M-token context, using DeepSeek-V4-Pro's published settings
n, m_csa, m_hca, top_k, window = 1_000_000, 4, 128, 1024, 128
print(f"{'layer type':14s} {'KV entries stored':>18s} {'entries attended/query':>24s}")
print(f"{'full attention':14s} {n:>18,} {n:>24,}")
print(f"{'CSA':14s} {n // m_csa + window:>18,} {top_k + window:>24,}")
print(f"{'HCA':14s} {n // m_hca + window:>18,} {n // m_hca + window:>24,}")

Output:

tokens: 16 -> compressed entries: 4
indexer picks compressed blocks: [0, 1]
layer type      KV entries stored   entries attended/query
full attention          1,000,000                1,000,000
CSA                       250,128                    1,152
HCA                         7,940                    7,940

Pause and think: Why does V4 need both CSA and HCA instead of just one of them?

They cover each other's blind spots. HCA gives every query a complete but coarse view of the whole context (nothing can be missed, but detail is blurred by 128× compression). CSA gives fine 4-token detail but only for the top-k blocks the indexer selects, which could miss something. Interleaving gives both coverage and precision, and the sliding window keeps recent tokens exact.

Manifold-Constrained Hyper-Connections (mHC)

A normal Transformer has a single residual stream: each layer reads the token vector, computes something, and adds it back (x + f(x)). Hyper-Connections (a 2024 idea from other researchers) widen this into several parallel streams, with learnable matrices that mix the streams before and after each layer. That adds expressiveness at little compute cost.

The problem: when unconstrained mixing matrices are multiplied across dozens of layers, signals can grow or shrink exponentially, making very large training runs unstable. DeepSeek's mHC constrains the stream-mixing matrix to be doubly stochastic: all entries non-negative, and every row and every column summing to 1. The set of such matrices is a geometric object (a "manifold", the Birkhoff polytope), hence the name. A few iterations of the Sinkhorn-Knopp algorithm (alternately normalising rows and columns) project the learned matrix onto this set.

Muon optimizer

An optimizer decides how to change weights given the gradients. Most LLMs have been trained with AdamW, which scales each weight's update individually. Muon (introduced in 2024 by Keller Jordan and collaborators) treats each weight matrix as a whole: it takes the momentum of the gradient and orthogonalises it with a few Newton-Schulz iterations, a cheap way of making the update push equally in all important directions instead of letting a few directions dominate.

Moonshot's Kimi K2 showed Muon could work at trillion-parameter scale. DeepSeek reports using Muon for V4, with a custom distributed implementation, for faster convergence and better stability than AdamW-style baselines. In typical Muon setups, AdamW is still used for parameters that are not ordinary matrices, such as embeddings and normalisation weights.

FP4 quantization-aware training

Quantization stores numbers with fewer bits. FP4 uses only 4 bits per number, a quarter of BF16. Quantizing a model after training often hurts quality. Quantization-aware training (QAT) instead simulates the low precision during training, so the model learns weights that still work when rounded.

V4 applies FP4 QAT to the MoE expert weights, which are the vast majority of its parameters, and to the indexer's query-key path. The released checkpoints store expert weights in FP4 and most other weights in FP8. For a 1.6T-parameter model, halving bytes per expert weight versus FP8 makes a large difference to how many GPUs are needed for serving.

Pause and think: Roughly how much memory would 1.5 trillion expert parameters take in FP4 compared with BF16?

BF16 uses 2 bytes per parameter: about 3 TB. FP4 uses 0.5 bytes: about 0.75 TB. That is a 4× saving (ignoring small overheads such as scaling factors).

Pre-training

Pre-training is the long first phase where the model learns to predict text from a huge corpus. DeepSeek reports pre-training V4 on more than 32 trillion tokens of diverse, filtered data. Like V3, it keeps Multi-Token Prediction (MTP): extra heads learn to predict tokens further ahead, which gives a richer training signal and can be reused at inference for speculative decoding. Long context is not bolted on at the end only; the attention design is built so that 1M-token sequences are affordable to train and serve.

Post-training: specialist training and on-policy distillation

After pre-training, V4 is shaped into an assistant in two stages:

  • Specialist training. Separate copies of the model are trained per domain (for example mathematics, coding, agentic tool use and instruction following) with supervised fine-tuning (SFT) and then reinforcement learning using GRPO (Group Relative Policy Optimization, the RL method DeepSeek used for R1, which scores a group of sampled answers against each other instead of training a separate value model).
  • On-policy distillation. A single student model then learns from all specialists. "On-policy" means the student generates its own answers and the relevant specialist grades each token of those answers (the report describes a reverse-KL objective). Learning on its own outputs avoids the mismatch of copying teacher text the student would never produce.

Reasoning modes

The three modes exposed by DeepSeek-V4
ModeBehaviourGood for
Non-thinkAnswers directly, no visible chain of thoughtFormat conversion, lookups, routine agent steps
Think HighReasons step by step before answeringDebugging, planning, moderate maths
Think MaxMaximum reasoning effort; very long thoughts (DeepSeek advises a context of at least ~384K tokens)The hardest problems where accuracy matters most

Using the modes in our coding agent Route routine steps such as "list files" or "rename this variable" to Non-think, the core bug investigation to Think High, and a tricky concurrency proof to Think Max. Cost and latency stay bounded while hard steps get more thought. In thinking mode with tool calls, V4 also keeps earlier reasoning across turns so the agent does not lose its train of thought.

Putting it all together

Each piece removes a specific bottleneck. Compressed hybrid attention cuts KV memory and FLOPs at 1M tokens. MoE keeps compute per token near 49B (Pro) or 13B (Flash) despite huge capacity. mHC and Muon keep a very large training run stable and efficient. FP4 QAT shrinks the memory for the experts. Specialist RL plus on-policy distillation gives one model strong skills in several domains, and reasoning modes let users trade cost for depth.

What to be careful about Compression is lossy: exact details from far back survive only if the indexer selects the right CSA blocks or the coarse HCA summary keeps them, so exact long-range recall should be tested on your task. Vendor efficiency numbers are relative to V3.2 at 1M tokens; at short contexts the gains are smaller. And a 1.6T-parameter model still needs a multi-GPU server; "efficient" does not mean "runs on a laptop".

Worked example, step by step

The earlier script counted entries at 1M tokens. Let us repeat the same simple count at three context sizes, using only the settings already listed for V4-Pro: 4× compression with top-k 1,024 for CSA, 128× compression for HCA, and a 128-token local window in both. As before, this count ignores head structure and storage precision.

The three formulas for a context of n tokens

  1. CSA entries stored: n ÷ 4 compressed entries, plus the 128 uncompressed window tokens.
  2. CSA entries attended: The smaller of n ÷ 4 and the top-k of 1,024, plus the 128 window tokens. Top-k cannot pick more entries than exist.
  3. HCA entries: n ÷ 128 compressed entries plus the window. All of them are attended, so stored and attended are the same number.
Simple entry count per layer type, V4-Pro settings (whole-number division)
Context nCSA storedCSA attended per queryHCA stored and attendedFull attention
2,0006286281432,000
32,0008,1281,15237832,000
1,000,000250,1281,1527,9401,000,000

Read the middle column from top to bottom. At 2,000 tokens there are only 500 compressed entries, fewer than the top-k, so in this count selection has nothing to drop and a query reads 628 entries: about 31% of full attention. At 32,000 tokens it reads 1,152, about 3.6%. At 1M tokens it still reads 1,152, about 0.12%. The work per query in a CSA layer stops growing once n ÷ 4 passes the top-k, while HCA keeps growing slowly, at 1/128 of the rate of full attention.

This is the arithmetic behind the earlier warning that gains are smaller at short contexts. It also shows what to test on our own task: place one exact detail far back in a long input and ask for it. Whether it comes back depends on the indexer picking its block, which no entry count can tell us.

Practice: try it yourself

We will build a toy to feel why pooling is lossy and why selection helps. We hide one “needle” token in 64 random tokens, merge keys with a plain mean, and check whether a query that resembles the needle still finds the right entry. This is a simplified simulation with random keys and plain averaging, not the model's learned pooling or indexer.

practice_pooling_recall.py

import numpy as np
rng = np.random.default_rng(0)
n, d = 64, 8                                   # toy context: 64 tokens, 8-dim keys
keys = rng.normal(0, 1, (n, d))
needle = 37                                    # the token our query is looking for
query = keys[needle] + rng.normal(0, 0.3, d)   # the query resembles the needle's key
def pooled_scores(m):
# merge every m keys into one entry (a plain mean in this toy), then score
entries = keys.reshape(n // m, m, d).mean(axis=1)
return entries @ query
for m in [1, 4, 16]:
s = pooled_scores(m)
best = int(s.argmax())
ranked = np.sort(s)
print(f"pool x{m:<2}: entries={n // m:>2} best entry={best:>2} "
f"holds needle={best == needle // m} lead over runner-up={ranked[-1] - ranked[-2]:.2f}")
# Sparse selection on the x4 entries: keep the top-k, plus a recent local window
k, window = 3, 8
top = np.argsort(pooled_scores(4))[-k:][::-1]
print("top-k x4 entries:", top.tolist(), "-> token ranges",
[(int(e) * 4, int(e) * 4 + 3) for e in top])
print("entries attended:", k + window, "instead of", n)

Output:

pool x1 : entries=64 best entry=37 holds needle=True lead over runner-up=10.12
pool x4 : entries=16 best entry= 5 holds needle=False lead over runner-up=0.66
pool x16: entries= 4 best entry= 1 holds needle=False lead over runner-up=0.55
top-k x4 entries: [5, 9, 15] -> token ranges [(20, 23), (36, 39), (60, 63)]
entries attended: 11 instead of 64

Now change it:

  • Set k = 1. Predict from the output above whether the needle's block is still selected.
  • Change the noise on line 7 from 0.3 to 0.0, so the query equals the needle's key. Predict whether pooling 16× now finds the right entry.
  • Set needle = 62. Predict whether pooling matters for this needle at all, given the local window of 8 recent tokens.

Pause and think: At 4× pooling the top-ranked entry was block 5, not the needle's block 9, yet top-k = 3 kept block 9. What does this tell us about choosing k?

After pooling, the needle's key is mixed with three unrelated keys, so its score is noisy and it may not rank first. A larger k is a safety margin: it keeps blocks that are relevant but not top-ranked. The price is more entries to attend to, so k trades recall against compute.

Pause and think: If the needle were token 60 of 64 and the local window covers the last 8 tokens, would compression affect it?

No. Tokens 56 to 63 are inside the local window, which is kept uncompressed, so the query can read token 60 exactly. Compression only affects tokens that have left the window. This is why recent detail stays sharp while old detail depends on pooling and selection.

Quick summary

  • Two open MoE models: V4-Pro (1.6T / 49B active) and V4-Flash (284B / 13B active), both 1M-token context.
  • CSA: pool 4 tokens per entry, an indexer picks top-k entries; HCA: pool 128 tokens per entry, attend to all; both add a 128-token local window.
  • mHC: multi-stream residuals with doubly stochastic mixing for stability.
  • Muon optimizer and FP4 QAT for experts make training and serving more efficient.
  • Domain specialists (SFT + GRPO) are merged by on-policy distillation; three reasoning modes at inference.

Key takeaways

  • DeepSeek-V4 comes as V4-Pro (1.6T / 49B active) and V4-Flash (284B / 13B active), both with 1M-token context.
  • Hybrid attention compresses the KV cache: CSA (4×, top-k) and HCA (128×, dense) plus a 128-token window.
  • At 1M tokens, V4-Pro needs about 27% of V3.2's per-token FLOPs and 10% of its KV cache.
  • mHC stabilises multi-stream residuals; Muon and FP4 QAT make large-scale training and serving cheaper.
  • Specialists trained with SFT + GRPO are merged by on-policy distillation; Non-think / Think High / Think Max trade cost for depth.

Key terms

  • CSA: Compressed Sparse Attention: pool every 4 tokens into one KV entry and attend to the top-k selected entries.
  • HCA: Heavily Compressed Attention: pool every 128 tokens into one KV entry and attend to all of them.
  • Lightning indexer: A cheap scoring module that picks which compressed entries each query attends to.
  • mHC: Manifold-Constrained Hyper-Connections: multi-stream residuals with doubly stochastic mixing matrices.
  • Muon: An optimizer that orthogonalises momentum updates for weight matrices via Newton-Schulz iterations.
  • Quantization-aware training: Training while simulating low-precision weights so the model stays accurate when quantized.
  • On-policy distillation: A student learns from teacher feedback on answers the student itself generated.

← 6.6 Flash Attention: Memory-Efficient Attention at Scale · 7.1 Small Language Models: Big Capability in Compact Form →