Lesson 17.4 · 28 min
Language Processing Units: A New Approach to LLM Inference
If a GPU can do a thousand trillion operations per second, why does a chatbot running on it still type at a few dozen words per second — and how did one chip design make that dramatically faster?
In short: An LPU (Language Processing Unit) is Groq's chip design for fast LLM inference. Token-by-token generation is limited by how fast weights can be read from memory, not by arithmetic, so the LPU keeps model weights in fast on-chip SRAM spread across many chips, and has a compiler plan every operation and every data transfer in advance so nothing ever waits or guesses. The result is very low latency per user, paid for with many chips per model and less flexibility than a GPU.
What is an LPU?
An LPU (Language Processing Unit) is a processor designed by the company Groq to run large language models (LLMs) quickly. It grew out of Groq's earlier Tensor Streaming Processor (TSP) architecture, described in a 2020 research paper. Groq's founder, Jonathan Ross, had previously worked on Google's TPU project.
The LPU is an inference chip: it runs already-trained models; it is not built for training them. Its headline promise is speed per user: very high tokens per second and low, predictable latency. Groq offers it mainly as a cloud service (GroqCloud) through an API, and as on-premises systems.
Think of it like a train timetable versus city traffic A GPU is like city traffic: lots of cars, traffic lights, junctions and drivers making decisions on the fly. Throughput is high, but any single trip has unpredictable delays. An LPU is like a railway with a fixed timetable planned months ahead: every train knows exactly which track it will be on, at which second. No one waits at junctions, and every trip takes exactly as long as scheduled.
Running example: our appliance shop's support assistant. Customers on the website hate waiting while an answer trickles out. We want to know whether an LPU-based API would help, and what it would cost us in flexibility.
How an LLM writes text, one token at a time
An LLM generates text autoregressively: it predicts one token (a word piece), appends it to the text, and runs again to predict the next. A 300-token answer means 300 sequential passes through the model. There are two phases:
The two phases of LLM inference
- Prefill: The whole prompt (say 500 tokens of instructions, ticket history and the customer's question) is processed in one parallel pass. Lots of maths per weight read: this phase is compute-heavy and GPUs handle it well. It produces the first token and fills the KV cache.
- Decode, step 1: To produce the next token, the model runs again on just the newest token, using the stored keys and values (the KV cache) for everything before it.
- Decode, step 2 … N: Repeat once per output token. Each step depends on the previous one, so they cannot run in parallel. This phase sets the 'typing speed' the user sees.
- Stop: Generation ends at an end-of-sequence token or a length limit.
Pause and think: Why can't we generate all 300 output tokens in parallel, the way prefill processes all 500 prompt tokens at once?
Because each new token depends on all the tokens before it, including ones not generated yet. Prompt tokens are all known up front, so they can be processed together; output tokens must be produced one after another.
The real bottleneck is memory, not math
During one decode step for one user, every weight of the model is read once and used for about two floating-point operations (one multiply, one add). A 70B-parameter model in 8-bit format means reading 70 GB to do 140 GFLOPs — about 2 FLOPs per byte. Modern accelerators can do hundreds of FLOPs in the time it takes to read one byte from their main memory. So the arithmetic units mostly sit idle, waiting for weights to arrive.
decode_bottleneck.py
import math
# Generating ONE token with batch size 1 means reading every weight once.
params = 70e9 # a 70B-parameter model
bytes_per_param = 1 # 8-bit weights
weight_bytes = params * bytes_per_param
flops_per_token = 2 * params # one multiply + one add per weight
print(f"weights to read per token: {weight_bytes / 1e9:.0f} GB")
print(f"math per token: {flops_per_token / 1e9:.0f} GFLOPs")
# Illustrative hardware figures (order of magnitude, not a benchmark)
gpu_bw, gpu_flops = 3.35e12, 990e12 # one HBM GPU: bytes/s, FLOP/s
t_mem = weight_bytes / gpu_bw
t_math = flops_per_token / gpu_flops
print(f"\nGPU: time to read weights {t_mem * 1e3:6.2f} ms"
f" | time to do the math {t_math * 1e3:5.3f} ms")
print(f"GPU: memory-bound ceiling ~{1 / t_mem:.0f} tokens/s for one user")
# On-chip SRAM design: weights are spread over many chips, each with
# a small but very fast memory, and all chips read their slice in parallel
sram_per_chip = 230e6 # bytes of on-chip SRAM per chip
sram_bw_per_chip = 80e12 # bytes/s of on-chip bandwidth per chip
chips = math.ceil(weight_bytes / sram_per_chip)
t_sram = (weight_bytes / chips) / sram_bw_per_chip
print(f"\nSRAM chips needed just to hold the weights: {chips}")
print(f"each chip reads its {weight_bytes / chips / 1e6:.0f} MB slice in "
f"{t_sram * 1e6:.2f} us")
print("(real speed is then set by chip-to-chip hops, the compiler's schedule,"
" and the KV cache, not by this raw read time)")Output:
weights to read per token: 70 GB math per token: 140 GFLOPs GPU: time to read weights 20.90 ms | time to do the math 0.141 ms GPU: memory-bound ceiling ~48 tokens/s for one user SRAM chips needed just to hold the weights: 305 each chip reads its 230 MB slice in 2.87 us (real speed is then set by chip-to-chip hops, the compiler's schedule, and the KV cache, not by this raw read time)
Why a GPU struggles here
GPUs are superb at throughput: given many requests at once, they batch them so each weight read serves many users, and their compute units fill up. But for the latency of one stream, several things work against them:
- Weights live in off-chip HBM. HBM is very fast by memory standards (several TB/s), but every decode step must stream the entire model through it.
- Dynamic hardware. Caches, schedulers that pick which warps to run, and memory controllers that arbitrate between requests all make timing vary from run to run. Software has to leave slack for the worst case.
- Kernel-by-kernel execution. A forward pass is many separate kernel launches; small per-step overheads add up when a step must take only milliseconds.
- Multi-GPU communication. Big models are split across GPUs, and each decode step needs synchronisation between them, adding latency.
- Batching trades latency for throughput. Serving systems batch users together to use the hardware efficiently, which can make each individual stream slower.
GPUs are not standing still Speculative decoding, better kernels, quantisation, faster interconnects and larger on-chip caches keep improving GPU decode speed. The LPU's advantage is about design trade-offs, not a law of nature.
Idea 1: keep the model on the chip
If reading weights from off-chip memory is the bottleneck, remove that memory from the critical path. The LPU stores model weights in SRAM (static RAM) built directly on the chip, right next to the compute units. Groq's first-generation chip has about 230 MB of SRAM with on-chip memory bandwidth Groq quotes at up to roughly 80 TB/s — over 20× the HBM bandwidth of a top GPU of the same era. Weights are not cached copies of something in DRAM: the SRAM is the main memory.
The problem with on-chip memory
SRAM is fast because each bit is stored in a handful of transistors right beside the logic. That also makes it expensive in chip area: a gigabyte of SRAM would need far more silicon than a single chip can hold. So each LPU chip holds only a sliver of a large model:
- A 70B-parameter model at 8 bits needs about 70 GB → roughly 300 chips at ~230 MB each, before counting the KV cache and activations (our code computed 305).
- Groq's public demos of 70B-class models ran on hundreds of chips across multiple racks.
- Long contexts and many concurrent users need KV-cache space too, which competes for the same scarce SRAM.
Spreading one model over hundreds of chips creates a new problem: the chips must pass activations to each other for every token, and any waiting at those hand-offs would destroy the speed gained from SRAM. The next two ideas address exactly that.
Pause and think: Our shop considers running a 7B model in 8-bit on LPUs. About how many ~230 MB chips are needed just for the weights?
7 GB ÷ 0.23 GB ≈ 30.4, so about 31 chips — before leaving room for the KV cache. The same model fits on a single GPU with room to spare, which shows the trade-off clearly.
Idea 2: remove all the guesswork
Conventional processors spend much of their chip area and energy reacting to the unknown: caches guess what data will be needed next, branch predictors guess which way code will go, schedulers decide at run time which work to run, and memory controllers arbitrate between competing requests. When a guess is wrong, the processor stalls.
The LPU takes the opposite approach: deterministic, software-scheduled execution. Neural-network inference is extremely predictable — the same operations on the same-shaped tensors every time — so Groq's compiler works out, ahead of time, exactly which functional unit does what on which clock cycle, and exactly when each piece of data arrives where it is needed. The hardware has no caches and no dynamic scheduling; it simply follows the plan.
- Predictable latency: a given program takes the same number of cycles every time, so performance can be known before running.
- More silicon for useful work: area that would hold caches and schedulers goes to compute and SRAM instead.
- No stalls from mispredictions or cache misses, because nothing is predicted and nothing misses.
Idea 3: a network that never waits
Determinism is extended across chips. In a typical cluster, chips send packets through switches that queue and route them dynamically, so arrival times vary, and receivers must wait and check. In Groq's design, chips are linked directly and the compiler also schedules chip-to-chip transfers: it knows, for each cycle, which link carries which data. Many chips can therefore behave like one very large, synchronised processor.
Because transfers are planned, there is no need for handshakes or for buffering 'just in case', and the compiler can overlap communication with computation so data arrives at the next chip exactly when that chip is ready to use it.
The cost of a fixed plan A static schedule is only as good as its assumptions. Changing the model, its size, or how it is split across chips means recompiling. Workloads with unpredictable behaviour cannot benefit as much, and the whole system must be engineered (clocks, links, failures) to keep the timetable reliable.
The assembly line
Inside a chip, the TSP/LPU layout looks less like a grid of independent cores and more like a factory floor. Functional units of the same kind are grouped into vertical slices — matrix-multiply units, vector units, memory units, data-reshaping (switch) units. Data streams horizontally across the chip, passing slice after slice. Each slice, following the compiler's instructions, reads the streams flowing past, works on them, and writes results back onto streams for the next slice.
Like a car factory In a car factory, each station does one job — weld, paint, fit wheels — and the conveyor brings the car to each station exactly when it is ready. No worker hunts for parts; no car waits in a queue. Every token's computation rides such a conveyor through the chips.
What happens when we send a prompt
One request to an LPU-backed API
- Ahead of time: compile and load: Before any user arrives, the model is compiled for a specific set of chips, and its weights are loaded and partitioned into the SRAM of all of them. They stay resident.
- Request arrives: Our support bot sends the prompt over the API. The host system tokenises it and sends the tokens into the chip network.
- Prefill: The prompt tokens stream through the layers spread across the chips, building the KV cache and producing the first output token.
- Decode on the conveyor: Each new token flows through every layer, chip to chip, on the pre-planned schedule. Weights are read from local SRAM; nothing waits for off-chip memory.
- Stream back: Tokens are sent back to our application as they are produced (streaming), so the customer sees the answer appear quickly.
Why an LPU is fast, all in one place
| Design choice | Effect on speed |
|---|---|
| Weights in on-chip SRAM | Removes the off-chip memory-bandwidth wall during decoding |
| Model spread over many chips | Aggregate SRAM bandwidth grows with the number of chips |
| Compiler-scheduled execution | No cache misses, mispredictions or run-time scheduling stalls; predictable latency |
| Scheduled chip-to-chip network | No queuing or handshakes at hand-offs between chips |
| Streaming 'assembly line' layout | Computation and data movement overlap continuously |
Independent benchmarks have generally shown LPU-based services among the highest output speeds per user for the open models they serve, often several times faster than typical GPU-based endpoints. Exact numbers change quickly with new models, software and GPU generations, so check current measurements before deciding.
Where an LPU works well, and where it does not
LPU vs GPU, and when to use which
- Use an LPU service when user-perceived speed is the product (voice, live chat, interactive agents), the model we need is offered, and an API fits our data rules.
- Use GPUs when we train or fine-tune, need a custom or very large model, must self-host on our own hardware, or care most about cost per token in large offline batches.
- Use both in a router: fast LPU-served models for interactive turns, GPU-served models for heavy or custom jobs.
Decision for our shop For the live website chat, an LPU-backed API serving an open model is attractive: answers appear almost instantly. For nightly summarisation of thousands of tickets and for fine-tuning on our own data, GPUs are the better fit. The landscape also shifts quickly, and other vendors (for example Cerebras, with wafer-scale on-chip memory) pursue related SRAM-heavy ideas.
Worked example, step by step
Faster decoding sounds like a pure win, but how much of it does a customer actually feel? Let us build a latency budget for our shop's assistant and compare two endpoints: A streams 40 tokens per second and B streams 400. Both have the same fixed costs: 0.1 s of network time and 0.3 s until the first token. Every figure here is illustrative, chosen to make the arithmetic easy; none is a measurement of any real service.
Three situations, two endpoints
- A 300-token answer on A: 0.1 + 0.3 + 300 / 40 = 0.4 + 7.5 = 7.9 s. Almost all of the wait is decoding.
- The same answer on B: 0.1 + 0.3 + 300 / 400 = 0.4 + 0.75 = 1.15 s. Decoding got 10× faster and the whole answer about 6.9× faster.
- A 20-token reply: A: 0.4 + 0.5 = 0.9 s. B: 0.4 + 0.05 = 0.45 s. Only 2× faster. The fixed 0.4 s is now most of the wait, and no decode speed can remove it.
- An agent that chains 6 calls: Each call writes 150 tokens and must finish before the next begins. A: 6 × (0.4 + 3.75) = 24.9 s. B: 6 × (0.4 + 0.375) = 4.65 s. Waiting adds up across the chain, so speed compounds.
- Ask who waits for the last token: A person reading a streamed answer mostly cares when the first words appear. A program that needs the complete output before it can act (the next agent step, a tool call, a sentence to be spoken) waits for the final token. That is where decode speed pays off most.
The practical rule: before paying for per-user speed, write down the budget. If most of the wait is fixed overhead, shorten the prompt or move closer to the server first. If most of it is output tokens, and something is waiting for the last one, faster decoding is the right lever.
Practice: try it yourself
We will turn the budget into a small calculator. It needs nothing but plain Python. We compute the wait for a single answer, for an agent chain, and then ask two sharper questions: how much of each wait is fixed overhead, and how much of a 10× faster decoder the user really feels.
practice_latency_budget.py
# A latency budget for our support assistant. All numbers are illustrative.
def answer_time(tokens, tokens_per_s, first_token_s=0.3, network_s=0.1):
"""Seconds until the full answer has arrived for one LLM call."""
return network_s + first_token_s + tokens / tokens_per_s
speeds = {"endpoint A (40 tok/s)": 40, "endpoint B (400 tok/s)": 400}
# 1) One chat answer of 300 tokens
print("one 300-token answer")
for name, tps in speeds.items():
print(f" {name:23s}: {answer_time(300, tps):5.2f} s")
# 2) An agent that chains 6 calls of 150 tokens; each waits for the last
print("agent workflow: 6 calls x 150 tokens, one after another")
for name, tps in speeds.items():
total = sum(answer_time(150, tps) for _ in range(6))
print(f" {name:23s}: {total:5.2f} s")
# 3) How much of the wait is NOT token generation?
print("share of the wait that is fixed overhead (network + first token)")
for tokens in (20, 300):
for name, tps in speeds.items():
total = answer_time(tokens, tps)
fixed = 0.3 + 0.1
print(f" {tokens:3d} tokens, {name:23s}: {fixed / total:4.0%}")
# 4) Speed-up actually felt by the user when the chip is 10x faster
for tokens in (20, 300, 2000):
gain = answer_time(tokens, 40) / answer_time(tokens, 400)
print(f"10x faster decoding, {tokens:4d}-token answer: {gain:4.1f}x faster overall")Output:
one 300-token answer endpoint A (40 tok/s) : 7.90 s endpoint B (400 tok/s) : 1.15 s agent workflow: 6 calls x 150 tokens, one after another endpoint A (40 tok/s) : 24.90 s endpoint B (400 tok/s) : 4.65 s share of the wait that is fixed overhead (network + first token) 20 tokens, endpoint A (40 tok/s) : 44% 20 tokens, endpoint B (400 tok/s) : 89% 300 tokens, endpoint A (40 tok/s) : 5% 300 tokens, endpoint B (400 tok/s) : 35% 10x faster decoding, 20-token answer: 2.0x faster overall 10x faster decoding, 300-token answer: 6.9x faster overall 10x faster decoding, 2000-token answer: 9.3x faster overall
Now change it:
- Change the default
first_token_sto1.5, as if every call carried a very long prompt. Predict whether the overall gain for a 300-token answer goes up or down from 6.9×. - Add a third entry,
"endpoint C (4000 tok/s)": 4000. Predict the time for one 300-token answer. Is C ten times better than B for the user? - Change the agent to 20 calls of 50 tokens each. Predict both totals. Does the fast endpoint's advantage grow or shrink compared with 6 calls of 150?
Pause and think: In the run above, a 10× faster decoder made a 20-token reply only 2.0× faster. Why, and what would we change to speed up short replies?
Only the 'tokens ÷ speed' part of the budget shrinks. For 20 tokens that part was 0.5 s out of 0.9 s; the other 0.4 s (network plus time to first token) did not move. Short replies are limited by fixed overhead, so the useful changes are a shorter prompt (less prefill work), a server closer to the user, or reusing connections. More decode speed gives little.
Pause and think: Every night we summarise 5,000 old tickets. Nobody is watching the output appear. Is tokens per second per user the right number to optimise?
No. Per-user speed matters when someone, or some program, is waiting on one stream. For an overnight batch what matters is total cost and total throughput: how many tokens the whole system produces per hour and per unit of money. Batching many requests together trades a slower individual stream for more total work, which is exactly the trade this job wants. That is why this lesson points such jobs to GPUs.
Key takeaways
- LLM decoding is sequential and memory-bandwidth-bound: each token reads all the weights for very little maths.
- The LPU keeps weights in on-chip SRAM, trading capacity for enormous bandwidth, so models span many chips.
- A compiler schedules every operation and every chip-to-chip transfer ahead of time: no caches, no guessing, no waiting.
- Data streams through slices of functional units like an assembly line, across chip boundaries.
- LPUs win on per-user speed for supported models; GPUs win on flexibility, training, memory capacity and batched cost.
Key terms
- LPU: Language Processing Unit: Groq's inference chip design for fast LLM token generation.
- Decode phase: The token-by-token part of LLM generation, where each step depends on the previous token.
- SRAM: Fast static memory built on the processor chip; high bandwidth but small capacity.
- HBM: High Bandwidth Memory: stacked DRAM next to a GPU or TPU, large but slower than on-chip SRAM.
- Deterministic execution: Running a program whose timing is fully planned in advance, so it takes the same cycles every time.
- Static scheduling: A compiler deciding ahead of time which unit does what on which cycle, instead of hardware deciding at run time.
- Tensor Streaming Processor: Groq's architecture in which data streams across slices of specialised functional units.
← 17.3 Google TPUs: Purpose-Built Hardware for Neural Networks · 17.5 Cloud vs Edge: Where Should Your Model Run? →