How LLMs Work: Tokens, Embeddings and Attention Explained
By Modern AI Engineering · · 8 min read
A large language model, or LLM, does one thing: it predicts the next piece of text. It reads what has been written so far, works out which piece is likely to come next, adds it, and repeats. A whole answer is built this way, one piece at a time.
To make each prediction, the model takes four steps. It splits text into tokens. It turns tokens into lists of numbers called embeddings. It passes them through layers of attention, so each token can use the others as context. Then it produces a probability for every possible next token and picks one. This article explains each step.
What is an LLM?
An LLM is a very large neural network trained on a huge amount of text. Large refers to the number of parameters, the adjustable numbers inside the network. Language model means it models which text tends to follow which.
It is not a database of sentences and it does not look answers up. What it learned during training is stored as numbers in its weights. When you send a prompt, it calculates a response from those numbers.
What is a token?
A model cannot read letters. It works with numbers, so text is first cut into small pieces called tokens, and each token gets an ID number.
A token is often a word, but not always. Common words tend to be one token. Rare or long words are split into several parts. Punctuation and spaces count too. This subword approach lets the model handle any text, including words it has never seen, with a vocabulary of fixed size. A common method for building the vocabulary is byte pair encoding, or BPE, which starts from small units and repeatedly merges the pairs that appear together most often.
Tokens matter in practice. Model limits and prices are counted in tokens, not words. The context window, the amount of text a model can consider at once, is also measured in tokens.
What are embeddings?
A token ID is only a label. The number says nothing about meaning. So the model looks up each ID in a table and replaces it with an embedding, a long list of numbers.
You can think of an embedding as a location in a space with many dimensions. During training, tokens used in similar ways end up near each other. Words for animals gather in one region, words for cities in another. Directions in the space carry meaning as well.
The model also needs to know the order of the tokens, because the same words in a different order can mean something else. Position information is added to the embeddings so the model can tell first from last.
Embeddings are useful outside the model too. Search systems and RAG use them to find text with a similar meaning, even when the wording is different.
What is attention and why does it matter?
An embedding gives a token one fixed meaning, but words depend on context. The word bank means one thing next to river and another next to loan. Attention is the mechanism that fixes this.
In self-attention, each token looks at the other tokens and decides how relevant each one is to it. It then pulls in information from the relevant ones, weighted by that relevance. After this step, the vector for bank is no longer generic. It has been shaped by the words around it.
The model does this with several attention heads at once. Each head can track a different kind of relationship, such as which noun a pronoun refers to, or which verb goes with which subject. This is called multi-head attention.
In a model that generates text, a token may only look at tokens before it, never after. This rule is called causal masking. It matches how the model is used: when writing, the future does not exist yet.
What happens inside a Transformer layer?
The architecture that combines these ideas is the Transformer. It is a stack of identical layers, and each layer has two main parts.
The first is attention, where tokens exchange information with each other. The second is a feed-forward network, which processes each token separately and holds much of what the model has learned.
The token vectors pass through layer after layer. With each one they become richer. Early layers tend to capture simple things such as grammar. Later layers capture more abstract meaning. By the top of the stack, the vector at the last position holds what the model needs to predict what comes next.
How does an LLM generate text?
At the end of the stack the model produces a score for every token in its vocabulary. These scores are turned into probabilities. Then one token is chosen.
How it is chosen is called sampling. Always taking the most likely token gives safe and repetitive text. Temperature controls how much the model favours likely tokens. A low temperature makes output focused and predictable. A high temperature makes it more varied. Top-k and top-p sampling cut off the unlikely tokens before choosing, so the model does not pick something absurd.
The chosen token is added to the input and the whole process runs again for the next one. This is called autoregressive generation. It is why answers appear word by word on the screen, and why longer answers take longer.
How is an LLM trained?
Training uses the same task as generation. The model is shown text with the next token hidden and asked to predict it. Its prediction is compared with the real token, the error is measured with a loss function, and backpropagation adjusts the weights slightly. Repeated over an enormous amount of text, this teaches grammar, facts and patterns of reasoning, because all of them help predict what comes next.
This first stage is called pre-training. Afterwards, models are usually trained further to follow instructions and to answer in ways people find helpful. Fine-tuning for a specific task is a smaller version of the same process.
What does this explain about LLM behaviour?
Once you know the mechanism, several familiar behaviours make sense.
- Hallucinations. The model produces likely text. Likely is not the same as true, and nothing inside checks facts.
- Different answers to the same prompt. Sampling involves chance unless the temperature is very low.
- Knowledge cutoff. The weights hold only what was in the training data. Newer facts must be supplied in the prompt, for example through RAG.
- Limits on input length. Attention compares tokens with each other, so cost grows quickly as the context gets longer.
- Trouble with exact letters or counting. The model sees tokens, not characters.
Frequently asked questions
How does an LLM work in simple terms?
It turns text into tokens, turns tokens into numbers, uses attention to understand each token in context, and predicts the next token. It repeats that prediction to write a full answer.
Is a token the same as a word?
No. A token can be a whole word, part of a word, a punctuation mark or a space. Common words are often one token, and rare words are split into several.
Do LLMs understand language?
They build rich internal representations that let them use language in useful ways. Whether that counts as understanding is debated. In practice, treat the output as a strong prediction and verify what matters.
What is the difference between an LLM and a Transformer?
The Transformer is the architecture, the design of the network. An LLM is a large model built with that design and trained on text.
Why do LLMs make things up?
They generate text that is statistically likely, and they have no built-in step that checks facts. Giving the model source documents and asking it to answer from them reduces the problem.