Modern AI Engineering

What Is a Large Language Model (LLM)? Explained Simply

By Modern AI Engineering · · 7 min read

A large language model, or LLM, is a computer program that has learned the patterns of language from a very large amount of text. You give it some text, called a prompt, and it continues with the text that is most likely to follow. Chat assistants, coding helpers and many search tools are built on LLMs.

This article explains what the three words mean, what an LLM can and cannot do, how one is made, and the main types. It stays at a simple level. A separate article on this site goes step by step through the inner workings.

What does LLM stand for?

LLM stands for large language model. Each word tells you something.

Model means a mathematical system that has learned from examples. It is a neural network: many layers of simple units joined by adjustable numbers called parameters or weights.

Language means its material is text. It learned from books, articles, websites and code, and it reads and writes text.

Large refers to size in two ways. The network has a very large number of parameters, and it was trained on a very large amount of text. Size matters because a larger network trained on more text can hold more patterns.

What does a large language model actually do?

It predicts the next piece of text. That is the whole task. Given the words so far, the model works out which piece is likely to come next, adds it, and repeats.

A simple comparison is the word suggestion on a phone keyboard. An LLM does the same kind of thing, but with a far larger network and far more context. It can take pages of text into account, not only the last few words.

It may seem strange that prediction alone produces useful answers. The reason is that predicting text well requires a lot. To continue a sentence about history, it helps to know the history. To continue a piece of code, it helps to know how the language works. Training pushes that knowledge into the weights.

How is a large language model trained?

Training happens in stages.

  1. Pre-training. The model is shown huge amounts of text with the next piece hidden, and asked to predict it. Each wrong guess leads to a small correction of the weights. This stage needs a lot of computing power and is done by a small number of organisations.
  2. Instruction training. The model is trained further on examples of questions and good answers, so it learns to follow a request instead of only continuing text.
  3. Alignment. People compare answers, and the model is adjusted towards the ones they prefer, so it becomes more helpful and safer.
  4. Fine-tuning, when needed. A company can train the model a little more on its own examples for a narrow task.

What can an LLM do, and what can it not do?

An LLM is strong wherever the task can be done by reading and writing text. It is weak wherever the task needs facts it never saw, exact calculation or a guarantee of truth.

Good atWeak at
Writing, rewriting and summarising textKnowing events after its training data ends
Answering questions on well-covered topicsKnowing your private documents, unless you supply them
Explaining and writing codeExact arithmetic with long numbers
Translating between languagesCounting letters or characters reliably
Pulling structured data out of messy textSaying when it does not know
Following instructions about tone and formatGiving the same answer every time

Why do LLMs make mistakes?

An LLM produces text that is likely, and likely is not the same as true. There is no step inside the model that checks a fact against a source. When the model lacks the information, it can still write a fluent answer that is wrong. This is called a hallucination.

The model also has a knowledge cutoff. Its weights hold only what was in the training text, so newer facts are missing.

Engineers work around both limits. They give the model source documents to answer from, a method called retrieval-augmented generation or RAG. They give it tools, such as a calculator or a search function. And they test its answers on a fixed set of questions before trusting it in a product.

What are the main types of LLM?

The word covers a family of models. A few distinctions come up often.

  • Closed and open models. Closed models are reached through an API run by the company that made them. Open models publish their weights, so you can run them on your own hardware.
  • Large and small models. Small language models have fewer parameters. They are cheaper and faster, and can run on a laptop or a phone, at some cost in ability.
  • Reasoning models. These spend extra steps working through a problem before they give the final answer. They help on hard tasks and cost more time.
  • Multimodal models. These accept images or audio as well as text.
  • Embedding models. These do not write text. They turn text into vectors that capture meaning, which search systems use.

LLM vs generative AI vs AI: how the terms relate

The three terms sit inside each other. AI is the widest. It covers any system that performs a task we link with intelligence. Generative AI is the part of AI that creates new content, such as text, images or audio. An LLM is one kind of generative AI, the kind that works with text.

So every LLM is generative AI, but an image generator is generative AI without being an LLM. And a chat product is not the same thing as the model. The product is an application built around an LLM. It adds a conversation history, instructions, tools and safety checks.

How to learn more about LLMs

If you want to use LLMs well, learn three things next: what a token is, what the context window is, and what temperature does. They explain most of the behaviour you see day to day.

If you want to build with them, go one layer deeper: embeddings, attention and the Transformer. In the AI Engineering Bootcamp, Module 4 covers how the model is built, Module 5 covers how it produces output, and Module 7 covers the different types of language model. The first lesson of the course is free with a free account and introduces LLMs alongside five other core terms.

Frequently asked questions

What is an LLM in simple words?

It is a program trained on a very large amount of text to predict what comes next. By repeating that prediction it can write answers, summaries, translations and code.

Is ChatGPT an LLM?

ChatGPT is a product built on LLMs. The model generates the text. The product around it adds the chat interface, memory of the conversation, tools and safety checks.

What is the difference between an LLM and generative AI?

Generative AI is any AI that creates content, including images, audio and video. An LLM is the kind of generative AI that reads and writes text.

Do LLMs learn from my conversations as I chat?

The weights of the model do not change while you chat. The model only sees the text placed in its context for that request. Whether a provider uses conversations for later training depends on its policy and your settings.

Why is it called a large language model?

Because the network has a very large number of parameters and was trained on a very large amount of text. Both are far beyond earlier language models.

Learn it properly: the AI Engineering Bootcamp

More articles