Modern AI Engineering

Lesson 4.1 · 27 min

Generative AI: Creating Instead of Classifying

How can a program that was only ever shown existing text, images and sound produce a sentence, picture or song that never existed before?

In short: Generative AI is AI that creates new content (text, images, audio, video, code) instead of only labelling or scoring existing content. It learns the patterns of a huge amount of example data, stores them as numbers called parameters, and then creates something new by sampling from what it learned, one small piece at a time. It is powerful and useful every day, but it can be confidently wrong, biased and out of date, so we must use it with care.

What is Generative AI?

Generative AI is a kind of artificial intelligence that creates new content. We give it a request, usually written in plain language, and it produces a reply: an email draft, an answer to a question, a picture of a cat in a spacesuit, a short melody, or a working Python function.

The name splits neatly into two parts. Generative means “able to produce or bring something into being”. AI (artificial intelligence) means software that does tasks we normally associate with human thinking, such as understanding language or recognising objects. Put together: software that produces new things in a way that looks intelligent.

Throughout this lesson we will follow one running example: a support chatbot for an online shop. A customer types “Where is my order?” and the bot writes a friendly, fresh reply. Nobody wrote that exact reply in advance. The bot generated it.

Think of it like a well-read storyteller Imagine someone who has read millions of books but memorised none of them word for word. Ask them for a story about a lost umbrella and they can tell a new one, because they absorbed how stories usually go: how sentences flow, what happens next, how they end. Generative AI works the same way. It learns the patterns of its examples, not a copy of each one.

What does “generate” mean here? It does not mean looking up a stored answer or copying and pasting. It means building an output piece by piece, where each piece is chosen because it is likely given everything before it. For text, a piece is a token: a word or a part of a word. For images, the model may refine a whole picture step by step. The output is new in the sense that it is freshly assembled, even though every pattern inside it came from training data.

How is Generative AI different from the old AI?

Before generative AI became popular, most AI in products was discriminative (also called predictive or analytical AI). A discriminative model looks at an input and gives back a label or a number: spam or not spam, which digit is in this photo, how likely is this card payment to be fraud, what price will this house sell for. Its output is small and comes from a fixed set of answers.

A generative model gives back a whole new piece of content. Its output is large and open-ended: a paragraph, an image with millions of pixels, a minute of audio. In probability terms, a discriminative model learns “given this input, which label?”, while a generative model learns “what does realistic data look like?” so that it can produce more of it.

Pause and think: Our shop wants to flag which incoming emails are angry so a human answers them first. Is that a generative or a discriminative task?

Discriminative. The output is a label (angry or not angry), chosen from a fixed set. A generative model could do it too if we ask it to answer with one word, but the task itself is classification.

How does Generative AI learn?

A generative model learns by looking at an enormous number of examples and slowly adjusting itself until it can predict those examples well. For a text model the main exercise is surprisingly simple: hide the next word and guess it. “The order has ___”. If the model guesses badly, a training algorithm nudges its internal numbers so that next time the right word gets a bit more probability. Repeat this across trillions of words and the model ends up absorbing grammar, facts, styles and reasoning patterns, because all of them help predict the next word.

This is called self-supervised learning: the labels come from the data itself (the real next word is the label), so nobody has to annotate anything by hand. That is why generative models can learn from so much data.

The usual training recipe for a text model

  1. Collect data: Gather a very large collection of text, such as web pages, books and code, and clean it (remove duplicates, spam and unsafe content).
  2. Tokenize: Cut the text into tokens (words or word pieces) and turn each token into a number the model can work with.
  3. Pre-train: Show the model billions of snippets and ask it to predict each next token. Measure how wrong it is with a loss function and adjust its parameters with gradient descent.
  4. Fine-tune: Train further on smaller, curated examples of good behaviour, such as helpful question-and-answer pairs, so it follows instructions instead of just continuing text.
  5. Align with feedback: Use human or AI preferences (for example, reinforcement learning from human feedback, RLHF) to make answers more helpful, honest and safe.

Image generators learn with a different exercise. A popular one is used by diffusion models: take a real picture, add random noise, and train the model to remove the noise. After learning to clean up noise at every level, the model can start from pure noise and clean it into a brand-new image. Audio and video models use similar ideas. The common thread is: learn the patterns of real data well enough to recreate data like it.

How does Generative AI actually create something new?

After training, a text model works like a very smart autocomplete. Given the text so far, it outputs a probability distribution: a list of every possible next token with a chance attached. For “Your order has” it might give been 0.55, shipped 0.30, arrived 0.10 and tiny chances to thousands of other tokens.

The program then samples: it picks one token at random, weighted by those chances, appends it to the text and asks the model again. Word by word, a new reply appears. Because the pick is random, asking the same question twice can give two different, equally reasonable answers. A setting called temperature controls how adventurous the pick is. Low temperature makes the most likely token win almost every time (safe, repetitive). High temperature flattens the chances so rarer tokens get picked more (creative, but riskier).

Why is the result new? Because the model combines patterns, not stored sentences. It learned that “order” often comes before “has shipped” and that “your order” starts many replies. Chaining these choices can produce a sentence that appears nowhere in its training data. The code later in this lesson shows exactly this happening with a tiny model.

Pause and think: We set temperature to 0 (always pick the most likely token) and ask the same question twice. What happens?

We get the same reply both times (in practice, almost always: tiny numerical differences on some hardware can still cause rare changes). With no randomness in the pick, the model's top choice wins at every step, so the output is effectively deterministic.

What can Generative AI create?

Anything that can be turned into numbers and has learnable patterns can, in principle, be generated. Today the main families are below. A model that handles several of these at once (for example, it reads an image and writes about it) is called multimodal.

Common kinds of generated content
ContentTypical model styleEveryday example
TextLarge language model (LLM) predicting next tokensChat assistants, email drafts, summaries
CodeLLM trained heavily on source codeCoding assistants that suggest the next lines
ImagesDiffusion model (noise → picture), guided by a text promptIllustrations, product mock-ups, image edits
Audio and speechToken-based or diffusion audio modelsText-to-speech voices, short music clips
VideoDiffusion or transformer models over framesShort clips from a text description
Structured dataLLM asked for JSON or tablesFilling forms, extracting fields from documents

A short history of generative AI

  1. GANs: Generative Adversarial Networks pit a generator against a critic and produce the first convincing synthetic faces.
  2. The Transformer: The paper “Attention Is All You Need” introduces the architecture behind nearly every modern language model.
  3. GPT-1, GPT-2, GPT-3: OpenAI shows that a bigger next-word predictor trained on more text gets steadily more capable; GPT-3 has 175 billion parameters.
  4. Diffusion models: Denoising diffusion becomes the leading way to generate images; Stable Diffusion is released openly in 2022.
  5. ChatGPT: A chat interface on an instruction-tuned LLM brings generative AI to hundreds of millions of people.
  6. Multimodal and open models: Models that read images and audio, plus strong open-weight families, make generative AI a standard building block.

What is a model in Generative AI?

A model is the trained program that does the generating. Under the hood it is a big mathematical function, almost always a neural network, whose behaviour is set by millions or billions of numbers called parameters (or weights). Training is simply the process of finding good values for those numbers.

A helpful way to think about it: the model file is a compressed summary of patterns in the training data. It does not contain a searchable copy of the internet. GPT-2 (2019) had 1.5 billion parameters and GPT-3 (2020) had 175 billion; many current models are of similar or larger size, though some vendors do not publish their counts. More parameters usually means more capacity to store patterns, but also more memory and compute to run.

  • Architecture: the shape of the network (for modern text models, the Transformer).
  • Parameters: the learned numbers inside that shape.
  • Inference: running the trained model to produce output (as opposed to training it).
  • Prompt: the input we give the model at inference time.
  • Foundation model: a large model pre-trained on broad data that can be adapted to many tasks.

Model vs product A chatbot app is a product built around a model. The product adds a user interface, safety filters, memory of the conversation, tools such as web search, and a system prompt. The same model can power many different products.

The complete flow of Generative AI

Let us follow our customer's message “Where is my order?” all the way through a text model. Click each node to read what happens there.

Image generation follows the same broad shape (prompt in, content out), but the middle is different: the prompt is encoded into vectors that guide a diffusion model as it removes noise over many steps.

Code: a tiny generative model you can run

Real LLMs are huge, but the core idea fits in 40 lines. We train a bigram model on five support sentences: it only learns which word tends to follow which. Then we generate by sampling one word at a time. Watch for outputs that were never in the training data.

tiny_generator.py

import random
from collections import defaultdict, Counter
# 1. Training data: a tiny "internet" of support-chat sentences
corpus = [
"the order has shipped today",
"the order is on the way",
"the refund has been sent",
"the refund is on the way",
"your order has been sent today",
]
# 2. Learning: count which word follows which (a bigram model)
counts = defaultdict(Counter)
for line in corpus:
words = ["<s>"] + line.split() + ["</s>"]
for prev, nxt in zip(words, words[1:]):
counts[prev][nxt] += 1
# The learned "parameters": P(next word | previous word)
def next_probs(word):
total = sum(counts[word].values())
return {w: round(c / total, 2) for w, c in counts[word].items()}
print("P(next | 'the')  =", next_probs("the"))
print("P(next | 'has')  =", next_probs("has"))
# 3. Generating: sample one word at a time until the end token
def generate(seed):
random.seed(seed)
word, out = "<s>", []
while True:
probs = next_probs(word)
word = random.choices(list(probs), weights=list(probs.values()))[0]
if word == "</s>":
return " ".join(out)
out.append(word)
for s in range(5):
text = generate(s)
tag = "(copied)" if text in corpus else "(NEW)"
print(f"sample {s}: {text:32s} {tag}")

Output:

P(next | 'the')  = {'order': 0.33, 'way': 0.33, 'refund': 0.33}
P(next | 'has')  = {'shipped': 0.33, 'been': 0.67}
sample 0: your order has shipped today     (NEW)
sample 1: the refund is on the way         (copied)
sample 2: your order has shipped today     (NEW)
sample 3: the way                          (NEW)
sample 4: the order has shipped today      (copied)

Look at the output. “your order has shipped today” is new: the training data only had “your order has been sent today” and “the order has shipped today”. The model recombined patterns. But “the way” is also new, and it is nonsense as a reply: every word pair in it was seen in training, yet the whole is useless. This is a tiny version of a real problem. A model can produce something that is locally fluent but wrong. Large models are vastly better, because they look at much more context than one word, yet the same risk remains.

Where do we use Generative AI every day?

  • Chat assistants that answer questions, explain topics and brainstorm.
  • Writing help: drafting emails, rewriting text in a different tone, translating, summarising long documents.
  • Coding assistants that complete lines, explain errors and write tests.
  • Customer support bots like ours that draft or send replies, often grounded in the company's help pages.
  • Search and research tools that read several sources and write a combined answer with links.
  • Creative tools for images, voice-overs, music sketches and video clips.
  • Office features such as meeting notes, auto-generated slide drafts and spreadsheet formulas from a description.

Our support bot in production A realistic setup: the shop's help articles are searched for the passage most related to the customer's question, that passage is pasted into the prompt, and the model writes a reply based on it. A human agent reviews replies about refunds before they are sent. This mix of generation, retrieved facts and human review is how many companies use generative AI safely.

The limitations we must know

Generative AI produces what is likely, not what is verified. Almost every limitation comes from that one fact.

  • Hallucination: the model can state false facts, fake citations or invented order numbers in a confident tone.
  • Knowledge cutoff: it only knows what was in its training data, which stops at some date, unless we give it fresh information in the prompt.
  • Bias: patterns in the training data, including unfair stereotypes, can show up in the output.
  • Inconsistency: the same prompt can give different answers, which makes testing harder.
  • Cost and speed: large models need powerful hardware; long outputs take time and money.
  • Privacy, copyright and misuse: sensitive data in prompts, ownership questions about training data and outputs, and deepfakes are real concerns that vary by country and are still evolving legally.

The most common mistake Treating fluent output as proof of truth. Never let a generative model be the only check for facts, numbers, legal or medical advice, or actions with real consequences (refunds, payments, deletions). Ground it in trusted data and keep a human or a rule-based check in the loop.

When not to use it: if a simple rule or a lookup gives an exact answer (today's order status from the database, a tax calculation), use that instead. Generative AI is best when the output is language or media, when some variation is fine, and when a mistake can be caught.

Worked example, step by step

Our bigram model printed “the way” as a reply. How likely was that, compared with a sensible reply? We can work it out by hand. The chance of a whole reply is the chance of each word given the word before it, all multiplied together. We use the counts from the five training sentences.

How likely is “your order has shipped today”?

  1. First word: Four training sentences start with “the” and one starts with “your”. So P(your | start) = 1/5 = 0.2.
  2. your → order: “your” was only ever followed by “order”. P = 1.0. Running total: 0.2.
  3. order → has: “order” was followed by “has” twice and “is” once. P = 2/3. Running total: 0.2 × 0.667 ≈ 0.133.
  4. has → shipped: “has” was followed by “been” twice and “shipped” once. P = 1/3. Running total: ≈ 0.044.
  5. shipped → today → end: Both steps had only one option in training, so each is 1.0. Final answer: about 0.044, or 4.4%.

Now the surprise. “the way” needs only three steps: P(the | start) = 4/5, P(way | the) = 2/6, and P(end | way) = 1.0. That gives 0.8 × 0.333 × 1.0 ≈ 0.267. The nonsense reply is the single most likely output of this model.

Reply probabilities under the bigram model, computed by hand from the training counts.
ReplyFactorsProbability
the way0.8 × 1/3 × 1≈ 0.267
the order has shipped today0.8 × 1/3 × 2/3 × 1/3 × 1 × 1≈ 0.059
your order has shipped today0.2 × 1 × 2/3 × 1/3 × 1 × 1≈ 0.044
the refund is on the way0.8 × 1/3 × 1/2 × 1 × 1 × 1/3 × 1≈ 0.044

Two lessons hide in this table. First, short outputs have fewer factors below 1, so a weak model favours them. Second, the model only sees one word back. After “the” it cannot tell whether “on” came before, so it treats “the way” at the start of a reply as normal. More context is the cure, and that is exactly what large models add.

Practice: try it yourself

We will build the sampling step on its own. We give four candidate next words a score, turn the scores into probabilities at three temperatures, and then draw 1,000 words each time to see what the customer would actually get.

practice_sampling.py

import math
import random
# Illustrative scores (logits) for the word after "Your order has"
logits = {"been": 2.0, "shipped": 1.4, "arrived": 0.3, "exploded": -1.5}
def softmax(scores, temperature):
# Divide each score by the temperature, then turn scores into probabilities
exps = {w: math.exp(s / temperature) for w, s in scores.items()}
total = sum(exps.values())
return {w: e / total for w, e in exps.items()}
def sample_counts(temperature, draws=1000, seed=7):
# Draw many next words and count how often each one is picked
random.seed(seed)
probs = softmax(logits, temperature)
picks = random.choices(list(probs), weights=list(probs.values()), k=draws)
return probs, {w: picks.count(w) for w in probs}
for t in [0.2, 1.0, 3.0]:
probs, counts = sample_counts(t)
print(f"temperature {t}")
for w in logits:
print(f"  {w:9s} p = {probs[w]:.3f}   picked {counts[w]:4d} / 1000")

Output:

temperature 0.2
been      p = 0.952   picked  946 / 1000
shipped   p = 0.047   picked   54 / 1000
arrived   p = 0.000   picked    0 / 1000
exploded  p = 0.000   picked    0 / 1000
temperature 1.0
been      p = 0.568   picked  608 / 1000
shipped   p = 0.312   picked  263 / 1000
arrived   p = 0.104   picked  115 / 1000
exploded  p = 0.017   picked   14 / 1000
temperature 3.0
been      p = 0.371   picked  404 / 1000
shipped   p = 0.304   picked  289 / 1000
arrived   p = 0.210   picked  187 / 1000
exploded  p = 0.115   picked  120 / 1000

Now change it:

  • Add 0.05 to the temperature list. Predict first: how many of the 1,000 picks will be “been”?
  • Change draws=1000 to draws=10 and fix the last print to match. Predict: will the counts still look like the probabilities?
  • Raise the score of “exploded” from -1.5 to 2.5. Predict which word wins at temperature 1.0, and whether a low temperature makes the bad word rarer or more common.

Pause and think: At temperature 3.0 the word “exploded” was picked 120 times in 1,000, but only 14 times at temperature 1.0. The model's scores did not change. What does this tell us about where risky output can come from?

It can come from the sampling settings, not only from the model. A high temperature flattens the probabilities, so words the model itself rated as unlikely get picked far more often. The same model can be safe or sloppy depending on how we sample from it.

Pause and think: At temperature 1.0, “been” has p = 0.568 but was picked 608 times, not 568. Is that a bug?

No. Sampling is random, so counts wobble around probability × draws. With 1,000 draws a gap of a few dozen is normal. With more draws the share gets closer to 0.568; with only 10 draws it can be far off.

Summary

Generative AI creates new content. It learns the patterns of huge amounts of data by practising prediction (for text: guess the next token), stores those patterns as parameters inside a model, and creates by sampling one piece at a time. It differs from classic discriminative AI, which only labels or scores inputs. We use it daily for writing, coding, support and media, and we keep its limits in view: it can be wrong, biased and out of date, so we ground it and check it.

Key takeaways

  • Generative AI creates new content; discriminative (classic) AI labels or scores existing input.
  • Models learn by practising prediction on huge data, mostly self-supervised, and store patterns in parameters.
  • Text models generate one token at a time by sampling from a probability distribution; temperature controls randomness.
  • New output comes from recombining learned patterns, which is also why it can be fluent but false.
  • Ground outputs in trusted data and keep checks in place for facts and consequential actions.

Key terms

  • Generative AI: AI that produces new content such as text, images, audio, video or code.
  • Discriminative model: A model that maps an input to a label or number, such as spam or not spam.
  • Token: A small piece of text, a word or part of a word, that a language model reads and writes.
  • Parameters: The learned numbers inside a model that determine its behaviour.
  • Sampling: Picking the next piece of output at random, weighted by the model's probabilities.
  • Temperature: A setting that sharpens (low) or flattens (high) the probabilities before sampling.
  • Hallucination: Fluent, confident output that is false or not supported by any source.

← 3.10 TensorFlow Explained: Static Graphs and Production ML · 4.2 Autoregressive Models: Predicting One Token at a Time →