Deep Learning Explained: Neural Networks to Transformers
By Modern AI Engineering · · 8 min read
Deep learning is a way of teaching computers with neural networks that have many layers. Each layer takes numbers in, transforms them and passes them on. The network learns by comparing its output with the right answer and adjusting its internal numbers a little, over and over, until the error is small.
This article walks from a single neuron to the Transformer, the design behind modern language models. No advanced math is needed to follow it.
What is a neural network?
Start with one neuron. It receives a few numbers as input. It multiplies each one by a weight, adds them up, and adds one more number called the bias. Then it passes the total through a simple function called an activation, which decides how strongly the neuron fires.
The weights say how much each input matters. The bias shifts the result up or down, so the neuron can fire even when the inputs are small. The activation adds a bend. Without that bend, any stack of layers would behave like a single straight-line model, however many layers it had.
A layer is many neurons side by side. A network is many layers in a row. The word deep only means there are many layers.
How does a neural network learn?
Learning happens in a loop with four parts.
- Forward pass. The input flows through the layers and the network produces a prediction.
- Loss. A loss function turns the difference between the prediction and the right answer into a single number. For classification the usual choice is cross-entropy loss.
- Backpropagation. The network works out how much each weight contributed to the error, moving backwards from the output to the input.
- Update. Gradient descent changes every weight a little in the direction that lowers the loss. The size of the change is set by the learning rate.
What are gradient descent and backpropagation?
Picture the loss as a hilly landscape. Every possible setting of the weights is a place on it, and the height is the error. Training means walking downhill. Gradient descent looks at the slope where you stand and takes a step in the steepest downward direction.
Backpropagation is how the slope is calculated. A network has a huge number of weights, and you need to know how the error changes with each one. Backpropagation uses the chain rule from calculus to pass the error backwards layer by layer, reusing work as it goes, so the whole calculation is fast.
People often mix the two up. Backpropagation computes the slopes. Gradient descent uses them to update the weights.
Why do deep networks need dropout and normalization?
Two problems appear as networks grow.
The first is overfitting. A large network can memorise its training data and then fail on new data. Dropout fights this by switching off a random share of neurons during each training step. The network cannot lean on any single neuron, so it learns patterns that hold more widely.
The second is unstable training. As numbers pass through many layers they can grow very large or shrink towards zero, and learning slows or breaks. Normalization layers keep the numbers in a steady range. Batch normalization does this across a batch of examples. Layer normalization does it within one example, which suits language models, and newer models often use a lighter version called RMSNorm.
From feed-forward networks to CNNs and RNNs
The plain network described so far is called a feed-forward network. It treats its input as one flat list of numbers. That works for tables, but it wastes the structure in images and text.
Convolutional neural networks, or CNNs, were designed for images. They slide small filters across the picture, so the same pattern detector is reused everywhere. They became the standard for computer vision.
Recurrent neural networks, or RNNs, were designed for sequences such as sentences. An RNN reads one item at a time and carries a hidden state forward as its memory of what came before. This fits language naturally, but it has two weaknesses. It must process tokens one after another, which is slow. And information from far back in the sequence fades by the time it is needed.
What is a Transformer and why did it replace RNNs?
The Transformer removed the step-by-step reading. Instead of passing a memory along the sequence, it lets every token look directly at every other token. This mechanism is called self-attention.
For each token, attention asks which other tokens are relevant right now and by how much, then blends their information in. A word at the end of a paragraph can use a word from the start in a single step. Nothing has to survive a long chain.
Because all tokens are processed together, training can run in parallel on GPUs. That made it practical to train on far more text than before, and it is the main reason large language models became possible. The same design now handles images too, by cutting a picture into patches and treating the patches as tokens.
How do you learn deep learning as a beginner?
Follow the order of this article. Begin with a single neuron and what the bias does. Work through gradient descent with small numbers. Follow one full example of backpropagation by hand. Then learn cross-entropy loss, dropout and normalization. After that, study RNNs briefly, so you see the problem that attention solves, and move on to the Transformer.
A little hands-on work helps a great deal. Build a tiny network in PyTorch, train it, and watch the loss fall. Then change the learning rate and see what breaks.
Module 3 of the AI Engineering Bootcamp teaches these topics in ten lessons, and Module 4 goes inside the Transformer piece by piece.
Frequently asked questions
What is the difference between a neural network and deep learning?
A neural network is the model. Deep learning is the practice of training neural networks with many layers. A network with only one or two layers is usually not called deep.
Do I need calculus to understand deep learning?
You need the idea of a slope and the chain rule. Both can be learned with small worked examples. You do not need a full calculus course before you start.
What is backpropagation in simple terms?
It is the method a network uses to find out how much each weight contributed to its error. It passes the error backwards from the output so that every weight can be adjusted.
Are Transformers a type of deep learning?
Yes. A Transformer is a deep neural network built around attention layers. It is trained with the same loop of loss, backpropagation and gradient descent as other networks.