Artificial Intelligence and Machine Learning · Kaplan Business School
A visual, example-led introduction — no prior mathematics assumed.
Use the ← and → arrow keys to move between slides. Press T for the table of contents.
Section 1
What Is a Neural Network?
1.1 The idea · 1.2 A dataset and the parts of a network · 1.3 How it learns
1.1
1.1 The Idea
Definition
A neural network is a computer model, loosely inspired by the brain, that learns patterns from examples rather than following rules written by a person.
It is built from many small units called neurons, connected together in layers. Each connection has an adjustable setting. Learning means tuning those settings until the network's answers are good.
Everyday analogy
Learning to judge a ripe avocado. At first you guess and get it wrong. Each time you cut one open, you adjust your sense of the right colour and firmness. After enough examples you judge well — without any rulebook. A neural network learns the same way, from examples.
Why it matters
This lets computers handle problems that are impossible to write as rules — recognising images, understanding text, forecasting demand — simply by learning from data.
Figure 1.1: A neural network is a web of simple units.
1.2
1.2 A Dataset and the Parts of a Network
To make this concrete, take a simple dataset: measurements of iris flowers, and the species each one belongs to. The network's job is to predict the species from the four measurements.
Petal L
Petal W
Sepal L
Sepal W
Species
1.4
0.2
5.1
3.5
Setosa
4.7
1.4
7.0
3.2
Versicolor
6.0
2.5
6.3
3.3
Virginica
Inputs and output
The four measurements are the inputs (also called features). The species is the output we want to predict.
Every arrow carries a weight. Every neuron adds a bias and applies an activation function. These are the parts we will meet in turn.
1.3
1.3 How It Learns
A fresh network starts with random settings, so its first guesses are poor. It improves through one simple loop, repeated many times:
Make a guess from the current settings.
Check the mistake against the known answer.
Adjust the settings a little to reduce the mistake.
Repeat thousands of times until the guesses are good.
Keep this loop in mind
Everything later in this workshop is a detail of one of these four steps. Return to this slide whenever a concept feels abstract.
Figure 1.3: The learning loop at the heart of every neural network.
1.Q
Knowledge Check — Section 1
Q1. How does a neural network differ from ordinary rule-based software?
A neural network is trained on examples and tunes its own settings, rather than executing rules a programmer wrote by hand.
Q2. In the iris dataset, the four measurements act as the network's:
The measurements are fed in as inputs; the species is the output the network predicts. Weights are the adjustable settings on the connections.
Section 2
The Neuron and Activation
2.1 A neuron as a decision-maker · 2.2 Inputs, weights, and bias (interactive) · 2.3 Activation functions · 2.4 The activation function explored (interactive) · 2.5 Practice calculation
2.1
2.1 A Neuron as a Decision-Maker
A single artificial neuron is a tiny decision-maker. It looks at a few pieces of information, decides how much each one matters, adds them up, and produces one number as its answer.
Everyday analogy
A loan officer deciding whether to approve an application. They weigh several factors: income matters a lot, a long credit history matters somewhat, a recent missed payment counts against. They combine these into an overall judgement. A neuron does exactly this, with numbers.
A neuron, in one sentence
It multiplies each input by an importance value (a weight), adds the results together with a baseline adjustment (a bias), and passes the total through a decision rule (an activation function) to produce its output.
The next slide lets you become that neuron: turn the dials yourself and watch the output change.
Figure 2.1: A neuron weighs its inputs, then decides.
2.2
2.2 Inputs, Weights, and Bias
Drag the sliders. The weights set how much each input matters (red = counts against). The bias nudges the result up or down. Watch the neuron respond.
Inputs
6
3
8
Weights (importance)
0.6
0.8
-0.5
Bias (baseline)
0.5
Weighted sum + bias: 6.5
Decision rule ReLU → 6.50
Decision rule Sigmoid → 0.998
Sigmoid output as a confidence dial (0 = no, 1 = yes):
2.3
2.3 Activation Functions
The weighted sum is just a number. The activation function is the decision rule that turns it into the neuron's output. Different rules suit different jobs.
Step
output 0 or 1
used in perceptrons
ReLU
max(0, x)
most common today
Sigmoid
squashes to 0–1
probabilities
tanh
squashes to −1–1
centred at zero
Function
Output range
Step
0 or 1
ReLU
0 to ∞
Sigmoid
0 to 1
tanh
−1 to 1
Why they matter
Without a non-linear rule, stacking neurons only ever produces a straight-line model. Activation functions are what let networks learn curved, non-linear patterns.
Worked example
A customer spends 500 → weighted sum 2.1 → ReLU(2.1) = 2.1 → likely to purchase. A customer spends 50 → weighted sum −0.3 → ReLU(−0.3) = 0 → unlikely to purchase.
2.4
2.4 The Activation Function Explored
Move the input slider and watch where the total lands on each rule.
Step
on / off (0 or 1)
ReLU
keep positive, else 0
Sigmoid
squash to 0–1
4.0
For an input total of 4.0:
Step output: 1
ReLU output: 4.00
Sigmoid output: 0.982
Notice
Step is all-or-nothing. ReLU passes positive signals through unchanged. Sigmoid gives a smooth confidence between 0 and 1. The choice shapes how the network behaves.
2.5
2.5 Practice: Compute a Neuron's Output
Given network
Formula you need
Weighted sum: z = (w₁ × x₁) + (w₂ × x₂) + b. Then pass z through the activation function.
Part A — ReLU, where ReLU(x) = max(0, x)
z = + 0.5 =
output = ReLU() =
Part B — Sigmoid, where Sigmoid(x) = 1 / (1 + e−x)
use the same z =
output = 1 / (1 + e−) =
2.Q
Knowledge Check — Section 2
Q1. In a neuron, what does a weight represent?
Each weight scales one input up or down. A large positive weight makes that input strongly influential; a negative weight makes it count against the result.
Q2. What is the role of the activation function?
Without a non-linear decision rule, stacking neurons still only produces straight-line relationships. The activation function is what gives networks their flexibility.
Q3. A neuron computes a total of −2. What does ReLU output?
ReLU keeps positive values and replaces anything negative with 0. Since −2 is negative, the output is 0.
Section 3
Network Architecture
3.1 The three kinds of layer · 3.2 What each layer does · 3.3 What happens when you add neurons and layers (interactive)
3.1
3.1 The Three Kinds of Layer
A neural network is organised into layers. Every network has an input layer and an output layer, with one or more hidden layers in between.
Layer
What it does
Input
Receives the raw features, one neuron per feature (e.g. the four iris measurements).
Hidden
Transforms and combines the inputs into more useful internal features. This is where the real work happens.
Output
Produces the final prediction, one neuron per possible answer (e.g. the three species).
Analogy
An assembly line. Raw parts enter (input), each station reshapes and combines them (hidden), and the finished product comes out the end (output).
3.2
3.2 What Each Layer Does
Fully connected, feedforward
Every neuron connects to every neuron in the next layer, and information flows one way, from input to output. This is called a multilayer perceptron.
Each hidden neuron detects one simple pattern in the inputs.
Later layers combine those into more complex patterns.
The output layer turns the final combination into a prediction.
3.3
3.3 Adding Neurons and Layers
More hidden neurons and more layers give a network more capacity — more freedom to shape its decision boundary. Here the blue points sit inside a ring of red points; no straight line can separate them. Increase the neurons and watch the boundary adapt.
6
With this many neurons
A smooth boundary that captures the ring — a good fit.
Neurons
Result
Too few
Boundary too simple — underfits
About right
Smooth boundary — good fit
Too many
Boundary contorts to each point — overfits
The trade-off
More capacity can capture more complex patterns, but it also risks overfitting and needs more data and computation. Use enough, not as much as possible.
3.Q
Knowledge Check — Section 3
Q1. What is the main job of the hidden layers?
Hidden layers do the real work: each neuron detects a simple pattern, and successive layers combine these into the complex features needed for the prediction.
Q2. A network with far too few hidden neurons for its task will most likely:
Too little capacity means the boundary stays too simple to fit the data, which is underfitting. You saw this at the low end of the slider.
Q3. What is a risk of adding far more neurons and layers than needed?
Excess capacity lets the network memorise noise (overfitting) and makes training slower and hungrier for data. Use enough capacity, not the maximum.
Section 4
How a Network Learns
4.1 Turning dials · 4.2 Make a guess · 4.3 Measure the mistake · 4.4 Find the way downhill · 4.5 Assign blame backward · 4.6 The learning curve · 4.7 Practice calculation
4.1
4.1 Learning Is Just Turning Dials
A fresh network starts with random weights, so its first guesses are poor. Learning means adjusting every weight and bias, a little at a time, to make the mistakes smaller. The four-step loop, now with proper names:
Plain step
Proper name
1. Make a guess
Forward pass
2. Measure the mistake
Loss (error)
3. Work out how to improve
Backpropagation
4. Take a small step
Gradient descent
The whole of training
Repeat these four steps thousands of times. Nothing more mysterious than that is happening inside a neural network as it learns.
Analogy: a sound engineer
A dozen dials on a mixing desk. The engineer plays a little, hears what is wrong, nudges a few dials, and listens again. Over many rounds the mix gets better. The network tunes its "dials" (weights) the same way, guided by its mistakes.
4.2
4.2 Step One: Make a Guess (Forward Pass)
The forward pass is the network doing what we practised in Section 2, layer by layer, until a prediction comes out the end. This is the "guess".
Forward pass
Feeding the inputs through the network, layer by layer, to produce a prediction. No learning happens yet — this is only the guess we will then judge.
Formula for one neuron
output = f ( (w₁ × x₁) + (w₂ × x₂) + ⋯ + b ), where f is the activation function. The forward pass applies this neuron by neuron, from left to right, until the output layer gives the prediction.
This example
Press "Run the guess" to send the inputs 2 and 3 through the network.
4.3
4.3 Step Two: Measure the Mistake
To improve, the network needs a number for how wrong it is. We compare the prediction to the correct answer and measure the gap. Drag the prediction and watch the penalty grow.
0.30
How far off: 0.70
Penalty (gap squared): 0.49
Why square the gap?
Squaring makes big mistakes hurt far more than small ones. A gap of 0.2 gives a penalty of 0.04; a gap of 0.8 gives 0.64 — sixteen times larger. This penalty is called the loss.
4.4
4.4 Step Three: Find the Way Downhill
Picture the loss as a valley. Each weight sits somewhere on the slope. The network feels which way is downhill (the gradient) and takes a step in that direction. The step size is the learning rate. Try different learning rates.
0.35
Current weight: 5.00
Current loss: 12.50
Steps taken: 0
Press "Run" and watch the ball roll toward the minimum.
Choosing the learning rate
Too small (0.05): tiny steps, very slow. Too large (1.8): the ball overshoots and bounces past the bottom. A well-chosen rate slides smoothly into the valley.
4.5
4.5 Step Four: Assign Blame Backward
The mistake appears at the output, but the weights that caused it are spread through the whole network. Backpropagation traces the error backward and works out how much each weight contributed, so each one can be nudged in the right direction.
Analogy: tracing a mistake
A team ships a product with a fault. The manager works backward through each stage to see who contributed to the problem and by how much, then adjusts each person's part a little. Backpropagation does this for weights.
The formulas
Output error: δout = (y − t) × f′(output). Hidden error: δh = δout × w × f′(zh). Gradient of a weight = (error at its neuron) × (its input). Then w ← w − η × gradient.
Weights being nudged
w₁: 0.50 w₂: 0.30 w₃: 0.80
Press the button to send the error backward.
4.6
4.6 Putting It Together: the Learning Curve
Running the four steps once barely helps. Running them over the whole dataset many times is what produces a good model. One full pass through the data is called an epoch.
Epoch: 0
Loss: 1.000
Reading the curve
The loss starts high and falls as the network learns. When the curve flattens, further training brings little improvement — the model has converged.
4.7
4.7 Practice: One Step of Backpropagation
Given
Hidden uses ReLU, so ReLU′(z) = 1 when z > 0 and 0 otherwise. Output is linear, so its slope is 1. Target t = 1; learning rate η = 0.1.
Forward pass
zh = w₁x₁ + w₂x₂ = h = ReLU(zh) =
y = w₃ × h =
Backward pass
δout = y − t = δh = δout × w₃ × ReLU′(zh) =
Updated weights — gradient of a weight = (error at its neuron) × (its input)
w₃ ← w₃ − η × (δout × h) =
w₁ ← w₁ − η × (δh × x₁) =
w₂ ← w₂ − η × (δh × x₂) =
4.Q
Knowledge Check — Section 4
Q1. What does the loss measure?
The loss is a single number capturing how wrong the network is. Training aims to make it as small as possible.
Q2. If the learning rate is set far too high, what tends to happen?
Large steps jump past the bottom of the valley. As you saw in the demo, the ball overshoots and the loss may increase rather than settle.
Q3. What is the job of backpropagation?
Backpropagation traces the output error backward through the network, assigning a share of the blame to each weight so gradient descent can update it.
Section 5
Making Training Work
5.1 Underfitting and overfitting · 5.2 Spotting overfitting (interactive) · 5.3 How to prevent it
5.1
5.1 The Central Challenge: Fitting Just Right
The real goal is not to do well on the training examples — it is to do well on new, unseen data. Two failures get in the way.
Underfitting
Too simple. Misses the real pattern. Poor on training and new data.
Good fit
Captures the trend without chasing every point. Works well on new data.
Overfitting
Passes through every point, memorising noise. Looks perfect in training, fails on new data.
The key idea
A model that memorises the training data has not learned the underlying pattern. Generalisation — performing well on data it has never seen — is what we are really after.
5.2
5.2 Spotting Overfitting as It Happens
Hold back some data the network never trains on — a validation set. Track the loss on both. While both fall, learning is healthy. When validation loss starts to rise, the network has begun memorising. Press Train.
Training loss: 1.00
Validation loss: 1.00
Both losses start high.
The tell-tale gap
Training loss keeps falling, but validation loss turns upward. The point where validation is lowest is the best model — the dashed line marks where training should stop.
5.3
5.3 How to Prevent Overfitting
Several practical techniques keep a network honest. You do not need the internals today — just what each one is for.
Technique
What it does, in plain terms
More training data
The more varied examples the network sees, the harder it is to memorise them all — it is pushed to learn the general pattern.
Early stopping
Stop training at the moment validation loss is lowest, before memorising begins. This is the dashed line from the previous slide.
Simpler network
Fewer neurons and layers means less capacity to memorise noise. Match the network's size to the difficulty of the problem.
Regularisation
Gently discourage the network from relying too heavily on any one weight, keeping the model smoother.
Dropout
Randomly ignore some neurons during each training step, so the network cannot lean on any single path and must spread what it learns.
Takeaway
The art of training is balancing capacity against generalisation: powerful enough to learn the pattern, disciplined enough not to memorise the noise.
5.Q
Knowledge Check — Section 5
Q1. A model scores almost perfectly on training data but poorly on new data. This is:
Excellent on training but weak on unseen data is the signature of overfitting: the model memorised rather than generalised.
Q2. Why do we hold back a validation set?
The validation set is an honest preview of performance on unseen data, letting us catch overfitting as it begins.
Q3. Which technique stops training when validation loss is lowest?
Early stopping halts training at the point of best validation performance, before the network starts to memorise noise.
Section 6
Autoencoders
6.1 The idea and the bottleneck (interactive) · 6.2 Image reconstruction · 6.3 Use cases
6.1
6.1 Autoencoders: Learning a Compressed Representation
Definition
An autoencoder is an unsupervised technique that uses a neural network for feature learning — it learns from the data alone, with no labels.
The network is built with a deliberate bottleneck in the middle — a layer with very few neurons.
The first half (the encoder) squeezes the input down into that small bottleneck: a compressed representation.
The second half (the decoder) tries to reconstruct the original input from that compressed form.
The trick
If a small middle layer can rebuild the input, it must have captured the input's most important features. That compact summary is the useful thing an autoencoder produces.
Figure 6.1: six inputs squeeze through three units, then rebuild six outputs.
6.2
6.2 A Simple Example: Image Reconstruction
A classic demonstration: train an autoencoder on images of handwritten digits. It learns to compress each image into a handful of numbers, then rebuild it. The rebuilt image is close to the original — proof the compressed code kept what mattered.
Analogy
Describing a photo to a friend over the phone in just a few words, then having them redraw it. If their drawing matches, your short description captured the essentials. The "code" is that short description.
Reference code
A worked Python example that reconstructs digit images is available at geeksforgeeks.org/auto-encoders. Your facilitator may demonstrate it in class.
6.3
6.3 What Autoencoders Are Used For
Use case
What the compressed representation gives us
Data compression
A smaller encoding that still captures the essentials of the input.
Dimension reduction
Fewer features for other models to work with — a learned alternative to PCA from Week 4.
Anomaly detection
Inputs that reconstruct poorly are flagged as unusual — useful for fraud and fault detection.
Image denoising
Rebuild a clean image from a noisy or partial one.
Generative tasks
Sample new, plausible data from the learned compressed space.
Link back to Week 4
Autoencoders and PCA both reduce dimensions. The difference: an autoencoder is a neural network, so it can capture curved, non-linear structure that PCA, being linear, cannot.
Where this fits
The autoencoder is your first example of an architecture designed for a purpose — the same neuron and the same training loop from earlier today, arranged into an encoder and decoder around a bottleneck.
6.Q
Knowledge Check — Section 6
Q1. What forces an autoencoder to learn a compressed representation?
The small bottleneck cannot hold everything, so the encoder must squeeze the input down to only its most important features.
Q2. An autoencoder is trained without labels. This makes it an example of:
It learns purely from the input data by trying to reconstruct it, with no target labels — that is unsupervised learning.
Q3. How can an autoencoder help detect anomalies?
An autoencoder rebuilds familiar patterns well. When an input is unlike anything in training, the reconstruction is poor, and that large error signals an anomaly.