DATA4800

Artificial Intelligence & Machine Learning

Workshop 6
Naïve Bayes  ·  Support Vector Machines  ·  Gradient Boosting
0.1

Where we are in the course

You have already met the first family of classifiers. Today we add three more powerful tools and learn when to reach for each.

Trees &Forests Clustering& PCA Assessment 1(in class) 6 You are here Naïve Bayes,SVM, Boosting NeuralNetworks DeepCNNs
Big picture
Naïve Bayes, SVM and Gradient Boosting are all supervised classification methods — they learn from labelled examples to predict a category (Yes/No, Malignant/Benign, Stay/Leave).
0.2

What you will be able to do

#By the end of this workshop you can…
1Apply Bayesian classification — use probabilities to predict a class.
2Investigate support vector machines — separate groups with the widest possible boundary.
3Understand gradient boosting — build a strong model from many small ones that learn from mistakes.
How to read these slides
Each tool follows the same path: a business problem → the idea in plain language → a worked example you can follow step by step → a hands-on activity in Orange / Python.
1
Section 1

Naïve Bayes Classifiers

Classifying things by asking: “Given what I observe, which outcome is most probable?”
1.1

1.1  Where this tool fits

Naïve Bayes is a form of supervised learning used for classification — predicting a label from labelled training data.

Machine Learning Supervised Unsupervised Reinforcement Regression Classification Naïve Bayes lives here Clustering Density est.
Figure 1.1 — Naïve Bayes, SVM and Gradient Boosting are all classification methods.
1.2

1.2  What is a Naïve Bayes classifier?

Business scenario
Your email provider sees a new message. It must decide: spam or not spam? It has no certainty — but it has seen thousands of past emails. Words like “free”, “winner”, “invoice” each shift the odds. Naïve Bayes combines those clues into one decision.

The core idea

Classify each item by its conditional probabilities — how likely each outcome is, given what we observe.

We write “the probability of A given B” as \(P(A \mid B)\).

Where it is used

  • Spam / email filtering
  • Sentiment analysis (positive vs negative reviews)
  • Classifying documents and articles by topic
Why “naïve”?
It makes a simplifying assumption: that the clues (features) are independent of one another. This is rarely perfectly true — but the method still works remarkably well, and it is fast.
1.3

1.3  Our running example: “Shall we play golf?”

We will learn the method through a small weather dataset. The goal: predict whether a person will play golf (Yes or No) from the day's weather.

OutlookTemp.HumidityWindyPlay?
0RainyHotHighFalseNo
1RainyHotHighTrueNo
2OvercastHotHighFalseYes
3SunnyMildHighFalseYes

(First four of 14 rows — full table appears in 1.10.)

The four clues (features)
  • Outlook: Sunny, Rainy, Overcast
  • Temperature: Hot, Mild, Cool
  • Humidity: High, Normal
  • Windy: True, False
The label to predict
Play Golf — Yes or No.
Dataset: geeksforgeeks.org/naive-bayes-classifiers
1.4

1.4  Building block: conditional probability

\(P(A \mid B)\) means: out of the times B happens, how often does A also happen?

Shop example
  • A = buys socks, \(P(A)=0.4\)
  • B = buys shoes, \(P(B)=0.5\)
  • Both, \(P(A \text{ and } B)=0.3\)

“Of the people who bought shoes, what fraction also bought socks?”

$$P(A \mid B)=\frac{P(A \text{ and } B)}{P(B)}=\frac{0.3}{0.5}=0.6$$

Socks (A) Shoes (B) A and B 0.3
The overlap is “both”. Dividing it by B's circle gives \(P(A\mid B)\).
1.5

1.5  Activity: calculate \(P(C\mid F)\)

Given
  • \(P(C)=0.25\) — you have a cough
  • \(P(F)=0.15\) — you have the flu
  • \(P(C \text{ and } F)=0.12\) — both
Try first, then reveal. Q1: state \(P(C\mid F)\) in words. Q2: compute it. Q3: can you find \(P(F\mid C)\)?
Step-by-step working
1In words: \(P(C\mid F)\) is “the probability you have a cough, given you have the flu.”
2Formula: $$P(C\mid F)=\frac{P(C \text{ and } F)}{P(F)}$$
3Substitute: \(=\dfrac{0.12}{0.15}\)
4Divide: \(=0.80\) → an 80% chance.
5Q3 — flip it: \(P(F\mid C)=\dfrac{P(C \text{ and } F)}{P(C)}=\dfrac{0.12}{0.25}=0.48\). “Given a cough, a 48% chance of flu.”
!Key insight: \(P(C\mid F)\neq P(F\mid C)\). The condition matters — this is exactly what Bayes' theorem lets us convert between.
1.6

1.6  Bayes' theorem — updating a belief with evidence

$$P(A \mid B)=\frac{P(B \mid A)\,P(A)}{P(B)}$$

TermName & meaning
\(P(A\mid B)\)Posterior — belief in A after seeing evidence B.
\(P(B\mid A)\)Likelihood — how well A explains the evidence.
\(P(A)\)Prior — belief in A before any evidence.
\(P(B)\)Evidence — how common B is overall.
Prior initial belief evidence B Posterior updated belief
Evidence turns a prior belief into a sharper posterior belief.
1.7

1.7  Setting up the golf problem

We restate the question in the language of classification.

Notation
One example day
\(X=(\text{Rainy, Hot, High, False})\), \(\;y=\text{No}\).
The question we can now ask: “What is the probability that someone will not play golf, given the weather is Rainy, Hot, High humidity, and no wind?”
1.8

1.8  The “naïve” independence assumption

To keep the maths simple, we assume the clues do not influence each other within a class. This lets us multiply their probabilities.

The assumption
$$P(x_1,x_2,\dots,x_n \mid y)=P(x_1\mid y)\times P(x_2\mid y)\times\cdots\times P(x_n\mid y)$$

Other assumptions of the model

  • Features are independent (the “naïve” part)
  • Discrete features follow simple category counts
  • All features treated as equally important
  • No missing data
In plain terms
Instead of studying every combination of weather together (which would need a huge dataset), we look at each clue on its own and combine them by multiplying. Simple, fast, surprisingly effective.
1.9

1.9  The decision rule

Because the bottom of Bayes' formula, \(P(B)\), is the same for every class, we can ignore it when comparing classes. That leaves a simple recipe:

$$P(y \mid x_1,\dots,x_n)\;\propto\;P(y)\prod_{i=1}^{n}P(x_i \mid y)$$

(the “\(\propto\)” symbol means “is proportional to”.)

In words — the whole method in one line
For each class (Yes and No): multiply the class's base rate by the probability of each clue in that class. Whichever class scores higher wins. $$\hat{y}=\arg\max_{y}\;P(y)\prod_{i=1}^{n}P(x_i \mid y)$$
1.10

1.10  The full training data (14 days)

OutlookTemp.HumidityWindyPlay
0RainyHotHighFalseNo
1RainyHotHighTrueNo
2OvercastHotHighFalseYes
3SunnyMildHighFalseYes
4SunnyCoolNormalFalseYes
5SunnyCoolNormalTrueNo
6OvercastCoolNormalTrueYes
7RainyMildHighFalseNo
8RainyCoolNormalFalseYes
9SunnyMildNormalFalseYes
10RainyMildNormalTrueYes
11OvercastMildHighTrueYes
12OvercastHotNormalFalseYes
13SunnyMildHighTrueNo
Our test day
We want to classify a new day: $$X=(\text{Sunny, Hot, Normal, False})$$
Count first
Of the 14 days: 9 are “Yes”, 5 are “No”. We will build all our probabilities by simply counting rows in this table.
Dataset: geeksforgeeks.org/naive-bayes-classifiers
1.11

1.11  Step 1 — class base rates (priors)

Before looking at today's weather, how often is golf played at all?

$$P(\text{Yes})=\frac{9}{14}\approx 0.64 \qquad P(\text{No})=\frac{5}{14}\approx 0.36$$

Frequency tables — count Yes/No for each clue value

OutlookYesNo
Sunny32
Overcast40
Rainy23
Total95
HumidityYesNo
Normal61
High34
Total95
WindyYesNo
False62
True33
Total95
Temp.YesNo
Hot22
Mild42
Cool31
These counts come straight from the 14 rows in 1.10 — you can verify each one yourself.
1.12

1.12  Step 2 — probabilities for our test day

Test day = (Sunny, Hot, Normal, False). Turn each count into a probability by dividing by the class total (9 for Yes, 5 for No).

FeatureValue\(P(\text{value}\mid \text{Yes})\)\(P(\text{value}\mid \text{No})\)
OutlookSunny3/9 ≈ 0.332/5 = 0.40
TemperatureHot2/9 ≈ 0.222/5 = 0.40
HumidityNormal6/9 ≈ 0.671/5 = 0.20
WindyFalse6/9 ≈ 0.672/5 = 0.40
Read the pattern
“Normal humidity” and “no wind” are much more common on Yes days (0.67 each) than on No days (0.20, 0.40). Those two clues will pull the answer toward Yes.
1.13

1.13  Step 3 — multiply out each class score

Multiply the base rate by all four clue probabilities, for each class.

Follow the arithmetic step by step
Y1Yes — set up: \(P(\text{Yes})\cdot P(\text{S}\mid Y)\cdot P(\text{H}\mid Y)\cdot P(\text{N}\mid Y)\cdot P(\text{F}\mid Y)\)
Y2Yes — substitute: \(=\dfrac{9}{14}\cdot\dfrac{3}{9}\cdot\dfrac{2}{9}\cdot\dfrac{6}{9}\cdot\dfrac{6}{9}\)
Y3Yes — multiply: \(\approx 0.643\times0.333\times0.222\times0.667\times0.667 \approx \mathbf{0.0212}\)
N1No — set up: \(P(\text{No})\cdot P(\text{S}\mid N)\cdot P(\text{H}\mid N)\cdot P(\text{N}\mid N)\cdot P(\text{F}\mid N)\)
N2No — substitute: \(=\dfrac{5}{14}\cdot\dfrac{2}{5}\cdot\dfrac{2}{5}\cdot\dfrac{1}{5}\cdot\dfrac{2}{5}\)
N3No — multiply: \(\approx 0.357\times0.40\times0.40\times0.20\times0.40 \approx \mathbf{0.0046}\)
Two raw scores:   Yes ≈ 0.0212   vs   No ≈ 0.0046.   Yes is already far larger — but let's turn them into clean percentages next.
1.14

1.14  Step 4 — normalise and decide

The two scores are proportional, not true probabilities. Divide each by their sum so they add to 100%.

$$P(\text{Yes}\mid\text{today})=\frac{0.0212}{0.0212+0.0046}\approx 0.82 \qquad P(\text{No}\mid\text{today})=\frac{0.0046}{0.0212+0.0046}\approx 0.18$$

Play — Yes
82%
Play — No
18%
Decision
Since \(0.82 > 0.18\), the model predicts: Yes — golf will be played. The clues “normal humidity” and “no wind” were decisive.
Note: every probability here was counted directly from the 14-row table, so your hand-calculation will match these figures exactly.
1.Q

Knowledge Check — Naïve Bayes

Q1. Why is the method called “naïve”?
The “naïve” assumption is feature independence — it lets us multiply individual clue probabilities together.
Q2. In our golf example, why could we ignore the denominator \(P(B)\) when comparing Yes vs No?
\(P(B)\) is a shared constant. Dropping it keeps the comparison valid; we only restore it (via normalising) to get clean percentages.
2
Section 2

Support Vector Machines

Drawing the widest, safest boundary between two groups.
2.1

2.1  What is a Support Vector Machine?

Business scenario
A clinic plots two patient groups — healthy and at-risk — using two measurements. They want a single dividing line so that future patients can be sorted automatically. But many lines could separate today's data. Which line is best?

The idea

An SVM finds the boundary (a hyperplane) that separates the classes with the largest possible gap on either side.

A wider gap means the boundary is more confident and generalises better to new data.

Good to know

  • Works for classification and regression.
  • In 2D the boundary is a line; in higher dimensions it is a “hyperplane”.
  • Milestone paper: Boser, Guyon & Vapnik (1992).
2.2

2.2  The hyperplane and the margin

Of all lines that separate the two groups, the SVM picks the one whose margin (the empty gap) is widest.

feature x₁ feature x₂ optimal hyperplane margin
Figure 2.2 — the solid line maximises the empty margin between the dashed edges.
2.3

2.3  Support vectors — the points that matter

Definition
The support vectors are the data points sitting closest to the boundary. They alone define where the boundary goes.
Why this is powerful
You could delete every other point and the boundary would not move. The SVM cares only about the hardest, borderline cases — not the easy ones far from the edge.
support vector support vector
Only the circled/outlined points define the boundary.
2.4

2.4  When a straight line won't work: the kernel trick

Sometimes no straight line can separate the groups. The SVM cleverly lifts the data into a higher dimension where a flat boundary does exist.

2D — no line can split these
kernel
3D — now a flat plane splits them decision surface
Figure 2.4 — the “kernel trick” projects data upward until a simple flat cut separates the classes.
2.5

2.5  A real use: normal vs cancer patients

Plot patients by two gene measurements. The SVM finds the line with the largest gap between the two patient groups.

Gene X Gene Y ← gap (margin) → Normal patients Cancer patients
Takeaway
The best boundary is the one furthest from the borderline patients on both sides — giving the most reliable predictions for new patients.
2.6

2.6  Activity — breast cancer prediction with SVM

Your role
You are an analyst at a cancer clinic. You are given a dataset with several measurements per tumour and a label: 0 = malignant, 1 = benign. Build an SVM to predict the label for new cases.

What to do

  1. Load the built-in breast-cancer dataset (scikit-learn / Orange).
  2. Split into training and test sets.
  3. Train an SVM classifier.
  4. Read off accuracy and the confusion matrix.
Resources
Notebook:
DATA4800_Week_6_Breast_Cancer_SVM_Gradient_Boosting.ipynb

In Orange: File → Data Sampler → SVM → Test & Score.
Source: datacamp.com/tutorial/svm-classification-scikit-learn-python
2.Q

Knowledge Check — SVM

Q1. What does an SVM try to maximise when choosing its boundary?
A wider margin gives a more confident boundary that generalises better to unseen data.
Q2. What is the “kernel trick” for?
When no straight line works, the kernel maps data upward where a simple separating surface exists.
3
Section 3

Ensemble Learning & Gradient Boosting

Many small models, working together, beat one big model.
3.1

3.1  What is ensemble learning?

Everyday analogy
Ask one expert and you get one opinion, possibly biased. Ask a diverse panel and combine their answers — the group is usually more accurate than any single member. Ensemble learning applies this “wisdom of the crowd” to models.
Definition
Ensemble learning combines many models to produce better predictions than any one alone. The two main strategies are bagging and boosting.
3.2

3.2  Bagging vs Boosting

Bagging — work in parallel

data model 1 model 2 model 3 vote

Train many models on different random slices, then average them. Reduces variance.
Example: Random Forest.

Boosting — work in sequence

data model 1 model 2 model 3 sum

Each new model fixes the mistakes of the one before, focusing on hard cases. Reduces bias.
Example: Gradient Boosting.

3.3

3.3  What is gradient boosting?

Definition
Gradient boosting combines many weak models (usually small decision trees) into one strong model. The trees are built one after another, each one trained to reduce the errors left by the previous ones.

Why “weak” models?

Each small tree on its own is barely better than guessing. But stacked in sequence, each correcting the last, they add up to a highly accurate model.

What it's good at

Capturing complex relationships in structured/tabular business data — churn, credit risk, demand — often achieving top accuracy.

3.4

3.4  How it works: learning from mistakes

Each round, the errors (red) that remain get more attention, and the next tree shrinks them further.

Round 1 many errors remain Round 2 fewer, smaller errors Round 3 almost all corrected strong model
In one sentence
Gradient boosting trains each new tree on the leftover errors of the trees so far — steadily correcting the model's hardest mistakes.
3.5

3.5  Gradient boosting — strengths & weaknesses

StrengthsWeaknesses
High accuracy on complex problems.Computationally intensive — slow to train on big data.
Handles missing data reasonably well.Can overfit if not carefully tuned, especially with noise.
Feature importance — shows which variables matter.Harder to interpret than a single tree.
Excels on tabular/business data.Needs tuning (learning rate, tree count, depth).
Robust to outliers vs a plain tree.Slower predictions — less ideal for real-time.
3.6

3.6  Gradient Boosting vs Random Forest

FeatureRandom Forest (bagging)Gradient Boosting
TrainingMany trees built independently and in parallel.Trees built sequentially, each fixing prior errors.
PerformanceStable across datasets.Often higher accuracy, more sensitive to noise.
OverfittingLess prone (averaging).More prone if not tuned.
SpeedFaster (parallel).Slower (sequential).
InterpretabilityEasier.Harder.
Best forLarge datasets, many features.Smaller, cleaner data needing high accuracy.
3.7

3.7  Activity — compare every model in Orange

Your role
You are an HR analyst predicting employee churn — who is at high risk of leaving. Dataset in Orange: “Employee attrition”.

The challenge

Build every classifier you have learned so far and compare their predictive performance side by side:

  • Logistic Regression
  • k-Nearest Neighbours (k-NN)
  • Decision Tree & Random Forest
  • Support Vector Machine
  • Gradient Boosting
In Orange
Wire: File → Test & Score, then attach each learner widget. Compare AUC, accuracy and F1 in one table, and inspect the confusion matrix for the best model.
Discussion
Which model wins on this data — and does the “best accuracy” model also give the clearest business explanation?
3.Q

Knowledge Check — Ensembles

Q1. The key difference between bagging and boosting is:
Random Forest (bagging) averages independent trees; Gradient Boosting (boosting) adds trees one at a time to correct mistakes.
Q2. Each new tree in gradient boosting is trained mainly on:
Boosting focuses attention on the remaining mistakes, shrinking them round after round.
4.1

4.1  Three tools, three jobs

MethodCore ideaReach for it when…
Naïve BayesCombine clue probabilities; pick the most likely class.Text/spam, fast baselines, many categorical features.
SVMWidest possible boundary between classes.Clear separation, medium-sized data, high-dimensional features.
Gradient BoostingMany small trees, each fixing the last one's errors.Structured business data where accuracy is the priority.
The habit to build
No single model wins everywhere. Try several, compare them fairly (as in the Orange activity), and weigh accuracy against interpretability for the business question at hand.
4.2

Coming next — Week 7

Introduction to Neural Networks

Questions? Bring your Orange workflows and the breast-cancer notebook to the lab.

Table of Contents

Press T or Escape to close