Naïve Bayes · Support Vector Machines · Gradient Boosting
0.1
Where we are in the course
You have already met the first family of classifiers. Today we add three more powerful tools and learn when to reach for each.
Big picture
Naïve Bayes, SVM and Gradient Boosting are all supervised classification methods — they learn from labelled examples to predict a category (Yes/No, Malignant/Benign, Stay/Leave).
0.2
What you will be able to do
#
By the end of this workshop you can…
1
Apply Bayesian classification — use probabilities to predict a class.
2
Investigate support vector machines — separate groups with the widest possible boundary.
3
Understand gradient boosting — build a strong model from many small ones that learn from mistakes.
How to read these slides
Each tool follows the same path: a business problem → the idea in plain language → a worked example you can follow step by step → a hands-on activity in Orange / Python.
1
Section 1
Naïve Bayes Classifiers
Classifying things by asking: “Given what I observe, which outcome is most probable?”
1.1
1.1 Where this tool fits
Naïve Bayes is a form of supervised learning used for classification — predicting a label from labelled training data.
Figure 1.1 — Naïve Bayes, SVM and Gradient Boosting are all classification methods.
1.2
1.2 What is a Naïve Bayes classifier?
Business scenario
Your email provider sees a new message. It must decide: spam or not spam? It has no certainty — but it has seen thousands of past emails. Words like “free”, “winner”, “invoice” each shift the odds. Naïve Bayes combines those clues into one decision.
The core idea
Classify each item by its conditional probabilities — how likely each outcome is, given what we observe.
We write “the probability of A given B” as \(P(A \mid B)\).
Where it is used
Spam / email filtering
Sentiment analysis (positive vs negative reviews)
Classifying documents and articles by topic
Why “naïve”?
It makes a simplifying assumption: that the clues (features) are independent of one another. This is rarely perfectly true — but the method still works remarkably well, and it is fast.
1.3
1.3 Our running example: “Shall we play golf?”
We will learn the method through a small weather dataset. The goal: predict whether a person will play golf (Yes or No) from the day's weather.
Outlook
Temp.
Humidity
Windy
Play?
0
Rainy
Hot
High
False
No
1
Rainy
Hot
High
True
No
2
Overcast
Hot
High
False
Yes
3
Sunny
Mild
High
False
Yes
(First four of 14 rows — full table appears in 1.10.)
\(P(A \mid B)\) means: out of the times B happens, how often does A also happen?
Shop example
A = buys socks, \(P(A)=0.4\)
B = buys shoes, \(P(B)=0.5\)
Both, \(P(A \text{ and } B)=0.3\)
“Of the people who bought shoes, what fraction also bought socks?”
$$P(A \mid B)=\frac{P(A \text{ and } B)}{P(B)}=\frac{0.3}{0.5}=0.6$$
The overlap is “both”. Dividing it by B's circle gives \(P(A\mid B)\).
1.5
1.5 Activity: calculate \(P(C\mid F)\)
Given
\(P(C)=0.25\) — you have a cough
\(P(F)=0.15\) — you have the flu
\(P(C \text{ and } F)=0.12\) — both
Try first, then reveal. Q1: state \(P(C\mid F)\) in words. Q2: compute it. Q3: can you find \(P(F\mid C)\)?
Step-by-step working
1In words: \(P(C\mid F)\) is “the probability you have a cough, given you have the flu.”
2Formula: $$P(C\mid F)=\frac{P(C \text{ and } F)}{P(F)}$$
3Substitute: \(=\dfrac{0.12}{0.15}\)
4Divide: \(=0.80\) → an 80% chance.
5Q3 — flip it: \(P(F\mid C)=\dfrac{P(C \text{ and } F)}{P(C)}=\dfrac{0.12}{0.25}=0.48\). “Given a cough, a 48% chance of flu.”
!Key insight: \(P(C\mid F)\neq P(F\mid C)\). The condition matters — this is exactly what Bayes' theorem lets us convert between.
1.6
1.6 Bayes' theorem — updating a belief with evidence
$$P(A \mid B)=\frac{P(B \mid A)\,P(A)}{P(B)}$$
Term
Name & meaning
\(P(A\mid B)\)
Posterior — belief in A after seeing evidence B.
\(P(B\mid A)\)
Likelihood — how well A explains the evidence.
\(P(A)\)
Prior — belief in A before any evidence.
\(P(B)\)
Evidence — how common B is overall.
Evidence turns a prior belief into a sharper posterior belief.
1.7
1.7 Setting up the golf problem
We restate the question in the language of classification.
Notation
\(y\) = the class label we want to predict — “Yes” or “No”.
\(X=(x_1,x_2,\dots,x_n)\) = the feature vector — the day's clues: (Outlook, Temperature, Humidity, Wind).
One example day
\(X=(\text{Rainy, Hot, High, False})\), \(\;y=\text{No}\). The question we can now ask: “What is the probability that someone will not play golf, given the weather is Rainy, Hot, High humidity, and no wind?”
1.8
1.8 The “naïve” independence assumption
To keep the maths simple, we assume the clues do not influence each other within a class. This lets us multiply their probabilities.
Instead of studying every combination of weather together (which would need a huge dataset), we look at each clue on its own and combine them by multiplying. Simple, fast, surprisingly effective.
1.9
1.9 The decision rule
Because the bottom of Bayes' formula, \(P(B)\), is the same for every class, we can ignore it when comparing classes. That leaves a simple recipe:
(the “\(\propto\)” symbol means “is proportional to”.)
In words — the whole method in one line
For each class (Yes and No): multiply the class's base rate by the probability of each clue in that class. Whichever class scores higher wins.
$$\hat{y}=\arg\max_{y}\;P(y)\prod_{i=1}^{n}P(x_i \mid y)$$
1.10
1.10 The full training data (14 days)
Outlook
Temp.
Humidity
Windy
Play
0
Rainy
Hot
High
False
No
1
Rainy
Hot
High
True
No
2
Overcast
Hot
High
False
Yes
3
Sunny
Mild
High
False
Yes
4
Sunny
Cool
Normal
False
Yes
5
Sunny
Cool
Normal
True
No
6
Overcast
Cool
Normal
True
Yes
7
Rainy
Mild
High
False
No
8
Rainy
Cool
Normal
False
Yes
9
Sunny
Mild
Normal
False
Yes
10
Rainy
Mild
Normal
True
Yes
11
Overcast
Mild
High
True
Yes
12
Overcast
Hot
Normal
False
Yes
13
Sunny
Mild
High
True
No
Our test day
We want to classify a new day:
$$X=(\text{Sunny, Hot, Normal, False})$$
Count first
Of the 14 days: 9 are “Yes”, 5 are “No”. We will build all our probabilities by simply counting rows in this table.
Frequency tables — count Yes/No for each clue value
Outlook
Yes
No
Sunny
3
2
Overcast
4
0
Rainy
2
3
Total
9
5
Humidity
Yes
No
Normal
6
1
High
3
4
Total
9
5
Windy
Yes
No
False
6
2
True
3
3
Total
9
5
Temp.
Yes
No
Hot
2
2
Mild
4
2
Cool
3
1
These counts come straight from the 14 rows in 1.10 — you can verify each one yourself.
1.12
1.12 Step 2 — probabilities for our test day
Test day = (Sunny, Hot, Normal, False). Turn each count into a probability by dividing by the class total (9 for Yes, 5 for No).
Feature
Value
\(P(\text{value}\mid \text{Yes})\)
\(P(\text{value}\mid \text{No})\)
Outlook
Sunny
3/9 ≈ 0.33
2/5 = 0.40
Temperature
Hot
2/9 ≈ 0.22
2/5 = 0.40
Humidity
Normal
6/9 ≈ 0.67
1/5 = 0.20
Windy
False
6/9 ≈ 0.67
2/5 = 0.40
Read the pattern
“Normal humidity” and “no wind” are much more common on Yes days (0.67 each) than on No days (0.20, 0.40). Those two clues will pull the answer toward Yes.
1.13
1.13 Step 3 — multiply out each class score
Multiply the base rate by all four clue probabilities, for each class.
Since \(0.82 > 0.18\), the model predicts: Yes — golf will be played. The clues “normal humidity” and “no wind” were decisive.
Note: every probability here was counted directly from the 14-row table, so your hand-calculation will match these figures exactly.
1.Q
Knowledge Check — Naïve Bayes
Q1. Why is the method called “naïve”?
The “naïve” assumption is feature independence — it lets us multiply individual clue probabilities together.
Q2. In our golf example, why could we ignore the denominator \(P(B)\) when comparing Yes vs No?
\(P(B)\) is a shared constant. Dropping it keeps the comparison valid; we only restore it (via normalising) to get clean percentages.
2
Section 2
Support Vector Machines
Drawing the widest, safest boundary between two groups.
2.1
2.1 What is a Support Vector Machine?
Business scenario
A clinic plots two patient groups — healthy and at-risk — using two measurements. They want a single dividing line so that future patients can be sorted automatically. But many lines could separate today's data. Which line is best?
The idea
An SVM finds the boundary (a hyperplane) that separates the classes with the largest possible gap on either side.
A wider gap means the boundary is more confident and generalises better to new data.
Good to know
Works for classification and regression.
In 2D the boundary is a line; in higher dimensions it is a “hyperplane”.
Milestone paper: Boser, Guyon & Vapnik (1992).
2.2
2.2 The hyperplane and the margin
Of all lines that separate the two groups, the SVM picks the one whose margin (the empty gap) is widest.
Figure 2.2 — the solid line maximises the empty margin between the dashed edges.
2.3
2.3 Support vectors — the points that matter
Definition
The support vectors are the data points sitting closest to the boundary. They alone define where the boundary goes.
Why this is powerful
You could delete every other point and the boundary would not move. The SVM cares only about the hardest, borderline cases — not the easy ones far from the edge.
Only the circled/outlined points define the boundary.
2.4
2.4 When a straight line won't work: the kernel trick
Sometimes no straight line can separate the groups. The SVM cleverly lifts the data into a higher dimension where a flat boundary does exist.
Figure 2.4 — the “kernel trick” projects data upward until a simple flat cut separates the classes.
2.5
2.5 A real use: normal vs cancer patients
Plot patients by two gene measurements. The SVM finds the line with the largest gap between the two patient groups.
Takeaway
The best boundary is the one furthest from the borderline patients on both sides — giving the most reliable predictions for new patients.
2.6
2.6 Activity — breast cancer prediction with SVM
Your role
You are an analyst at a cancer clinic. You are given a dataset with several measurements per tumour and a label: 0 = malignant, 1 = benign. Build an SVM to predict the label for new cases.
What to do
Load the built-in breast-cancer dataset (scikit-learn / Orange).
Q1. What does an SVM try to maximise when choosing its boundary?
A wider margin gives a more confident boundary that generalises better to unseen data.
Q2. What is the “kernel trick” for?
When no straight line works, the kernel maps data upward where a simple separating surface exists.
3
Section 3
Ensemble Learning & Gradient Boosting
Many small models, working together, beat one big model.
3.1
3.1 What is ensemble learning?
Everyday analogy
Ask one expert and you get one opinion, possibly biased. Ask a diverse panel and combine their answers — the group is usually more accurate than any single member. Ensemble learning applies this “wisdom of the crowd” to models.
Definition
Ensemble learning combines many models to produce better predictions than any one alone. The two main strategies are bagging and boosting.
3.2
3.2 Bagging vs Boosting
Bagging — work in parallel
Train many models on different random slices, then average them. Reduces variance. Example: Random Forest.
Boosting — work in sequence
Each new model fixes the mistakes of the one before, focusing on hard cases. Reduces bias. Example: Gradient Boosting.
3.3
3.3 What is gradient boosting?
Definition
Gradient boosting combines many weak models (usually small decision trees) into one strong model. The trees are built one after another, each one trained to reduce the errors left by the previous ones.
Why “weak” models?
Each small tree on its own is barely better than guessing. But stacked in sequence, each correcting the last, they add up to a highly accurate model.
What it's good at
Capturing complex relationships in structured/tabular business data — churn, credit risk, demand — often achieving top accuracy.
3.4
3.4 How it works: learning from mistakes
Each round, the errors (red) that remain get more attention, and the next tree shrinks them further.
In one sentence
Gradient boosting trains each new tree on the leftover errors of the trees so far — steadily correcting the model's hardest mistakes.
3.5
3.5 Gradient boosting — strengths & weaknesses
Strengths
Weaknesses
High accuracy on complex problems.
Computationally intensive — slow to train on big data.
Handles missing data reasonably well.
Can overfit if not carefully tuned, especially with noise.
Feature importance — shows which variables matter.
Harder to interpret than a single tree.
Excels on tabular/business data.
Needs tuning (learning rate, tree count, depth).
Robust to outliers vs a plain tree.
Slower predictions — less ideal for real-time.
3.6
3.6 Gradient Boosting vs Random Forest
Feature
Random Forest (bagging)
Gradient Boosting
Training
Many trees built independently and in parallel.
Trees built sequentially, each fixing prior errors.
Performance
Stable across datasets.
Often higher accuracy, more sensitive to noise.
Overfitting
Less prone (averaging).
More prone if not tuned.
Speed
Faster (parallel).
Slower (sequential).
Interpretability
Easier.
Harder.
Best for
Large datasets, many features.
Smaller, cleaner data needing high accuracy.
3.7
3.7 Activity — compare every model in Orange
Your role
You are an HR analyst predicting employee churn — who is at high risk of leaving. Dataset in Orange: “Employee attrition”.
The challenge
Build every classifier you have learned so far and compare their predictive performance side by side:
Logistic Regression
k-Nearest Neighbours (k-NN)
Decision Tree & Random Forest
Support Vector Machine
Gradient Boosting
In Orange
Wire: File → Test & Score, then attach each learner widget. Compare AUC, accuracy and F1 in one table, and inspect the confusion matrix for the best model.
Discussion
Which model wins on this data — and does the “best accuracy” model also give the clearest business explanation?
3.Q
Knowledge Check — Ensembles
Q1. The key difference between bagging and boosting is:
Random Forest (bagging) averages independent trees; Gradient Boosting (boosting) adds trees one at a time to correct mistakes.
Q2. Each new tree in gradient boosting is trained mainly on:
Boosting focuses attention on the remaining mistakes, shrinking them round after round.
4.1
4.1 Three tools, three jobs
Method
Core idea
Reach for it when…
Naïve Bayes
Combine clue probabilities; pick the most likely class.
Text/spam, fast baselines, many categorical features.
Many small trees, each fixing the last one's errors.
Structured business data where accuracy is the priority.
The habit to build
No single model wins everywhere. Try several, compare them fairly (as in the Orange activity), and weigh accuracy against interpretability for the business question at hand.
4.2
Coming next — Week 7
Introduction to Neural Networks
Questions? Bring your Orange workflows and the breast-cancer notebook to the lab.