Kaplan Business School · Lesson 11 of 12
| Week | Topic |
|---|---|
| 9 | Strategy optimisation: SEO, A/B tests, crisis handling |
| 10 | Customer lifetime value and churn: what they are, how to measure them |
| 11 | Predictive modelling: forecast sales, predict who leaves, decide what to do |
| 12 | Assessment |
Tomorrow will be 24 degrees.
BrewLab: how much will we sell next week?
There is a 60% chance of rain.
BrewLab: how likely is this member to leave?
Take an umbrella, or don't.
BrewLab: what do we do about it?
BrewLab is a specialty coffee shop in Melbourne. Dana runs a loyalty app: members earn a free drink after every ten purchases. Four thousand people have joined.
Dana is not a data person. She wants answers she can act on, in words she understands.
How much will BrewLab sell next week? Regression is a recipe that turns known facts into a sales estimate.
Look at two years of BrewLab weeks. Weeks with more marketing spend sold more. The dots do not sit on a perfect line, but a line through the middle captures the pattern.
That line is the model. Give it a spend, it gives back a sales estimate.
Start with base sales of about $1,800. Every extra $100 of marketing adds about $12.
| Number | What it means for Dana |
|---|---|
| 1,800 | Base sales when nothing special happens: no marketing, no promo, no rain, not December |
| +12 per $100 spend | Each extra $100 of marketing adds about $12 in sales |
| +650 if promo | A promotion week sells about $650 more than one without |
| −90 per rainy day | Each rainy day costs about $90 |
| +400 if December | A December week sells about $400 more |
Marketing spend $1,200 · promotion on · 2 rainy days forecast · not December
"Expect about $2,400 next week. The promotion is worth roughly $650 of that. If the rain clears, add another $180."
Typing 1,200 instead of 12 for spend. The forecast jumps to $16,000 for a coffee shop. If a number feels silly, it usually is.
Dana asks: "If I spend $9,000 in one week, the recipe says $3,530. Let's do it."
Her maths is right. Her trust is wrong. The model only ever saw weeks with spend between $300 and $2,000. It has no idea what happens at $9,000.
A coffee shop fills up. Extra advertising past a point brings nobody new through the door. The straight line keeps climbing anyway, because a straight line does not know about full shops.
Trust a forecast inside the data range. Outside it, the model is guessing, and so are you.
Answers "how much?" Sales, price, visits, spend.
Output: any number, like $2,414.
Part 1
Answers "will it happen?" Will they buy, will they leave, will they click.
Output: a chance between 0 and 1, like 0.82.
Part 2
Ridge, Lasso and ElasticNet are linear regression with a brake that stops the model over-reacting to noisy data. Polynomial regression lets the line bend. Quantile regression forecasts a "worst case" or "best case" instead of the middle. All are variations on the same recipe idea. You will meet them properly in DATA4400.
Which members are about to leave? Logistic regression gives every customer a number between 0 and 1, like a chance of rain.
Try the Part 1 recipe on a yes/no question. Leaving = 1, staying = 0. The straight line quickly predicts values like 1.4 or −0.3. There is no such thing as "140% leaving".
We want a curve that starts near 0, ends near 1, and never goes outside. That curve is logistic regression.
The app never says "140% chance of rain". It squeezes everything it knows into a number between 0% and 100%.
High score, high chance. Score 0 lands exactly on 0.50. The curve never leaves 0 to 1.
| Score | −3 | −2 | −1 | 0 | 1 | 2 | 3 |
|---|---|---|---|---|---|---|---|
| Chance | 0.05 | 0.12 | 0.27 | 0.50 | 0.73 | 0.88 | 0.95 |
Mia: 1 month, 2 orders, no complaint.
Tom: 3 months, 1 order, complaint.
Priya: same as Tom but no complaint. One complaint moves Priya from 0.62 to Tom's 0.82.
The model labels anyone above 0.50 "at risk". That is a shortcut. Priya at 0.62 and Tom at 0.82 both get the same label, but Tom is the more urgent call.
A chance of 0.62 means: out of 100 members like Priya, about 62 would leave and 38 would stay. Priya is not certain to leave.
Set it at 0.70 if she can only afford to contact the most likely leavers.
Set it at 0.30 if she wants to catch people early, before they drift.
Instead of adding points, a decision tree asks one question at a time and splits the customers. It ends in groups, each with its own leaving rate.
Trees are easy to read out loud. Logistic regression gives a smoother, more precise chance. In the lab today you will run both on the same data and compare.
Use the tree to explain the story to staff. Use logistic regression to rank who to contact first.
Data: WA_Fn-UseC_-Telco-Customer-Churn.csv (7,000 phone customers, with a column saying who left).
DATA4500-cluster-analysis-on-customer-churn.ipynb
Groups similar customers together so you can see which groups leave most. No target needed; the computer finds the groups.
DATA4500_Telco_Churn_prediction_and_explanation.ipynb
Builds both Part 2 models on the same data and shows which columns matter most.
While it runs, ask: which three columns would you expect to push the chance of leaving up? Check whether the model agrees.
The umbrella moment. A list of at-risk members is worth nothing until Dana does something with it.
The app checkout freezes. The coffee is inconsistent. The queue is too long at 8am.
Fixable
A complaint went unanswered. Staff were rude on a bad day. A refund took three weeks.
Fixable
They moved suburbs. They stopped drinking coffee. The product never suited them.
Let them go gracefully
The churn model tells Dana who. Only the reason tells her what to do. A voucher does not fix a frozen checkout.
A PwC survey asked people when they would stop dealing with a brand they love.
Every frozen checkout at BrewLab is not one annoyed customer. It is a coin flip on whether they ever come back.
Just after the first purchase, and when visits start to drop. Listen before they decide.
Spot the problem first. A short "sorry, here is a free drink" beats a long apology later.
"How likely are you to recommend us?" Promoters (9 to 10) minus detractors (0 to 6).
Onboarding for new members. Extra recognition for the loyal ones.
"Free coffee every fortnight", not "10-stamp digital card".
Promise what you can deliver, then deliver it. Over-promising creates leavers.
In pairs. Watch the short clip My Starbucks Rewards: Now on Android and iOS, then discuss and be ready to share.
Question 3 is the one that matters. Every "feature" in a loyalty app is also a data column.
Imagine every staff member keeping notes on customers in their own head. Nothing is shared, nothing is followed up. A CRM system is one shared notebook for the whole business.
A good CRM helps Dana:
Common tools: Salesforce, Microsoft Dynamics, Zoho, SAS Customer Intelligence.
The churn model, the voucher, the follow-up call and the result all live in one place. Next month, Dana can see whether the umbrella worked.
Source: Payne and Frow (2005), A Strategic Framework for Customer Relationship Management, Journal of Marketing 69(4).
Watch Salesforce: Philips is a Trailblazer, a short video about how the Dutch health-technology company uses CRM.
Because the pattern is the same at any size. Philips connects devices, patients and hospitals in one customer view. BrewLab connects an app, a till and a barista. Both are trying to notice a problem before the customer walks away.
Linear regression is a recipe. Each number is a plain sentence. Trust it inside the data range only.
Logistic regression adds up risk points and reads a chance off an S-curve. Use the chance to rank, not just the label to sort.
Find the reason people leave. Choose the action that fixes it. Measure against a comparison group.
Complete the three pen-and-paper exercises on MyKBS (Dana's sales recipe, the risk checklist, the umbrella decision). Week 12 is the final assessment.
Press T or Esc to close