This material has been reproduced and communicated to you by or on behalf of Kaplan Business School pursuant to Part VB of the Copyright Act 1968 (the Act).
The material in this communication may be subject to copyright under the Act. Any further reproduction or communication of this material by you may be the subject of copyright protection under the Act.
Do not remove this notice.
| Week | Topic |
|---|---|
| 1 | Types of variable. Visualisation for categorical variables. |
| 2 | Population and sample. Descriptive statistics. Visualisation for discrete data. |
| 3 | Visualisation for continuous data: histogram, cfp, boxplots. |
| 4 | Pictograms, scatter plots, correlation. |
| 5 | Group presentations. Linear trend line. |
| 6 | Visualisation for time series. |
| 7 | Visualisation for multivariate data. |
| 8 | Missing values: imputation and interpolation. |
| 9 | Aggregates and joining tables. |
| 10–11 | Data visualisation with ggplot. |
| 12 | Individual presentations. |
Six outcomes. Each one is assessable.
The word counts variables, not charts and not observations. Getting that definition right is what stops you from mistaking a wall of scatterplots for a multivariate analysis.
Every dataset in this unit is a rectangle. Rows are observations. Columns are variables. The word multivariate counts columns that you read together.
| Term | Columns read together | The question it answers | Typical chart |
|---|---|---|---|
| Univariate | 1 | How is this one variable distributed? Where is its centre and its spread? | Histogram, boxplot, bar chart |
| Bivariate | 2 | Do these two variables move together? In which direction, and how strongly? | Scatterplot, two-way table |
| Multivariate | 3 or more | How do these variables behave jointly? Does the relationship between two of them depend on the others? | Everything in this lesson |
The iris dataset: 150 flowers, four measurements, three species.
Below, the same two columns are plotted twice. The right-hand panel adds species.
| Group | n | r |
|---|---|---|
| All flowers pooled | 150 | −0.118 |
| setosa only | 50 | +0.743 |
| versicolor only | 50 | +0.526 |
| virginica only | 50 | +0.457 |
Pearson correlation of sepal length against sepal width, computed in R.
This is Simpson's paradox, which you met in Week 4. It is the single strongest argument for multivariate visualisation: a bivariate summary can be correct, reproducible, and point in exactly the wrong direction.
A sheet of paper and a screen both give you two position axes. Once a third variable arrives you have to spend something else to show it.
Position is read most accurately. Colour hue is read least accurately. The ranking comes from experiments on how precisely people judge each channel.
| Strategy | What it does | Charts in this lesson | Cost |
|---|---|---|---|
| Encode | Keep two position axes and spend colour, size or shape on the extra variables. | Bubble plot, coloured scatterplot, heat map, glyph plot, Chernoff faces | Every variable past the second is read less accurately. |
| Repeat | Split the page into panels, one panel per level of a variable, or one panel per pair. | Scatterplot matrix, small multiples | Panels shrink as p grows. Comparison across panels is harder than within one. |
| Project | Redraw every observation on a shared set of axes, or map three axes onto the plane. | Parallel coordinate plot, 3D scatterplot, radar chart | Overlap, occlusion, and dependence on the order or angle you happened to choose. |
One procedure, six steps, applied in the same order to every chart in this lesson. Use it in your presentation and use it in the exam.
Work through these in order. Steps 1 to 3 are about the chart. Steps 4 to 6 are about the data. Most misreadings happen because a reader jumps straight to step 4.
Identify what a single mark on the page stands for. One row of the data? One group? One time point?
Map every variable to the channel that carries it: horizontal position, vertical position, colour, size, shape, panel, angle.
Read the units and ranges. Check whether values were standardised, whether zero is on the axis, and whether panels share limits.
Describe the dominant pattern across all marks, before you look at any individual mark. Direction, strength, form.
Now split by the categorical channel. Do groups separate, overlap, or reverse? Name the marks that break the pattern.
Write one sentence you can defend, then one sentence naming what this chart cannot establish.
mtcars.# 32 cars, 4 variables, one plot
sz <- 1 + 3 * sqrt(mtcars$disp / max(mtcars$disp))
pal <- c("4" = "#0072B2", "6" = "#E69F00", "8" = "#D55E00")
plot(mtcars$wt, mtcars$mpg,
cex = sz, pch = 21, lwd = 1, col = "white",
bg = adjustcolor(pal[as.character(mtcars$cyl)], 0.7),
xlab = "Weight (1000 lbs)",
ylab = "Fuel economy (mpg)")
legend("topright", bty = "n", pch = 21, col = "white",
pt.bg = pal, legend = c("4 cyl", "6 cyl", "8 cyl"))
# radius <- sqrt(value) so that AREA is proportional
# to the value. Setting radius <- value would
# exaggerate large engines by a factor of disp.| The misreading | Why it happens | The fix |
|---|---|---|
| Treating hue as ordered | A rainbow palette assigned to a numeric variable invites the reader to rank the colours. Hue has no natural order, so the ranking they invent is arbitrary. | Use a sequential palette (one hue, varying lightness) for ordered variables. Reserve distinct hues for nominal categories. |
| Comparing panels with different axes | Software defaults often fit each panel to its own data. Two panels then look similar while describing very different ranges. | Set xlim and ylim explicitly and identically across panels.
State on the slide that the axes are shared. |
| Reading a value off an unlabelled scale | Streamgraph thickness, violin width and radar area all look quantitative. Two of them carry no units and the third is distorted by axis order. | Ask step 3 of the template before step 4. If the scale has no units, restrict yourself to claims about shape and ordering. |
Scatterplot matrix, heat map, 3D scatterplot, parallel coordinates, glyph plots and small multiples. Each slide gives the chart, the six-step read, and the base R that produces it.
# All 6 pairs of 4 variables, coloured by species
cols <- c("#0072B2", "#D55E00", "#009E73")[iris$Species]
pairs(iris[, 1:4],
col = cols,
pch = 19,
cex = 0.7,
main = "Iris: four measurements, six pairs")
# p variables produce choose(p, 2) distinct pairs:
choose(4, 2) # 6 -> comfortable
choose(7, 2) # 21 -> crowded
choose(10, 2) # 45 -> use a heat map insteadzlim. Without a fixed zlim the palette rescales to whatever is in the data, and a maximum of 0.3 would be drawn in the colour that 1.0 should own.v <- c("mpg","cyl","disp","hp","drat","wt","qsec")
M <- cor(mtcars[, v])
pal <- colorRampPalette(c("#2166AC","white","#B2182B"))(101)
image(1:7, 1:7, M[, 7:1], col = pal,
zlim = c(-1, 1), # fix the scale
axes = FALSE, xlab = "", ylab = "")
axis(1, 1:7, v, las = 2)
axis(2, 1:7, rev(v), las = 1)
text(rep(1:7, times = 7), rep(1:7, each = 7),
round(as.vector(M[, 7:1]), 2), cex = 0.7)
box()
# M[, 7:1] reverses the columns so that the first
# variable appears at the TOP of the vertical axis.# Base R has no scatter3d(). persp() draws the box
# and returns a projection matrix for trans3d().
zf <- matrix(min(mtcars$mpg), 2, 2)
pm <- persp(x = range(mtcars$wt),
y = range(mtcars$hp), z = zf,
zlim = range(mtcars$mpg),
theta = 40, phi = 22, d = 2,
xlab = "Weight", ylab = "Horsepower",
zlab = "mpg", ticktype = "detailed",
border = "grey80", col = NA)
pt <- trans3d(mtcars$wt, mtcars$hp, mtcars$mpg, pm)
fl <- trans3d(mtcars$wt, mtcars$hp,
rep(min(mtcars$mpg), nrow(mtcars)), pm)
segments(pt$x, pt$y, fl$x, fl$y,
col = "grey70", lty = 3) # drop lines
points(pt, pch = 21, bg = "#0072B2",
col = "white", cex = 1.3)
# theta and phi are the two rotation angles.
# Change them and re-read: if your finding moves,
# it was an artefact of the angle.X <- iris[, 1:4]
# min-max scale each column to [0, 1]
Xs <- apply(X, 2,
function(z) (z - min(z)) / (max(z) - min(z)))
cols <- c("#0072B2", "#D55E00", "#009E73")[iris$Species]
matplot(1:4, t(Xs), type = "l", lty = 1, lwd = 1,
col = adjustcolor(cols, 0.5),
axes = FALSE, xlab = "",
ylab = "Scaled value (0 to 1)")
axis(1, at = 1:4, labels = names(X))
axis(2)
abline(v = 1:4, col = "grey30")
# Reading rule for two ADJACENT axes:
# lines roughly parallel -> positive relationship
# lines crossing heavily -> negative relationshipUSArrests.# Base R ships stars(): the same idea, without
# the face metaphor. One glyph per row.
st <- c("Georgia","Hawaii","Mississippi","New Jersey",
"North Dakota","Texas","Vermont","California")
stars(USArrests[st, ],
draw.segments = TRUE,
key.loc = c(6.5, 1.6), len = 0.9,
main = "One glyph per state, four variables each")
# Chernoff faces themselves need a package:
# install.packages("aplpack")
# aplpack::faces(USArrests[st, ])
# The package is outside the base R used in this
# unit, so stars() is the version to submit.op <- par(mfrow = c(1, 3), mar = c(4, 4, 3, 1))
for (s in levels(iris$Species)) {
d <- subset(iris, Species == s)
plot(d$Sepal.Length, d$Sepal.Width,
pch = 19, col = "#0072B2",
xlim = range(iris$Sepal.Length), # SHARED
ylim = range(iris$Sepal.Width), # SHARED
xlab = "Sepal length",
ylab = "Sepal width", main = s)
abline(lm(Sepal.Width ~ Sepal.Length, data = d),
lwd = 2)
}
par(op) # always restore the graphics settings
# Dropping xlim and ylim lets R fit each panel to
# its own data. The panels then look alike while
# describing different ranges.zlim = c(-1, 1). What can go wrong?Violin, rose, radar and streamgraph. Each answers a narrow question well and invites a specific misreading. Learn both halves.
sp <- levels(iris$Species)
pal <- c("#0072B2", "#D55E00", "#009E73")
plot(NA, xlim = c(0.5, 3.5),
ylim = range(iris$Sepal.Length),
xaxt = "n", xlab = "",
ylab = "Sepal length (cm)")
axis(1, at = 1:3, labels = sp)
for (i in seq_along(sp)) {
y <- iris$Sepal.Length[iris$Species == sp[i]]
d <- density(y) # kernel density
w <- 0.38 * d$y / max(d$y) # scale to slot
polygon(c(i - w, rev(i + w)), # mirror it
c(d$x, rev(d$x)),
col = adjustcolor(pal[i], 0.3),
border = pal[i], lwd = 2)
boxplot(y, at = i, add = TRUE, boxwex = 0.10,
axes = FALSE, col = "black",
medcol = "white", outline = FALSE)
}
# Base R has no violin function. Building it from
# density() and polygon() is the exercise.set.seed(3100).set.seed(3100)
dirs <- c("N","NNE","NE","ENE","E","ESE","SE","SSE",
"S","SSW","SW","WSW","W","WNW","NW","NNW")
cnt <- c(38,25,18,15,22,41,66,84,
72,45,28,19,26,35,44,40)
cnt <- pmax(5, cnt + round(rnorm(16, 0, 4)))
rad <- sqrt(cnt / max(cnt)) # AREA proportional
plot(NA, xlim = c(-1.25, 1.25), ylim = c(-1.25, 1.25),
asp = 1, axes = FALSE, xlab = "", ylab = "")
for (i in seq_along(cnt)) {
a <- seq(-pi/16, pi/16, length.out = 24) +
(i - 1) * 2 * pi / 16
polygon(c(0, rad[i] * sin(a)),
c(0, rad[i] * cos(a)),
col = adjustcolor("#0072B2", 0.62),
border = "white")
}
ang <- (seq_along(dirs) - 1) * 2 * pi / 16
text(1.14 * sin(ang), 1.14 * cos(ang), dirs, cex = 0.8)
# sin() for x and cos() for y puts 0 at the top and
# runs clockwise, which is what a compass needs.ax <- c("Sales","Marketing","Development",
"Support","IT","Admin")
bud <- c(45, 15, 58, 18, 40, 12)
act <- c(52, 40, 57, 18, 33, 17)
vmax <- 60
ang <- (seq_along(ax) - 1) * 2 * pi / length(ax)
plot(NA, xlim = c(-1.4, 1.4), ylim = c(-1.4, 1.4),
asp = 1, axes = FALSE, xlab = "", ylab = "")
for (g in seq(10, 60, by = 10)) # grid rings
polygon((g/vmax) * sin(ang), (g/vmax) * cos(ang),
border = "grey85")
segments(0, 0, sin(ang), cos(ang), col = "grey85")
ring <- function(v, colr) {
r <- v / vmax
polygon(r * sin(ang), r * cos(ang), border = colr,
lwd = 2.5, col = adjustcolor(colr, 0.12))
points(r * sin(ang), r * cos(ang), pch = 21,
bg = colr, col = "white")
}
ring(bud, "#0072B2")
ring(act, "#CC0000")
text(1.22 * sin(ang), 1.22 * cos(ang), ax,
cex = 0.85, font = 2)Nothing has been added or removed. The only change is that both series now sit on a common vertical position scale with a shared baseline at zero.
set.seed(3100).set.seed(3100)
gen <- c("Rock","Electronic","Jazz","Hip-hop","Classical")
m <- 24
Y <- sapply(1:5, function(k)
pmax(5, round(c(120,90,40,70,30)[k] +
c(45,55,18,60,12)[k] *
sin(2*pi*(1:m)/12 + c(0,2,4,1,3)[k]) +
rnorm(m, 0, 8))))
colnames(Y) <- gen
tot <- rowSums(Y)
base <- -0.5 * tot # the "themeriver" baseline
plot(NA, xlim = c(1, m),
ylim = range(c(base, base + tot)),
axes = FALSE, xlab = "Month", ylab = "")
axis(1)
pal <- c("#0072B2","#D55E00","#009E73",
"#E69F00","#6A51A3")
for (k in 1:5) {
top <- base + Y[, k]
polygon(c(1:m, m:1), c(base, rev(top)),
col = pal[k], border = "white")
base <- top
}
# base <- rep(0, m) gives an ordinary stacked area
# chart, where the bottom band sits on a flat axis
# and can be read accurately.The question comes first, the chart second. Then you name the limitation before someone in the audience names it for you.
| The question you want to answer | Chart | Watch for |
|---|---|---|
| Which pairs among my numeric variables move together? | Scatterplot matrix (p ≤ 4), correlation heat map (p > 4) | Panel size; a fixed zlim |
| Does the relationship between two variables depend on a third? | Scatter with colour or size; small multiples | Shared axis limits across panels |
| Do my observations fall into groups across many variables at once? | Parallel coordinate plot | Axis order; overplotting |
| How does one observation compare with another on a fixed set of variables? | Glyph plot (stars), radar chart |
Arbitrary feature and axis assignment |
| How does the shape of a distribution differ between groups? | Violin plot, side-by-side boxplots | Bandwidth; how width was scaled |
| How is a directional or cyclic variable distributed? | Rose plot | Radius scaled by the square root |
| How do the parts of a total change over time? | Stacked area chart; streamgraph for shape only | No vertical axis on a streamgraph |
Parallel coordinates, heat maps of raw values and glyph plots all place variables measured in different units on a shared visual scale. Rescale first, or the variable with the largest raw numbers dominates the picture.
Maps every value into \([0, 1]\). Use it for parallel coordinates, where each axis needs a fixed top and bottom.
Centres each variable on zero and expresses it in standard deviations. Use it for heat maps of raw values, where a diverging palette needs a meaningful midpoint.
# z-scores: scale() does both steps
z <- scale(mtcars[, c("mpg", "hp", "wt")])
round(colMeans(z), 10)
#> mpg hp wt
#> 0 0 0
round(apply(z, 2, sd), 6)
#> mpg hp wt
#> 1 1 1
# min-max: write it out
mm <- function(v) (v - min(v)) / (max(v) - min(v))
Xs <- apply(iris[, 1:4], 2, mm)
round(range(Xs), 3)
#> 0 1
| Chart | The choice you are making | How to defend it |
|---|---|---|
| Parallel coordinates | Which variables sit next to each other. Only neighbours can be compared directly, so the order determines which relationships a reader can see. | Order by correlation, or place the variables your question is about side by side. State the rule you used. |
| Heat map | The order of rows and columns. Alphabetical order scatters a block structure; a sorted order reveals it. | Order by cluster or by the first principal component. Alphabetical order needs a reason. |
| Radar chart | The sequence of axes around the circle, which sets the polygon's shape and area. | Fix the order by convention and keep it identical across every radar chart in the report. |
| 3D scatterplot | The two rotation angles, theta and phi. |
Redraw at several angles. If your finding survives all of them, report it; if it does not, it was an artefact. |
| Streamgraph | The stacking order of the bands and the shape of the baseline. | Put the band you most want read at the bottom, or drop the wiggle and use a flat baseline. |
Multivariate charts fail first through density. 150 polylines already crowd a parallel coordinate plot; 5,000 turn it into a solid block.
adjustcolor(col, alpha.f = 0.2). Overlap then
reads as darkness, which restores density information.cex, or use pch = "."
once n runs into the thousands.jitter(x) for variables recorded to the nearest
whole number, where many points land on identical coordinates.smoothScatter(x, y) replaces the points with a
density surface.# transparency: the first thing to try
plot(iris$Sepal.Length, iris$Sepal.Width,
pch = 19, cex = 1.4,
col = adjustcolor("#0072B2", alpha.f = 0.25))
# jitter: for values rounded to 0.1 cm
plot(jitter(iris$Sepal.Length, amount = 0.03),
jitter(iris$Sepal.Width, amount = 0.03),
pch = 19, cex = 0.8)
# density surface: for large n
smoothScatter(rnorm(20000), rnorm(20000),
xlab = "x", ylab = "y")
Step 6 of the template produces two sentences. Here is the skeleton, and one completed example from this lesson.
Across [unit, n], [variable A]
[rises / falls] as [variable B]
[rises / falls], and this pattern
[holds / reverses] within [grouping variable].
This chart cannot establish [the limit].
Across 150 iris flowers, sepal width falls slightly as sepal length rises, and this pattern
reverses within every one of the three species.
This chart cannot establish why the species differ in level, because it contains no variable that
would explain the shift.
Both sentences are marked. A claim with no limit reads as overconfidence; a limit with no claim reads as no finding at all.
Unit → encoding → scale → structure → groups and exceptions → claim and limit. Six steps, in that order, on every chart.
| Scatterplot matrix | All pairs at once, p ≤ 4 |
| Correlation heat map | All pairs when p is large |
| Bubble plot | Two on position, two more encoded |
| 3D scatterplot | Three on position, if rotatable |
| Parallel coordinates | Whole rows, group structure |
| Glyph plot / faces | Row-by-row comparison |
| Small multiples | One relationship, once per group |
| Violin plot | Distribution shape by group |
| Rose plot | Directional or cyclic frequency |
| Radar chart | Shape recognition on fixed axes |
| Streamgraph | Shape and timing of parts of a total |
zlim, the axis order.airquality, which has six variables and missing values. Build a
correlation heat map and a scatterplot matrix of it. Note what R does with the missing
entries.The airquality exercise sets up next week directly. cor() returns
NA by default when any value in a pair is missing, and the argument that changes
that behaviour is the first decision in the imputation topic.
airquality heat map, and a one-line note of what cor(airquality)
returned before you did anything about it.Press T or Escape to close