TECH3100 · Data Visualisation in R

Plots for multivariate data

and other specialised visualisations
Lesson 7
Kaplan Business School Australia
All figures generated in base R 4.3.3 from built-in datasets
0.1

0.1 Copyright notice

Commonwealth of Australia · Copyright Regulations 1969

This material has been reproduced and communicated to you by or on behalf of Kaplan Business School pursuant to Part VB of the Copyright Act 1968 (the Act).

The material in this communication may be subject to copyright under the Act. Any further reproduction or communication of this material by you may be the subject of copyright protection under the Act.

Do not remove this notice.

0.2

0.2 Where this week sits

WeekTopic
1Types of variable. Visualisation for categorical variables.
2Population and sample. Descriptive statistics. Visualisation for discrete data.
3Visualisation for continuous data: histogram, cfp, boxplots.
4Pictograms, scatter plots, correlation.
5Group presentations. Linear trend line.
6Visualisation for time series.
7Visualisation for multivariate data.
8Missing values: imputation and interpolation.
9Aggregates and joining tables.
10–11Data visualisation with ggplot.
12Individual presentations.
0.3

0.3 What you should be able to do by the end

Six outcomes. Each one is assessable.

  1. State what makes an analysis multivariate rather than univariate or bivariate.
  2. Explain why a relationship between two variables can change sign once a third is added.
  3. Name the visual channels available to you and rank them by how accurately they are read.
  4. Apply a fixed six-step procedure to read any multivariate chart.
  5. Produce a scatterplot matrix, heat map, bubble plot, parallel coordinate plot, glyph plot, violin plot, rose plot, radar chart and streamgraph in base R.
  6. Choose the right plot for a stated question and state one limitation of your choice.
Assessment link Outcome 4 and outcome 6 are what the Week 12 individual presentation marks hardest. A correct chart with no defensible reading of it scores in the lower bands.
1
Section 1

What multivariate means

The word counts variables, not charts and not observations. Getting that definition right is what stops you from mistaking a wall of scatterplots for a multivariate analysis.

1.1

1.1 Start from the data matrix

Every dataset in this unit is a rectangle. Rows are observations. Columns are variables. The word multivariate counts columns that you read together.

Univariateone column, read togetherHeightWeightAgeSteps172683181401655924112501818445532015852389670distribution ofone variableBivariatetwo columns, read togetherHeightWeightAgeSteps172683181401655924112501818445532015852389670relationship betweentwo variablesMultivariatethree or more columns, read togetherHeightWeightAgeSteps172683181401655924112501818445532015852389670joint behaviour ofall four at onceOne row = one observation. One column = one variable. The count of columns you read together is what changes.
1.2

1.2 One definition, three cases

TermColumns read together The question it answersTypical chart
Univariate1 How is this one variable distributed? Where is its centre and its spread? Histogram, boxplot, bar chart
Bivariate2 Do these two variables move together? In which direction, and how strongly? Scatterplot, two-way table
Multivariate3 or more How do these variables behave jointly? Does the relationship between two of them depend on the others? Everything in this lesson
A common error Drawing ten separate scatterplots is ten bivariate analyses, not one multivariate analysis. The analysis becomes multivariate only when the variables are present in the same view at the same time, so that the reader can see how they interact.
Notation Write \(n\) for the number of observations and \(p\) for the number of variables. A dataset is an \(n \times p\) matrix. Univariate means \(p = 1\); bivariate means \(p = 2\); multivariate means \(p \geq 3\). Choosing a chart is mostly a question of how large \(p\) is.
1.3

1.3 Why the third column changes the answer

The iris dataset: 150 flowers, four measurements, three species. Below, the same two columns are plotted twice. The right-hand panel adds species.

Bivariate viewr = −0.118 across all 150 flowers5678234Sepal length (cm)Sepal width (cm)Multivariate viewr = +0.74 / +0.53 / +0.46 within species5678234Sepal length (cm)Sepal width (cm)Same two columns. Adding the third column reverses the sign of the relationship.
1.4

1.4 What just happened

The numbers

Groupnr
All flowers pooled150−0.118
setosa only50+0.743
versicolor only50+0.526
virginica only50+0.457

Pearson correlation of sepal length against sepal width, computed in R.

The mechanism Species shifts each group to a different place on the plane. Virginica has long sepals and middling width; setosa has short sepals and wide ones. Pooling the groups traces a line through the group centres, and that line runs the opposite way to the line inside each group.

The name

This is Simpson's paradox, which you met in Week 4. It is the single strongest argument for multivariate visualisation: a bivariate summary can be correct, reproducible, and point in exactly the wrong direction.

What follows from it

  • A correlation coefficient computed on pooled data carries an unstated assumption: that no grouping variable shifts the groups apart.
  • You cannot detect the problem from the two columns alone. You have to look at a third.
  • The reversal is not rare and it is not a curiosity. It appears in medical trials, salary audits, admissions data and school results.
Practical rule Before you report a correlation, ask which categorical variable could split your sample. Then plot it and look.
1.5

1.5 The page has only two dimensions

A sheet of paper and a screen both give you two position axes. Once a third variable arrives you have to spend something else to show it.

p = 1one axis is enoughvariable Ap = 2two axes, one planevariable Avariable Bp = 3the page has no third axisvariable Avariable Bvariable C → size and colour
Consequence Every multivariate chart in this lesson is a scheme for spending a non-position channel, splitting the page into panels, or squashing several dimensions into two. Those are the only three options.
1.6

1.6 The channels you can spend, ranked

Position is read most accurately. Colour hue is read least accurately. The ranking comes from experiments on how precisely people judge each channel.

Positionmost accurateLengthAngle / slopeAreaColour valueColour hueleast accurateJudged accuratelyJudged roughlyPut the variable you most need read precisely on position. Push the least critical variable to hue.Every multivariate chart is a decision about which variable loses precision.
1.7

1.7 Three strategies for p ≥ 3

StrategyWhat it does Charts in this lessonCost
Encode Keep two position axes and spend colour, size or shape on the extra variables. Bubble plot, coloured scatterplot, heat map, glyph plot, Chernoff faces Every variable past the second is read less accurately.
Repeat Split the page into panels, one panel per level of a variable, or one panel per pair. Scatterplot matrix, small multiples Panels shrink as p grows. Comparison across panels is harder than within one.
Project Redraw every observation on a shared set of axes, or map three axes onto the plane. Parallel coordinate plot, 3D scatterplot, radar chart Overlap, occlusion, and dependence on the order or angle you happened to choose.
Sizing rule of thumb For \(p \leq 4\) a scatterplot matrix shows all \(\binom{p}{2}\) pairs at a readable size (6 pairs). For \(5 \leq p \leq 10\) prefer a correlation heat map or parallel coordinates (21 pairs at \(p=7\), 45 at \(p=10\)). Beyond that, reduce the variables before you plot them.
1.Q

1.Q Knowledge check: Section 1

Q1. A student plots mpg against weight, then plots mpg against horsepower in a second figure. Is this a multivariate analysis?
Multivariate means the variables are present in one view, so their joint behaviour is visible. Two separate bivariate plots cannot show whether the weight effect changes at different horsepower levels.
Q2. You have one variable that a client will read exact values from, and two that provide context. Which channel should carry the first variable?
Position sits at the top of the accuracy ranking. Hue is nominal and carries no natural order, so a reader cannot extract a number from it.
Q3. Sepal length and sepal width correlate at −0.118 across all 150 iris flowers, and positively within each of the three species. What does that tell you?
This is Simpson's paradox. All four correlations are correct. The pooled figure describes a mixture of three populations, and the mixture has a structure that no single group has.
2
Section 2

How to read a multivariate chart

One procedure, six steps, applied in the same order to every chart in this lesson. Use it in your presentation and use it in the exam.

2.1

2.1 The six-step read

Work through these in order. Steps 1 to 3 are about the chart. Steps 4 to 6 are about the data. Most misreadings happen because a reader jumps straight to step 4.

STEP 01

Unit

Identify what a single mark on the page stands for. One row of the data? One group? One time point?

What is one dot, line, cell or face?
STEP 02

Encoding

Map every variable to the channel that carries it: horizontal position, vertical position, colour, size, shape, panel, angle.

Which variable sits on which channel?
STEP 03

Scale

Read the units and ranges. Check whether values were standardised, whether zero is on the axis, and whether panels share limits.

Raw or rescaled? Shared or separate?
STEP 04

Structure

Describe the dominant pattern across all marks, before you look at any individual mark. Direction, strength, form.

What does the bulk of the data do?
STEP 05

Groups & exceptions

Now split by the categorical channel. Do groups separate, overlap, or reverse? Name the marks that break the pattern.

Does the pattern hold in every group?
STEP 06

Claim & limit

Write one sentence you can defend, then one sentence naming what this chart cannot establish.

What can I say, and what can I not say?
Anchor Read the legend before you read the picture. A multivariate chart cannot be read faster than its encoding key can be read, and the key is where the extra variables are hiding.
2.2

2.2 The template applied: two variables on position, a third by colour and size

2345102030Toyota CorollaLincoln ContinentalWeight (1000 lbs)Fuel economy (mpg)Colour: cylinders4 cyl6 cyl8 cylArea: displacement (cu. in.)80250460
Figure 2.2. mtcars: 32 car models. Weight and fuel economy on position, engine displacement on area, cylinder count on colour. Four variables in one view.
  1. UnitOne circle is one 1974 car model. There are 32 circles, one per row of mtcars.
  2. EncodingHorizontal position: weight. Vertical position: fuel economy. Circle area: engine displacement. Fill colour: cylinder count, used as a nominal grouping with three levels (4, 6, 8).
  3. ScaleWeight runs 1.51 to 5.42 thousand lbs; fuel economy 10.4 to 33.9 mpg. Neither axis starts at zero, so read differences and not ratios. Radius is set to the square root of displacement, so area is proportional to the value.
  4. StructureFuel economy falls as weight rises. The cloud runs from top-left to bottom-right with little scatter around that path; the correlation is −0.868.
  5. Groups & exceptionsColour and size move with position: heavy cars are also 8-cylinder and large-displacement. Toyota Corolla sits at the light and efficient extreme (1.84, 33.9); Lincoln Continental at the heavy extreme (5.42, 10.4).
  6. Claim & limitAmong these 32 models, heavier cars return lower fuel economy, and weight, displacement and cylinder count rise together. The chart cannot separate the effect of weight from the effect of displacement, because the two correlate at 0.888.
# 32 cars, 4 variables, one plot
sz  <- 1 + 3 * sqrt(mtcars$disp / max(mtcars$disp))
pal <- c("4" = "#0072B2", "6" = "#E69F00", "8" = "#D55E00")

plot(mtcars$wt, mtcars$mpg,
     cex = sz, pch = 21, lwd = 1, col = "white",
     bg  = adjustcolor(pal[as.character(mtcars$cyl)], 0.7),
     xlab = "Weight (1000 lbs)",
     ylab = "Fuel economy (mpg)")

legend("topright", bty = "n", pch = 21, col = "white",
       pt.bg = pal, legend = c("4 cyl", "6 cyl", "8 cyl"))

# radius <- sqrt(value) so that AREA is proportional
# to the value. Setting radius <- value would
# exaggerate large engines by a factor of disp.
2.3

2.3 Three ways students misread these charts

The misreadingWhy it happens The fix
Treating hue as ordered A rainbow palette assigned to a numeric variable invites the reader to rank the colours. Hue has no natural order, so the ranking they invent is arbitrary. Use a sequential palette (one hue, varying lightness) for ordered variables. Reserve distinct hues for nominal categories.
Comparing panels with different axes Software defaults often fit each panel to its own data. Two panels then look similar while describing very different ranges. Set xlim and ylim explicitly and identically across panels. State on the slide that the axes are shared.
Reading a value off an unlabelled scale Streamgraph thickness, violin width and radar area all look quantitative. Two of them carry no units and the third is distorted by axis order. Ask step 3 of the template before step 4. If the scale has no units, restrict yourself to claims about shape and ordering.
In your presentation State steps 1, 2 and 3 out loud before you state your finding. The marker is checking whether you know what your own chart encodes.
3
Section 3

The core multivariate plots

Scatterplot matrix, heat map, 3D scatterplot, parallel coordinates, glyph plots and small multiples. Each slide gives the chart, the six-step read, and the base R that produces it.

3.1

3.1 Scatterplot matrix

SepalLength4.3 – 7.9SepalWidth2 – 4.4PetalLength1 – 6.9PetalWidth0.1 – 2.5Sepal.LengthSepal.WidthPetal.LengthPetal.WidthSepal.LengthSepal.WidthPetal.LengthPetal.Widthsetosaversicolorvirginica16 panels, 6 distinct pairs, each shown twice (above and below the diagonal).
Figure 3.1. iris: 150 flowers, four measurements. Every pair of variables appears twice, once above and once below the diagonal.
  1. UnitOne point in any off-diagonal panel is one flower. The same 150 flowers appear in all 12 off-diagonal panels.
  2. EncodingA panel's row gives the vertical variable and its column gives the horizontal variable. Colour gives species. The diagonal names the variable and its range.
  3. ScaleEvery variable keeps its own centimetre range. Ranges differ between panels, so compare the shape of a cloud across panels rather than distances.
  4. StructureThe petal length against petal width panel shows a tight positive band (r = 0.963). The sepal length against sepal width panel shows a loose cloud with no clear direction.
  5. Groups & exceptionsSetosa separates completely from the other two species in every panel that involves a petal measurement. Versicolor and virginica overlap in all panels.
  6. Claim & limitPetal measurements separate the three iris species; sepal measurements do not. The matrix shows pairs only, so any pattern that requires three variables at once remains invisible here.
# All 6 pairs of 4 variables, coloured by species
cols <- c("#0072B2", "#D55E00", "#009E73")[iris$Species]

pairs(iris[, 1:4],
      col  = cols,
      pch  = 19,
      cex  = 0.7,
      main = "Iris: four measurements, six pairs")

# p variables produce choose(p, 2) distinct pairs:
choose(4,  2)   #  6  -> comfortable
choose(7,  2)   # 21  -> crowded
choose(10, 2)   # 45  -> use a heat map instead
3.2

3.2 Correlation heat map

1.00-.85-.85-.78.68-.87.42mpgmpg-.851.00.90.83-.70.78-.59cylcyl-.85.901.00.79-.71.89-.43dispdisp-.78.83.791.00-.45.66-.71hphp.68-.70-.71-.451.00-.71.09dratdrat-.87.78.89.66-.711.00-.17wtwt.42-.59-.43-.71.09-.171.00qsecqsec−10+1Pearson correlation32 cars, 7 numeric variables, 21 distinct pairs on one screen.
Figure 3.2. mtcars: 7 numeric variables, 21 distinct pairs. Colour and printed value carry the same number.
  1. UnitOne cell is one pair of variables. Seven variables give 49 cells, of which 21 are distinct pairs and 7 are the trivial diagonal.
  2. EncodingRow and column identify the pair. Fill colour gives the Pearson correlation. The printed number gives the same value exactly, which removes the need to estimate from the colour.
  3. ScaleA diverging palette centred on zero and fixed to the range −1 to +1 by zlim. Without a fixed zlim the palette rescales to whatever is in the data, and a maximum of 0.3 would be drawn in the colour that 1.0 should own.
  4. StructureA block of strong positive correlations joins cyl, disp, hp and wt. That same block runs strongly negative against mpg.
  5. Groups & exceptionsqsec sits apart from the block. Its strongest association is with hp (−0.708) and it is close to unrelated to drat (0.091).
  6. Claim & limitEngine size, weight, power and cylinder count form one cluster that runs against fuel economy. Pearson correlation measures linear association only, so a strong curved relationship would appear here as a small number.
v   <- c("mpg","cyl","disp","hp","drat","wt","qsec")
M   <- cor(mtcars[, v])
pal <- colorRampPalette(c("#2166AC","white","#B2182B"))(101)

image(1:7, 1:7, M[, 7:1], col = pal,
      zlim = c(-1, 1),          # fix the scale
      axes = FALSE, xlab = "", ylab = "")

axis(1, 1:7, v,       las = 2)
axis(2, 1:7, rev(v),  las = 1)
text(rep(1:7, times = 7), rep(1:7, each = 7),
     round(as.vector(M[, 7:1]), 2), cex = 0.7)
box()

# M[, 7:1] reverses the columns so that the first
# variable appears at the TOP of the vertical axis.
3.3

3.3 3D scatterplot

Weight (1000 lbs)low → highHorsepowerlow → highFuel economy (mpg)Drop lines tie every point to the floor.Without them, the depth of a point on the page is a guess.
Figure 3.3. mtcars: weight, horsepower and fuel economy, all three on position axes. Drop lines mark each point's floor position.
  1. UnitOne sphere is one car. The grey dot below it is the same car's position on the floor of the box.
  2. EncodingThree position axes and nothing else: weight, horsepower, fuel economy. No colour, size or shape is spent.
  3. ScaleEach axis spans the observed range. The projection is fixed by a viewing angle that you choose, so depth and distance are confounded on the page.
  4. StructureThe cloud slopes down: cars that are both heavy and powerful sit near the bottom of the fuel-economy axis.
  5. Groups & exceptionsTwo isolated points sit at the far horsepower extreme and low on the mpg axis. Without the drop lines you could not tell whether they are far back or simply low.
  6. Claim & limitFuel economy declines with both weight and horsepower. The claim depends on the angle: rotate the box and the same points can appear to have no pattern. Prefer a scatterplot matrix or a bubble plot unless the reader can rotate the figure themselves.
# Base R has no scatter3d(). persp() draws the box
# and returns a projection matrix for trans3d().
zf <- matrix(min(mtcars$mpg), 2, 2)

pm <- persp(x = range(mtcars$wt),
            y = range(mtcars$hp), z = zf,
            zlim = range(mtcars$mpg),
            theta = 40, phi = 22, d = 2,
            xlab = "Weight", ylab = "Horsepower",
            zlab = "mpg", ticktype = "detailed",
            border = "grey80", col = NA)

pt <- trans3d(mtcars$wt, mtcars$hp, mtcars$mpg, pm)
fl <- trans3d(mtcars$wt, mtcars$hp,
              rep(min(mtcars$mpg), nrow(mtcars)), pm)

segments(pt$x, pt$y, fl$x, fl$y,
         col = "grey70", lty = 3)      # drop lines
points(pt, pch = 21, bg = "#0072B2",
       col = "white", cex = 1.3)

# theta and phi are the two rotation angles.
# Change them and re-read: if your finding moves,
# it was an artefact of the angle.
3.4

3.4 Parallel coordinate plot

SepalLengthmax 7.9min 4.3SepalWidthmax 4.4min 2PetalLengthmax 6.9min 1PetalWidthmax 2.5min 0.1setosaversicolorvirginicaEvery variable is rescaled to its own 0–1 range before drawing.Axis order is a choice. Neighbouring axes are the only pairs you can compare directly.
Figure 3.4. iris: each of the 150 flowers is one polyline crossing four axes. Every axis is rescaled to its own 0 to 1 range.
  1. UnitOne polyline is one flower. It crosses all four axes, so the whole row of the data matrix is visible at once.
  2. EncodingHorizontal position selects the variable. Vertical position on that axis gives the flower's value for it. Colour gives species.
  3. ScaleEach axis is rescaled independently to 0 to 1 by min-max scaling. Halfway up the petal-width axis is not the same number of centimetres as halfway up the sepal-length axis.
  4. StructureBetween the two petal axes the lines run roughly parallel, which signals a positive relationship. Between sepal length and sepal width the lines cross heavily, which signals a negative or weak one.
  5. Groups & exceptionsSetosa forms a separate bundle: high on sepal width, low on both petal axes, and it stays separate across all four axes. The other two species interleave.
  6. Claim & limitSetosa is separable from the other two species on petal measurements alone. Only neighbouring axes can be compared directly, so any conclusion about sepal length against petal width requires the axes to be reordered and the plot redrawn.
Reading rule Between two adjacent axes, parallel lines mean a positive relationship and crossing lines mean a negative one. This only works for neighbours, which is why axis order is an analytical decision.
X  <- iris[, 1:4]
# min-max scale each column to [0, 1]
Xs <- apply(X, 2,
            function(z) (z - min(z)) / (max(z) - min(z)))

cols <- c("#0072B2", "#D55E00", "#009E73")[iris$Species]

matplot(1:4, t(Xs), type = "l", lty = 1, lwd = 1,
        col  = adjustcolor(cols, 0.5),
        axes = FALSE, xlab = "",
        ylab = "Scaled value (0 to 1)")

axis(1, at = 1:4, labels = names(X))
axis(2)
abline(v = 1:4, col = "grey30")

# Reading rule for two ADJACENT axes:
#   lines roughly parallel -> positive relationship
#   lines crossing heavily -> negative relationship
3.5

3.5 Glyph plots and Chernoff faces

Georgiamurder 17.4Hawaiimurder 5.3Mississippimurder 16.1New Jerseymurder 7.4North Dakotamurder 0.8Texasmurder 12.7Vermontmurder 2.2Californiamurder 9Face width = murder rate · Eye size = assault rateMouth curve = urban population % · Face height and nose = rape rate
Figure 3.5. USArrests: 8 of 50 states. Four crime and population variables mapped to four facial features.
  1. UnitOne face is one state. Eight faces are shown from the 50 rows of USArrests.
  2. EncodingFace width: murder rate. Eye size: assault rate. Mouth curvature: percentage of the population living in urban areas. Face height and nose length: rape rate.
  3. ScaleEach variable is min-max scaled across all 50 states before it is mapped to a feature. The assignment of variable to feature is chosen by the analyst and carries no meaning of its own.
  4. StructureWide faces with large eyes mark high violent-crime rates. Georgia and Mississippi read as wide; North Dakota and Vermont read as narrow.
  5. Groups & exceptionsGeorgia's face is the widest because its murder rate of 17.4 is the maximum in the dataset. California is tall with large eyes and a broad smile: high assault, high rape rate, 91% urban.
  6. Claim & limitThe southern states shown carry higher murder and assault rates than the northern states shown. The reading depends on which variable was assigned to which feature. Swap murder and urban population and the same data produces a different visual story. Faces recruit face recognition, which is fast but not calibrated.
# Base R ships stars(): the same idea, without
# the face metaphor. One glyph per row.
st <- c("Georgia","Hawaii","Mississippi","New Jersey",
        "North Dakota","Texas","Vermont","California")

stars(USArrests[st, ],
      draw.segments = TRUE,
      key.loc = c(6.5, 1.6), len = 0.9,
      main = "One glyph per state, four variables each")

# Chernoff faces themselves need a package:
#   install.packages("aplpack")
#   aplpack::faces(USArrests[st, ])
# The package is outside the base R used in this
# unit, so stars() is the version to submit.
3.6

3.6 Small multiples

setosar = +0.7435678234Sepal widthversicolorr = +0.5265678virginicar = +0.4575678Sepal length (cm)Identical axes on all three panels. Only that makes the panels comparable.Species is encoded by which panel a point sits in.
Figure 3.6. iris: the same two axes repeated once per species, on shared limits. This is the chart that resolves slide 1.3.
  1. UnitOne point is one flower. One panel is one species.
  2. EncodingHorizontal position: sepal length. Vertical position: sepal width. Panel membership: species. The third variable is carried by which panel a point sits in.
  3. ScaleAll three panels use identical x and y limits. That shared scale is the only thing that makes the panels comparable, and it has to be set explicitly.
  4. StructureEvery panel shows a positive slope. The fitted lines have slopes of the same sign in all three.
  5. Groups & exceptionsThe clouds sit at different heights. At any given sepal length, a setosa flower has a wider sepal than a virginica flower.
  6. Claim & limitWithin each species, wider sepals accompany longer ones, and the three species sit at different levels. The panels are separate, so a reader cannot judge the pooled relationship from this view. That is precisely why the pooled r of −0.118 was surprising.
op <- par(mfrow = c(1, 3), mar = c(4, 4, 3, 1))

for (s in levels(iris$Species)) {
  d <- subset(iris, Species == s)
  plot(d$Sepal.Length, d$Sepal.Width,
       pch  = 19, col = "#0072B2",
       xlim = range(iris$Sepal.Length),   # SHARED
       ylim = range(iris$Sepal.Width),    # SHARED
       xlab = "Sepal length",
       ylab = "Sepal width", main = s)
  abline(lm(Sepal.Width ~ Sepal.Length, data = d),
         lwd = 2)
}

par(op)   # always restore the graphics settings

# Dropping xlim and ylim lets R fit each panel to
# its own data. The panels then look alike while
# describing different ranges.
3.Q

3.Q Knowledge check: Section 3

Q1. In a parallel coordinate plot, two neighbouring axes show many crossing lines. What does that suggest about those two variables?
A line that is high on the left axis and low on the right must cross a line that does the opposite. Many crossings between neighbours therefore signal that one variable falls as the other rises.
Q2. A correlation heat map is drawn without setting zlim = c(-1, 1). What can go wrong?
The numbers are still correct; the colours are misleading. Fixing zlim keeps the meaning of each colour constant across every heat map you produce.
Q3. Why does a scatterplot matrix of 10 variables become hard to use?
choose(10, 2) is 45. The matrix scales with the square of p, so panel size falls away quickly. At that size, switch to a correlation heat map.
4
Section 4

Other specialist plots

Violin, rose, radar and streamgraph. Each answers a narrow question well and invites a specific misreading. Learn both halves.

4.1

4.1 Violin plot

45678setosan = 50 median 5versicolorn = 50 median 5.9virginican = 50 median 6.5Sepal length (cm)The black box is the boxplot. The coloured outline is the kernel density, mirrored.Width has no units. Only its shape carries meaning.
Figure 4.1. iris sepal length by species. The coloured outline is a mirrored kernel density; the black box inside is a standard boxplot.
  1. UnitThe outline summarises 50 flowers per species. Unlike a scatterplot, no single mark corresponds to a single observation.
  2. EncodingVertical position: sepal length in centimetres. Horizontal half-width: the estimated density at that value. The mirroring is decorative, since both halves carry the same information.
  3. ScaleThe vertical axis is in centimetres. The horizontal width has no units. Check whether widths are scaled per group or to one common maximum, because the two choices tell different stories about relative sample size.
  4. StructureAll three distributions are single-peaked and roughly symmetric. The whole distribution shifts upward from setosa to versicolor to virginica.
  5. Groups & exceptionsSetosa is the narrowest in spread and the lowest in level (median 5.0, quartiles 4.8 and 5.2). Virginica has a median of 6.5 and a wider body.
  6. Claim & limitSepal length increases across the three species, with substantial overlap between versicolor and virginica. The outline is a smoothed estimate; the bandwidth is a choice, and a different bandwidth changes how many bumps appear.
sp   <- levels(iris$Species)
pal  <- c("#0072B2", "#D55E00", "#009E73")

plot(NA, xlim = c(0.5, 3.5),
     ylim = range(iris$Sepal.Length),
     xaxt = "n", xlab = "",
     ylab = "Sepal length (cm)")
axis(1, at = 1:3, labels = sp)

for (i in seq_along(sp)) {
  y <- iris$Sepal.Length[iris$Species == sp[i]]
  d <- density(y)                  # kernel density
  w <- 0.38 * d$y / max(d$y)       # scale to slot

  polygon(c(i - w, rev(i + w)),    # mirror it
          c(d$x,   rev(d$x)),
          col = adjustcolor(pal[i], 0.3),
          border = pal[i], lwd = 2)

  boxplot(y, at = i, add = TRUE, boxwex = 0.10,
          axes = FALSE, col = "black",
          medcol = "white", outline = FALSE)
}

# Base R has no violin function. Building it from
# density() and polygon() is the exercise.
4.2

4.2 Rose plot

N42E22SE70SSE89S64W21NW45Area, not radiusRadius is drawnproportional to√count, so thewedge area isproportional tothe count itself.Scaling radiusdirectly inflateslarge sectors.Synthetic wind-direction record, 16 compass sectors, set.seed(3100).
Figure 4.2. synthetic wind-direction record, 16 compass sectors, generated with set.seed(3100).
  1. UnitOne wedge is one compass sector. The whole rose is one station over one recording period.
  2. EncodingAngle gives direction and is fixed at 22.5° per sector, so angle carries no quantitative information. Radius carries the count.
  3. ScaleRadius is drawn proportional to the square root of the count, so that wedge area is proportional to the count. Grid rings mark 25, 50, 75 and 100 per cent of the largest sector. There is no common baseline.
  4. StructureThe rose is elongated toward the south and south-southeast and pulled in on the western side.
  5. Groups & exceptionsSSE is the modal sector with 89 observations. W is the least frequent with 21. The ratio between them is a little over four to one.
  6. Claim & limitWind at this station came predominantly from the south-southeast. The plot records the frequency of direction only. It says nothing about speed, so a sector with rare but severe wind would appear small.
set.seed(3100)
dirs <- c("N","NNE","NE","ENE","E","ESE","SE","SSE",
          "S","SSW","SW","WSW","W","WNW","NW","NNW")
cnt  <- c(38,25,18,15,22,41,66,84,
          72,45,28,19,26,35,44,40)
cnt  <- pmax(5, cnt + round(rnorm(16, 0, 4)))

rad  <- sqrt(cnt / max(cnt))   # AREA proportional

plot(NA, xlim = c(-1.25, 1.25), ylim = c(-1.25, 1.25),
     asp = 1, axes = FALSE, xlab = "", ylab = "")

for (i in seq_along(cnt)) {
  a <- seq(-pi/16, pi/16, length.out = 24) +
       (i - 1) * 2 * pi / 16
  polygon(c(0, rad[i] * sin(a)),
          c(0, rad[i] * cos(a)),
          col = adjustcolor("#0072B2", 0.62),
          border = "white")
}

ang <- (seq_along(dirs) - 1) * 2 * pi / 16
text(1.14 * sin(ang), 1.14 * cos(ang), dirs, cex = 0.8)

# sin() for x and cos() for y puts 0 at the top and
# runs clockwise, which is what a compass needs.
4.3

4.3 Radar (spider) chart

$10$20$30$40$50$60SalesMarketingDevelopmentCustomerSupportInformationTechnologyAdministrationAllocated budgetActual spendingFigures in $000s
Figure 4.3. departmental budget against actual spending across six departments. Illustrative figures in thousands of dollars.
  1. UnitOne polygon is one series. Two polygons are shown: allocated budget and actual spending. Each has six vertices, one per department.
  2. EncodingAngle selects the department. Department is nominal, so the order around the circle was chosen by whoever drew the chart. Radius from the centre carries dollars.
  3. ScaleOne common radial scale runs from $0 at the centre to $60 at the outer ring, shared by all six axes. Check this on any radar chart you meet: many give each axis its own range, which makes the polygon meaningless.
  4. StructureBoth polygons stretch toward Development and Sales and pull in toward Administration and Customer Support.
  5. Groups & exceptionsMarketing is where the two polygons separate most: $40 spent against $15 allocated. Information Technology is the one department that came in under its allocation.
  6. Claim & limitMarketing overspent its allocation by the largest margin; IT underspent. The enclosed area has no meaning. Reordering the six departments changes the area of both polygons without changing a single number.
ax   <- c("Sales","Marketing","Development",
            "Support","IT","Admin")
bud  <- c(45, 15, 58, 18, 40, 12)
act  <- c(52, 40, 57, 18, 33, 17)
vmax <- 60
ang  <- (seq_along(ax) - 1) * 2 * pi / length(ax)

plot(NA, xlim = c(-1.4, 1.4), ylim = c(-1.4, 1.4),
     asp = 1, axes = FALSE, xlab = "", ylab = "")

for (g in seq(10, 60, by = 10))          # grid rings
  polygon((g/vmax) * sin(ang), (g/vmax) * cos(ang),
          border = "grey85")
segments(0, 0, sin(ang), cos(ang), col = "grey85")

ring <- function(v, colr) {
  r <- v / vmax
  polygon(r * sin(ang), r * cos(ang), border = colr,
          lwd = 2.5, col = adjustcolor(colr, 0.12))
  points(r * sin(ang), r * cos(ang), pch = 21,
         bg = colr, col = "white")
}
ring(bud, "#0072B2")
ring(act, "#CC0000")

text(1.22 * sin(ang), 1.22 * cos(ang), ax,
     cex = 0.85, font = 2)
4.4

4.4 What the radar chart hides

$0$20$40$60+7Sales+25Marketing-1Development+0CustomerSupport-7InformationTechnology+5AdministrationSame six numbers. The +25 gap at Marketing is legible here and buried in the polygon.Bars keep every value on one common vertical position scale.

The same six pairs of numbers

Nothing has been added or removed. The only change is that both series now sit on a common vertical position scale with a shared baseline at zero.

What became visible

  • The Marketing gap of +25 is the largest movement in the data. On the radar it is one vertex pushed outward among six, competing with the shape of the whole polygon.
  • Customer Support is unchanged at +0. On the radar the two polygons touch there, which reads as a coincidence rather than a fact.
  • Development is the largest line item in both series. The radar shows this, because it is the longest spoke.
When a radar chart is the right choice When the axes are genuinely commensurable, few in number, and fixed by convention, and when the reader's task is to recognise a familiar shape rather than to extract a value. Player profiles and skills audits fit that description. Budget variance does not.
4.5

4.5 Streamgraph

RockElectronicJazzHip-hopClassicalM1M7M13M19M24MonthThere is no y axis. Band thickness is the value; vertical position carries nothing.Total height at any month is the sum across all five genres.Synthetic listening counts, 5 genres, 24 months, set.seed(3100).
Figure 4.5. synthetic monthly listening counts across five genres, 24 months, generated with set.seed(3100).
  1. UnitOne coloured band is one genre followed across 24 months. A vertical slice through the figure is one month.
  2. EncodingHorizontal position: month. Band thickness: that genre's count in that month. Colour: genre. Vertical position carries nothing at all.
  3. ScaleThere is no vertical axis. The baseline is set to minus one half of the monthly total so the shape is centred on the page, which is an aesthetic decision with no analytical content.
  4. StructureTotal thickness rises and falls on a cycle of roughly twelve months. The whole shape narrows around months 6 to 8 and again around months 18 to 20.
  5. Groups & exceptionsRock and Electronic run out of phase: Rock is thickest when Electronic is thinnest. Classical stays thin throughout and never competes for the reader's attention.
  6. Claim & limitTotal listening is seasonal, and Rock and Electronic peak at opposite points in the year. A thickness read against a curved baseline is judged far less accurately than the same value read as a line against a fixed axis. Do not use a streamgraph to support a claim about a specific value.
set.seed(3100)
gen <- c("Rock","Electronic","Jazz","Hip-hop","Classical")
m   <- 24

Y <- sapply(1:5, function(k)
  pmax(5, round(c(120,90,40,70,30)[k] +
                c(45,55,18,60,12)[k] *
                sin(2*pi*(1:m)/12 + c(0,2,4,1,3)[k]) +
                rnorm(m, 0, 8))))
colnames(Y) <- gen

tot  <- rowSums(Y)
base <- -0.5 * tot          # the "themeriver" baseline

plot(NA, xlim = c(1, m),
     ylim = range(c(base, base + tot)),
     axes = FALSE, xlab = "Month", ylab = "")
axis(1)

pal <- c("#0072B2","#D55E00","#009E73",
         "#E69F00","#6A51A3")
for (k in 1:5) {
  top <- base + Y[, k]
  polygon(c(1:m, m:1), c(base, rev(top)),
          col = pal[k], border = "white")
  base <- top
}

# base <- rep(0, m) gives an ordinary stacked area
# chart, where the bottom band sits on a flat axis
# and can be read accurately.
4.Q

4.Q Knowledge check: Section 4

Q1. In a radar chart, what does the area enclosed by the polygon measure?
Area depends on the sequence of vertices around the circle, and that sequence is arbitrary for nominal categories. Two orderings of the same six numbers produce polygons of different area.
Q2. A rose plot sets each wedge's radius directly equal to its count instead of to the square root of the count. What is the effect?
Doubling the radius quadruples the area. Since readers judge these wedges by area, radius-proportional scaling inflates the modal direction.
Q3. Which claim is a streamgraph least able to support?
Band thickness on a curved baseline is read poorly, and the figure has no vertical axis to read against. Streamgraphs support claims about shape, cycle and relative timing.
5
Section 5

Choosing and defending a chart

The question comes first, the chart second. Then you name the limitation before someone in the audience names it for you.

5.1

5.1 Question first, chart second

The question you want to answerChart Watch for
Which pairs among my numeric variables move together? Scatterplot matrix (p ≤ 4), correlation heat map (p > 4) Panel size; a fixed zlim
Does the relationship between two variables depend on a third? Scatter with colour or size; small multiples Shared axis limits across panels
Do my observations fall into groups across many variables at once? Parallel coordinate plot Axis order; overplotting
How does one observation compare with another on a fixed set of variables? Glyph plot (stars), radar chart Arbitrary feature and axis assignment
How does the shape of a distribution differ between groups? Violin plot, side-by-side boxplots Bandwidth; how width was scaled
How is a directional or cyclic variable distributed? Rose plot Radius scaled by the square root
How do the parts of a total change over time? Stacked area chart; streamgraph for shape only No vertical axis on a streamgraph
5.2

5.2 Standardise before you compare

Parallel coordinates, heat maps of raw values and glyph plots all place variables measured in different units on a shared visual scale. Rescale first, or the variable with the largest raw numbers dominates the picture.

Min-max scaling

$$x_i^{\ast} = \frac{x_i - \min(x)}{\max(x) - \min(x)}$$

Maps every value into \([0, 1]\). Use it for parallel coordinates, where each axis needs a fixed top and bottom.

z-scores

$$z_i = \frac{x_i - \bar{x}}{s}$$

Centres each variable on zero and expresses it in standard deviations. Use it for heat maps of raw values, where a diverging palette needs a meaningful midpoint.

In R

# z-scores: scale() does both steps
z <- scale(mtcars[, c("mpg", "hp", "wt")])

round(colMeans(z), 10)
#> mpg  hp  wt
#>   0   0   0

round(apply(z, 2, sd), 6)
#> mpg  hp  wt
#>   1   1   1

# min-max: write it out
mm <- function(v) (v - min(v)) / (max(v) - min(v))
Xs <- apply(iris[, 1:4], 2, mm)

round(range(Xs), 3)
#> 0 1
Report it Say on the slide which scaling you applied. A reader who assumes raw units on a rescaled axis will misread every value on it.
5.3

5.3 Order and angle are analytical choices

ChartThe choice you are making How to defend it
Parallel coordinates Which variables sit next to each other. Only neighbours can be compared directly, so the order determines which relationships a reader can see. Order by correlation, or place the variables your question is about side by side. State the rule you used.
Heat map The order of rows and columns. Alphabetical order scatters a block structure; a sorted order reveals it. Order by cluster or by the first principal component. Alphabetical order needs a reason.
Radar chart The sequence of axes around the circle, which sets the polygon's shape and area. Fix the order by convention and keep it identical across every radar chart in the report.
3D scatterplot The two rotation angles, theta and phi. Redraw at several angles. If your finding survives all of them, report it; if it does not, it was an artefact.
Streamgraph The stacking order of the bands and the shape of the baseline. Put the band you most want read at the bottom, or drop the wiggle and use a flat baseline.
5.4

5.4 Overplotting

Multivariate charts fail first through density. 150 polylines already crowd a parallel coordinate plot; 5,000 turn it into a solid block.

Five remedies in base R

  • Transparency. adjustcolor(col, alpha.f = 0.2). Overlap then reads as darkness, which restores density information.
  • Smaller marks. Reduce cex, or use pch = "." once n runs into the thousands.
  • Jitter. jitter(x) for variables recorded to the nearest whole number, where many points land on identical coordinates.
  • Sample. Plot a random subset and say so on the slide.
  • Bin. smoothScatter(x, y) replaces the points with a density surface.

In R

# transparency: the first thing to try
plot(iris$Sepal.Length, iris$Sepal.Width,
     pch = 19, cex = 1.4,
     col = adjustcolor("#0072B2", alpha.f = 0.25))

# jitter: for values rounded to 0.1 cm
plot(jitter(iris$Sepal.Length, amount = 0.03),
     jitter(iris$Sepal.Width,  amount = 0.03),
     pch = 19, cex = 0.8)

# density surface: for large n
smoothScatter(rnorm(20000), rnorm(20000),
              xlab = "x", ylab = "y")
Order of operations Try transparency first. Move to binning only when transparency alone still leaves a solid block, because binning discards the individual observations.
5.5

5.5 Write the claim, then write the limit

Step 6 of the template produces two sentences. Here is the skeleton, and one completed example from this lesson.

The skeleton

Across [unit, n], [variable A] [rises / falls] as [variable B] [rises / falls], and this pattern [holds / reverses] within [grouping variable].
This chart cannot establish [the limit].

Completed, from slide 1.3

Across 150 iris flowers, sepal width falls slightly as sepal length rises, and this pattern reverses within every one of the three species.
This chart cannot establish why the species differ in level, because it contains no variable that would explain the shift.

Both sentences are marked. A claim with no limit reads as overconfidence; a limit with no claim reads as no finding at all.

5.6

5.6 Summary

Concepts

  • Multivariate counts columns read together, not the number of charts you drew.
  • A bivariate relationship can reverse once a third variable is added. Check before you report a correlation.
  • The page offers two position axes. Everything past the second variable is spent on a channel that is read less accurately.
  • Encode, repeat, or project. Those are the three strategies, and every chart here is one of them.

Procedure

Unit → encoding → scale → structure → groups and exceptions → claim and limit. Six steps, in that order, on every chart.

Charts and their narrow strengths

Scatterplot matrixAll pairs at once, p ≤ 4
Correlation heat mapAll pairs when p is large
Bubble plotTwo on position, two more encoded
3D scatterplotThree on position, if rotatable
Parallel coordinatesWhole rows, group structure
Glyph plot / facesRow-by-row comparison
Small multiplesOne relationship, once per group
Violin plotDistribution shape by group
Rose plotDirectional or cyclic frequency
Radar chartShape recognition on fixed axes
StreamgraphShape and timing of parts of a total
5.7

5.7 Before next week

Practise

  • Reproduce Figures 3.1, 3.2 and 3.4 from the code panels. Run them and change one argument each time: the palette, the zlim, the axis order.
  • Take airquality, which has six variables and missing values. Build a correlation heat map and a scatterplot matrix of it. Note what R does with the missing entries.
  • Apply the six-step read to one chart you find in a news article this week. Write the two sentences from slide 5.5.

Week 8: Missing values

The airquality exercise sets up next week directly. cor() returns NA by default when any value in a pair is missing, and the argument that changes that behaviour is the first decision in the imputation topic.

Bring Your airquality heat map, and a one-line note of what cor(airquality) returned before you did anything about it.
TECH3100 · Lesson 7
← → navigate · T contents

Contents

Press T or Escape to close