TECH3100 · Data Visualisation in R

ggplot, part 2

Distributions, faceting, and finishing a figure
Lesson 11 · plus the Assessment 3 walkthrough
Kaplan Business School Australia
All figures and numbers computed in R 4.3.3 with ggplot2 3.4.4
0.1

Copyright notice

Commonwealth of Australia · Copyright Regulations 1969

This material has been reproduced and communicated to you by or on behalf of Kaplan Business School pursuant to Part VB of the Copyright Act 1968 (the Act).

The material in this communication may be subject to copyright under the Act. Any further reproduction or communication of this material by you may be the subject of copyright protection under the Act.

Do not remove this notice.

0.2

Two halves today

First half: the rest of ggplot

  • Showing a distribution, and what each geom hides.
  • Two variables when there are too many points to draw.
  • Faceting, which is the one thing from Week 7 that ggplot does better than anything you have used so far.
  • Scales against coordinates, annotation, and turning a working plot into one you can present.

Second half: Assessment 3

  • What is due, when, and in what format.
  • The forty marks, criterion by criterion, and what separates a pass from a distinction.
  • How to choose a case study that can actually be answered with the data you can get.
  • A structure for the talk, and a checklist for the submission.
Bring away two things A decision you want to investigate, and a dataset that could answer it. Those two are next week's presentation.
0.3

What you should be able to do by the end

  1. Choose between a histogram, a density curve, a boxplot and a violin, and say what each one throws away.
  2. Set a bin width on purpose and defend the number you chose.
  3. Split a plot into panels with facet_wrap() and facet_grid(), and decide between shared and free scales.
  4. Tell the difference between a scale limit and a coordinate zoom, and predict which one changes a fitted line.
  5. Annotate a chart so a reader knows where to look.
  6. Plan, deliver and submit Assessment 3 against the published rubric.
The one that will cost you marks if you skip it Outcome 4. Zooming with ylim() silently deletes rows, and any statistic drawn on top of it is then computed from what survived.
0.4

Setup and the datasets

library(ggplot2)      # already installed from Week 10

# today's datasets, all built in
faithful   # 272 Old Faithful eruptions
mpg        # 234 cars, from ggplot2
mtcars     # 32 cars
iris       # 150 flowers

str(faithful)
#> 'data.frame': 272 obs. of  2 variables:
#>  $ eruptions: num  3.6 1.8 3.33 2.28 4.53 ...
#>  $ waiting  : num  79 54 74 62 85 ...
One optional package geom_hex() needs install.packages("hexbin"). Everything else today runs in a fresh session with ggplot2 alone.

faithful, and why we start there

Old Faithful is a geyser in Wyoming. The dataset records how long each eruption lasted and how long the wait was before it.

Statisticeruptions (minutes)
n272
minimum1.600
first quartile2.163
median4.000
third quartile4.454
maximum5.100
mean3.488

Those seven numbers describe this variable badly, and the next slide shows why.

Section 1

Showing a distribution

Four geoms answer the same question. They disagree, and the disagreement is the lesson.
1.1

Four views of one variable

geom_histogram()counts in bins2345eruptionscountgeom_density()a smoothed estimate2345eruptionsdensitygeom_boxplot()five numbers only2345eruptionsgeom_violin()shape and summary2345The same 272 Old Faithful eruption times, four ways. The first, second and fourth show two separate groups.The third does not.The boxplot reports a median of 4.00 and a box running from 2.16 to 4.45. Only 7 of the 272 eruptionsfall between 3.0 and 3.5, so the middle of that box is close to empty.A boxplot summarises. Ask what it summarised away before you publish one.
Figure 1.1. The 272 eruption times from faithful, drawn four ways in ggplot2 3.4.4.
Analysis
R code
  • Unit
    One bar is a bin of eruption times; one box is all 272 of them; the violin and the density curve are smoothed estimates over the same 272.
  • Encoding
    The first, second and fourth put eruption length on a position axis and frequency on the other. The third puts five summary numbers on a position axis and nothing else.
  • Scale
    All four cover 1.6 to 5.1 minutes, so the differences are the geom, not the range.
  • Structure
    Three of the four show two separate groups of eruptions, one around two minutes and one around four and a half.
  • Groups and exceptions
    The boxplot shows one box with a median of 4.00 and a box from 2.163 to 4.454. It gives no sign that the data has two modes.
  • Claim and limit
    A boxplot is a summary of five numbers, and five numbers cannot describe two groups. Only 7 of the 272 eruptions fall between 3.0 and 3.5, so the centre of that box sits in a region where there is almost no data.
# counts in bins
ggplot(faithful, aes(eruptions)) +
  geom_histogram(binwidth = 0.25)

# a smoothed estimate of the same thing
ggplot(faithful, aes(eruptions)) +
  geom_density(fill = "steelblue", alpha = 0.3)

# five numbers
ggplot(faithful, aes(y = eruptions)) +
  geom_boxplot()

# shape and summary together
ggplot(faithful, aes(x = "", y = eruptions)) +
  geom_violin(fill = "grey90") +
  geom_boxplot(width = 0.1)

# or show the data itself alongside
ggplot(faithful, aes(x = "", y = eruptions)) +
  geom_violin(fill = "grey90") +
  geom_jitter(width = 0.08, alpha = 0.4)

summary(faithful$eruptions)
#>  Min. 1st Qu.  Median    Mean 3rd Qu.    Max.
#> 1.600   2.163   4.000   3.488   4.454   5.100
1.2

What the box threw away

2345eruption length (min)what you showwhat you havethe shape7 of 272 eruptions land in this bandThe median is 4.00 and the box spans 2.16 to 4.45, so the box covers a gap rather than a centre.97 eruptions are short and 175 are long. There is no typical eruption.Every summary throws something away. This one threw away the only thing that mattered.Plot the distribution before you summarise it.
Figure 1.2. The same 272 eruptions as a boxplot, as raw jittered points, and as a density curve, on one shared axis.
Analysis
R code
  • Unit
    One point in the middle panel is one eruption. The box and the curve each summarise all 272.
  • Encoding
    Vertical position is eruption length in every panel, so the three are directly comparable.
  • Scale
    One shared vertical axis from 1.4 to 5.3 minutes.
  • Structure
    The raw points fall into two clouds with a visible gap between them. The density curve has peaks at 1.98 and 4.37 minutes.
  • Groups and exceptions
    The shaded band marks 3.0 to 3.5 minutes. Seven of the 272 eruptions land there. 97 eruptions are shorter than three minutes and 175 are longer.
  • Claim and limit
    There is no typical eruption. The median of 4.00 is a real number and it describes nothing. Reporting it alone would tell a reader that eruptions last about four minutes, when in fact more than a third of them last about two.
The habit Draw the distribution before you summarise it. It costs one line and it is the only way to find out whether your summary is honest.
# the summary
ggplot(faithful, aes(y = eruptions)) + geom_boxplot()

# the data behind it
ggplot(faithful, aes(x = "", y = eruptions)) +
  geom_jitter(width = 0.15, alpha = 0.45)

# both together, which is what you should draw
ggplot(faithful, aes(x = "", y = eruptions)) +
  geom_violin(fill = "grey92") +
  geom_jitter(width = 0.08, alpha = 0.4) +
  labs(x = NULL, y = "Eruption length (minutes)")

# the numbers behind the claim
sum(faithful$eruptions > 3 & faithful$eruptions < 3.5)
#> 7
sum(faithful$eruptions < 3)    #> 97
sum(faithful$eruptions >= 3)   #> 175
median(faithful$eruptions)     #> 4
1.3

Bin width is a decision

bins = 52 peaks, barelyeruptions82bins = 122 clear peakseruptions49bins = 30 (the default)5 local peakseruptions23bins = 8022 local peakseruptions12The same 272 numbers every time. The only thing changing is how many bins they are dropped into.Too few bins and the two groups merge. Too many and every small wobble looks like a group of its own.ggplot's default of 30 bins is a placeholder, and it prints a message saying so. Set binwidth yourself,in the units of the variable, and say what you set it to.
Figure 1.3. The same 272 numbers dropped into 5, 12, 30 and 80 bins.
Analysis
R code
  • Unit
    One bar is one bin. The number of bins is the only thing changing across the four panels.
  • Encoding
    Horizontal position is eruption length, bar height is how many eruptions fell in that bin.
  • Scale
    Each panel is scaled to its own tallest bar, which is shown at the top left. 82, then 49, then 23, then 12.
  • Structure
    The two groups are present in all four panels, but they are easiest to see in the second.
  • Groups and exceptions
    At 5 bins the two groups nearly merge into one wide block. At 80 bins there are 22 local peaks, and 21 of them are noise.
  • Claim and limit
    The bin width decides what the reader sees. ggplot's default is 30 bins, which is a placeholder rather than a choice, and it prints a message telling you so. Choose a width in the units of the variable and state it.
# the default, with the message it prints
ggplot(faithful, aes(eruptions)) + geom_histogram()
#> `stat_bin()` using `bins = 30`. Pick better value
#> with `binwidth`.

# say what you mean, in minutes
ggplot(faithful, aes(eruptions)) +
  geom_histogram(binwidth = 0.25)

# or in bins, if the count is what matters
ggplot(faithful, aes(eruptions)) +
  geom_histogram(bins = 12)

# two groups on one axis
ggplot(mpg, aes(hwy, fill = drv)) +
  geom_histogram(binwidth = 2, position = "identity",
                 alpha = 0.5)

# position = "identity" overlaps them.
# The default, "stack", makes the upper groups
# unreadable because they no longer start at zero.
1.4

Which of the four to draw

GeomShowsHides Reach for it when
geom_histogram()the shape, as counts you can add up anything narrower than the bin width one variable, and the reader cares about how many
geom_density()the shape, smoothed the sample size, and the gaps between observations comparing the shape of two or three groups
geom_boxplot()median, quartiles, outliers the shape entirely, including extra modes many groups side by side, where shape is secondary
geom_violin()the shape, mirrored, per group the individual observations a handful of groups where shape matters
+ geom_jitter()every observation nothing, which is the point n is under about 500, so add it to any of the above
The ranking is about how much is thrown away Jitter throws away nothing. The histogram throws away detail below the bin width. The boxplot throws away the shape. The further down you go, the more you must already know about the data to be sure the summary is safe.
Boxplots are still the right choice often Twelve groups side by side is a job for boxplots and nothing else. The rule is not to avoid them; it is to look at the shape once, yourself, before you publish the summary.
1.Q

Knowledge check: Section 1

Q1. A boxplot of faithful$eruptions shows a median of 4.00 and a box from 2.16 to 4.45. What does that tell you about the shape?
The data is centred near 4 with moderate spread.
Nothing. A boxplot cannot show that the data has two separate groups.
The data is right-skewed.
Only 7 of 272 eruptions fall between 3.0 and 3.5, so the middle of that box is close to empty. Five summary numbers cannot represent two modes, and no boxplot will ever warn you.
Q2. geom_histogram() prints a message about bins = 30. What is it telling you?
The plot will fail without a binwidth.
30 is a placeholder, not a choice, and the shape you see depends on it.
The data has 30 distinct values.
The same 272 numbers show two clear groups at 12 bins and 22 local peaks at 80. The message is asking you to make the decision rather than inherit it.
Q3. You have nine groups and want to compare their spread on one chart. Which geom fits best?
Nine density curves on shared axes.
Boxplots side by side, after you have looked at the shapes yourself.
A histogram with fill mapped to group.
Nine density curves overlap into noise and nine stacked histograms are unreadable. Boxplots handle many groups well. The caution from this section is to check the shapes once before you commit to the summary.
Section 2

Two variables at once

What to draw when a scatterplot has more marks than the page can separate, and how to put a Week 9 aggregate straight onto a chart.
2.1

Counting instead of drawing

Every row as a markgeom_point()246203040displhwyRows counted into binsgeom_bin2d(binwidth = c(0.5, 2))246203040displhwycount171234 cars occupy only 126 distinct positions, so the left panel hides 108 of them. The right panel countsthem into 53 bins and colours each by how many landed there. The busiest bin holds 17 cars.Binning trades the individual observation for an honest count. Use it once transparency stops coping.
Figure 2.1. 234 cars as individual points, and the same 234 counted into rectangular bins.
Analysis
R code
  • Unit
    Left: one mark is one car. Right: one rectangle is a bin of engine size and economy, coloured by how many cars fell in it.
  • Encoding
    Position gives engine size and highway economy in both. The right panel adds colour for the count.
  • Scale
    Identical axes. The colour scale runs from 1 car to 17, using ggplot's default blue gradient.
  • Structure
    Economy falls as engine size rises, in both panels.
  • Groups and exceptions
    The 234 cars occupy only 126 distinct positions, so 108 marks in the left panel are drawn on top of another mark. The busiest bin on the right holds 17 cars.
  • Claim and limit
    Binning trades the individual observation for an honest count. The right panel cannot show you one particular car, and the bin width you choose changes the picture exactly as it does for a histogram.
# the problem, measured before you draw
nrow(mpg)                                 #> 234
nrow(unique(mpg[, c("displ", "hwy")]))    #> 126

# rectangular bins, width in data units
ggplot(mpg, aes(displ, hwy)) +
  geom_bin2d(binwidth = c(0.5, 2))

# hexagonal bins, which tile more evenly
# needs install.packages("hexbin")
ggplot(mpg, aes(displ, hwy)) +
  geom_hex(bins = 20)

# what ggplot computed
p <- ggplot(mpg, aes(displ, hwy)) +
       geom_bin2d(binwidth = c(0.5, 2))
d <- ggplot_build(p)$data[[1]]
nrow(d)         #> 53   bins drawn
max(d$count)    #> 17   busiest bin
sum(d$count)    #> 234  every car accounted for

# the escalation, in order
#   1. alpha       up to a few thousand points
#   2. geom_bin2d  tens of thousands
#   3. geom_hex    same, better tiling
2.2

A Week 9 table, straight onto a chart

drv4fr2seater005compact12350midsize3380minivan0110pickup3300subcompact4229suv51011class9 of the 21 cells hold no cars at all.geom_tile() draws one rectangle per row of an aggregate table, so this is a Week 9 resultwith a colour scale on top. It needs one row per cell, including the empty ones.A missing row and a zero look different. Decide which you mean, and make the chart say it.
Figure 2.2. Counts of cars by class and drive type, drawn with geom_tile().
Analysis
R code
  • Unit
    One rectangle is one combination of class and drive type.
  • Encoding
    Vertical position gives the class, horizontal position the drive type, and fill colour the count. The number is printed as well, because colour alone cannot be read to a value.
  • Scale
    A single blue gradient from the smallest count to the largest, 51.
  • Structure
    Most classes use one drive type almost exclusively. Pickups are all four-wheel drive; minivans are all front-wheel drive.
  • Groups and exceptions
    Nine of the 21 cells hold no cars at all. subcompact is the only class that appears in all three drive types.
  • Claim and limit
    Drive type is close to determined by vehicle class in this dataset. The empty cells are the finding, so the chart has to show them. A table with only the twelve non-empty rows would have hidden the pattern.
Zero and missing look the same A white tile can mean "no cars" or "no row in the table". Decide which you mean, and use geom_text() or a distinct fill so the reader knows.
# Week 9 produced the table
tab <- as.data.frame(table(mpg$class, mpg$drv))
names(tab) <- c("class", "drv", "n")
nrow(tab)          #> 21   every combination
sum(tab$n == 0)    #>  9   including the empty ones

# Week 11 draws it
ggplot(tab, aes(x = drv, y = class, fill = n)) +
  geom_tile(colour = "white") +
  geom_text(aes(label = n), colour = "white",
            size = 3) +
  scale_fill_gradient(low = "#132B43",
                      high = "#56B1F7") +
  labs(x = "Drive type", y = NULL, fill = "Cars")

# geom_tile needs one row per cell.
# aggregate() and table() differ here:
#   table()     returns every combination, zeros included
#   aggregate() returns only the combinations present
# If you aggregate first, the empty cells vanish
# from the chart without any warning.
2.3

Three labels that do not match their charts

The problem
The fix, in code

Three examples in circulation for this topic carry titles that describe a different chart from the one the code draws. They are worth a minute because the same mistake is the easiest way to lose marks in Assessment 3.

The label saysThe code actually does Why it matters
"Car horsepower and torque" aes(x = hp, y = wt), and wt is weight in thousands of pounds A reader who knows cars will assume you do not.
"Temperature variations by month" ggplot(mtcars, aes(factor(cyl), mpg)). mtcars has no temperature and no month. The axis labels contradict the title on the same slide.
"Stacked bar plots" geom_boxplot() Naming the geom wrongly suggests you do not know which one you used.
The check Read your title, then read your aes() call, then read your axis labels. All three have to name the same variables. It takes ten seconds per chart and it is worth marks under two separate criteria.

Each one is a one-line correction. The point is the habit, not the syntax.

# say what the axes actually are
ggplot(mtcars, aes(x = hp, y = wt)) +
  geom_bin2d(binwidth = c(50, 0.5)) +
  labs(title = "Heavier cars carry bigger engines",
       x = "Horsepower", y = "Weight (1000 lbs)")

# a title that names the variables in the aes()
ggplot(mtcars, aes(factor(cyl), mpg)) +
  geom_boxplot() +
  labs(title = "Fuel economy falls as cylinders rise",
       x = "Cylinders", y = "Miles per gallon")

# and if you want a genuinely stacked bar
ggplot(mpg, aes(class, fill = drv)) +
  geom_bar(position = "stack") +
  labs(title = "Drive type by vehicle class",
       x = NULL, y = "Cars", fill = "Drive")
Write the title last Draw the chart, read it, then write the title as the sentence you would say out loud about it. A title written before the chart describes what you hoped to find.
2.Q

Knowledge check: Section 2

Q1. Your scatterplot of 50,000 rows is a solid black mass. What is the first thing to try?
geom_bin2d() or geom_hex(), because alpha will still be saturated at that size.
Reduce the point size to 0.1.
Take a random sample of 500 rows.
Transparency copes with a few thousand points. At 50,000 even alpha = 0.01 saturates in the dense region. Binning gives you an honest count per region. Sampling is a last resort and has to be disclosed.
Q2. You build a tile heat map from aggregate() and three cells are missing from the chart entirely. Why?
geom_tile() drops cells with a count of zero.
aggregate() returns only the combinations that occur, so those rows never existed.
The fill scale needs an explicit limit.
table() returns every combination including the zeros; aggregate() returns only what is present. geom_tile draws one rectangle per row, so a missing row is a missing tile, and nothing warns you.
Q3. A chart is titled "Horsepower and torque" and its aes() reads aes(x = hp, y = wt). Under the Assessment 3 rubric, which criteria does that damage?
Only data visualisation design.
Both data visualisation design and analysis, because the interpretation now describes a variable that is not on the chart.
Neither, as long as the chart itself is correctly drawn.
The rubric marks design and technical execution separately from analysis and interpretation. A mislabelled axis makes the chart wrong under the first and makes every sentence you write about it wrong under the second.
Section 3

Faceting

One more layer splits a plot into a grid of panels. This is the Week 7 idea of small multiples, and it is the fastest way to ask whether a pattern holds everywhere.
3.1

One panel per group

2seater2462040compact246midsize246minivan246pickup2462040subcompact246suv246facet_wrap(~ class)one panel per level, wrapped into a gridSeven panels, 234 points shared between them, and every panel on the same axes.That shared scale is what makes the panels comparable, and ggplot does it by default. Week 7 called this small multiples.
Figure 3.1. mpg split into seven panels by vehicle class, on shared axes.
Analysis
R code
  • Unit
    One point is one car. The 234 cars are divided between the seven panels and none is drawn twice.
  • Encoding
    Position gives engine size and highway economy inside every panel. The panel itself encodes the class.
  • Scale
    Every panel runs 1.3 to 7.2 litres and 10 to 46 mpg. ggplot shares the scales by default, which is what makes the panels comparable.
  • Structure
    Economy falls with engine size inside most panels, and the panels sit at visibly different heights.
  • Groups and exceptions
    2seater holds 5 cars in the top right of its panel. pickup and suv sit low. compact and subcompact sit high and reach 44 mpg.
  • Claim and limit
    The relationship holds within classes, and the class itself accounts for much of the spread. With 5 cars in one panel and 62 in another, the panels are not equally reliable, and the chart does not show the counts.
# one panel per level of one variable
ggplot(mpg, aes(displ, hwy)) +
  geom_point() +
  facet_wrap(~ class)

# control the layout
facet_wrap(~ class, ncol = 4)
facet_wrap(~ class, nrow = 2)

# the tilde means "by". It is a formula, so the
# variable name is not quoted.

# how many panels, and how many points in each
table(mpg$class)
#>    2seater    compact    midsize    minivan
#>          5         47         41         11
#>     pickup subcompact        suv
#>         33         35         62

# facet on two variables with wrap, and you get
# one panel per observed combination
facet_wrap(~ class + drv)

# Shared scales are the default and they are
# what makes panels comparable. Changing that
# is the next slide.
3.2

Shared scales against free ones

Shared scalesfacet_wrap(~ class)2seater2040compact2040pickup2040Free scalesfacet_wrap(~ class, scales = "free")2seater22.52527.5compact3040pickup101520Left: every panel runs 10 to 46 mpg, so the pickup cloud sits visibly below the compact one.Right: each panel fits its own data, so all three clouds fill their panel and look alike.2seater covers 23 to 26 mpg and compact covers 23 to 44. On free scales those two look equally spread.Use free scales to see shape inside a panel. Use shared scales whenever the reader will compare panels,which is almost always. Say which you used.
Figure 3.2. Three of the seven classes, first on shared axes and then with scales = "free".
Analysis
R code
  • Unit
    One point is one car in all six panels. The same cars appear on both sides.
  • Encoding
    Position gives engine size and economy throughout. Only the axis limits change between the two halves.
  • Scale
    Left: every panel runs 10 to 46 mpg. Right: each panel is fitted to its own data, so 2seater runs 22 to 27 and compact runs 22 to 45.
  • Structure
    On the left the three clouds sit at different heights and the pickup cloud is clearly lowest. On the right all three fill their panel.
  • Groups and exceptions
    2seater covers 23 to 26 mpg and compact covers 23 to 44. On free scales those two look equally spread, and they are not.
  • Claim and limit
    Free scales show shape inside a panel and destroy comparison between panels. Almost every reason to facet is a comparison, so shared scales are almost always right. When you do use free scales, say so on the slide.
# the default: every panel on the same axes
facet_wrap(~ class)

# each panel fitted to its own data
facet_wrap(~ class, scales = "free")

# free on one axis only
facet_wrap(~ class, scales = "free_y")
facet_wrap(~ class, scales = "free_x")

# the ranges that produce the difference
aggregate(hwy ~ class, mpg, range)
#>        class  hwy.1 hwy.2
#>      2seater     23    26
#>      compact     23    44
#>       pickup     12    22

# A reader assumes panels are comparable unless
# told otherwise, because that is the default.
# Using free scales without saying so is the
# faceting equivalent of truncating an axis.
3.3

A matrix of panels

cyl4568empty4femptyemptyrdrvfacet_grid(drv ~ cyl)one panel per combination, laid out as a matrixTwelve panels for three drive types crossed with four cylinder counts. Three are empty.facet_grid keeps the empty cells, and that is the point: no rear-wheel four-cylinder car exists here,and the gap in the matrix says so. facet_wrap would have hidden it.
Figure 3.3. mpg split by drive type down the rows and cylinder count across the columns.
Analysis
R code
  • Unit
    One point is one car; one panel is one combination of drive type and cylinder count.
  • Encoding
    Position gives engine size and economy inside each panel. Row position encodes drive type and column position encodes cylinders.
  • Scale
    Shared across all twelve panels, so any panel can be compared with any other.
  • Structure
    Cars move down and to the right as cylinders increase: more cylinders means bigger engines and worse economy, in every drive type.
  • Groups and exceptions
    Three of the twelve panels are empty. There is no rear-wheel four-cylinder car and no rear-wheel five-cylinder car in this dataset, and only one front-wheel eight-cylinder.
  • Claim and limit
    The empty cells are a finding, not a gap. facet_grid keeps them, so the matrix shows you which combinations do not exist. facet_wrap would have silently dropped all three.
# rows ~ columns
ggplot(mpg, aes(displ, hwy)) +
  geom_point() +
  facet_grid(drv ~ cyl)

# one row, several columns
facet_grid(. ~ drv)

# several rows, one column
facet_grid(drv ~ .)

# the combinations that exist
table(mpg$drv, mpg$cyl)
#>      4  5  6  8
#>  4  23  0 32 48
#>  f  58  4 43  1
#>  r   0  0  4 21

# facet_grid  keeps every combination, so gaps show
# facet_wrap  keeps only what occurs, so gaps vanish
#
# Use grid when the absence of a combination is
# part of what you are reporting. Use wrap when
# you only have one variable, or when the empty
# cells would waste the page.
3.4

When to facet, and when not to

FACET WHEN

The question is "does it hold everywhere"

Week 7's reversal, Week 9's grouped means and Week 10's three fitted lines were all versions of this question. Faceting answers it directly.

This is the main use
FACET WHEN

Colour has run out

Above about six categories, colours stop being distinguishable. Panels have no such limit.

Seven classes is already too many for colour
FACET WHEN

The groups overlap heavily

Three overlapping clouds in one panel hide each other. Three panels separate them without changing a single value.

Cheaper than transparency
DO NOT WHEN

The comparison is the point and the groups are few

Two or three groups compared directly are easier to read in one panel with colour. Panels force the eye to travel.

Overlaying beats faceting at n = 2
DO NOT WHEN

Some panels would hold almost no data

A panel with 5 points invites a conclusion it cannot support. Report the counts, or combine the small groups.

2seater has 5 cars
DO NOT WHEN

You would need free scales to make it work

If the panels only look reasonable on free scales, the faceting variable is probably doing something other than what you think.

Ask why the ranges differ
3.Q

Knowledge check: Section 3

Q1. facet_wrap(~ class) gives seven panels and every one has the same axes. Did you ask for that?
Yes, shared scales must be requested explicitly.
No. Shared scales are the default, and they are what makes the panels comparable.
No, and it is a bug that requires scales = "fixed".
ggplot shares the scales unless you say otherwise. That default is doing real work: a reader assumes panels are comparable, and they are, until somebody adds scales = "free".
Q2. facet_grid(drv ~ cyl) produces twelve panels and three are empty. What should you do?
Switch to facet_wrap so the empty panels disappear.
Leave them. The empty cells report that those combinations do not exist.
Filter the data so only complete combinations remain.
No rear-wheel four-cylinder car exists in mpg. The gap in the matrix is the finding. facet_wrap would hide it by drawing only the nine combinations that occur.
Q3. On free scales, 2seater (23 to 26 mpg) and compact (23 to 44 mpg) look equally spread. What went wrong?
Nothing. Free scales are designed to show shape within a panel.
Nothing technically, but the reader will compare the panels anyway, so the spread is being misreported.
The facet variable should have been cylinders.
Free scales are a legitimate tool for seeing shape. The problem is that panels side by side invite comparison, and the reader has no way to know the axes differ unless you say so on the slide.
Section 4

Finishing a figure

Everything between a plot that works and a plot you would put in front of a client. Two of these will be marked directly next week.
4.1

The clustered bar chart, done properly

Two geom_bar layersone aes(y=) per layer0246setosaversicolovirginicaSpeciesThe second layer sits on top of the first.Nothing is dodged, because the two barsare in different layers and neither oneknows the other exists.One layer, reshaped firstaes(fill = measure), position = dodge02461.460.25setosa4.261.33versicolo5.552.03virginicaSpeciesPetal.LengthPetal.WidthBoth bars visible, side by side, on ashared scale.position = "dodge" separates groups inside one layer. It cannot separate two layers.Reshape to long form first, so the thing you want side by side is a column you can map to fill.
Figure 4.1. Average iris petal length and width by species, drawn with two geom_bar layers and then with one layer on reshaped data.
Analysis
R code
  • Unit
    One bar is the mean of one measurement for one species, over 50 flowers.
  • Encoding
    Horizontal position gives the species. Bar height gives the mean in centimetres. On the right, fill gives which measurement.
  • Scale
    Both panels run 0 to 6.4 centimetres, so the bars are directly comparable.
  • Structure
    Both petal length and petal width rise across the three species.
  • Groups and exceptions
    On the left the width bars are hidden behind the length bars, because a second layer draws over the first. On the right both are visible: setosa 1.46 and 0.25, versicolor 4.26 and 1.33, virginica 5.55 and 2.03.
  • Claim and limit
    Petals do not just get bigger, they change shape. The length-to-width ratio falls from 5.94 in setosa to 3.21 and then 2.74, so virginica petals are proportionally much wider. The left panel makes that impossible to see.
The general rule Anything you want side by side has to be a column you can map to an aesthetic. If it is currently two columns, reshape first. This is why tidy data matters.
# What does not work, and why
ggplot(iris_summary, aes(x = Species)) +
  geom_bar(aes(y = Avg_Petal_Length), stat = "identity",
           position = "dodge") +
  geom_bar(aes(y = Avg_Petal_Width), stat = "identity",
           position = "dodge")
# position = "dodge" separates groups WITHIN one
# layer. It cannot separate two layers, so the
# second simply draws on top of the first.

# Step 1: aggregate (Week 9)
ag <- aggregate(cbind(Petal.Length, Petal.Width) ~ Species,
                data = iris, FUN = mean)

# Step 2: reshape to one row per bar
lg <- reshape(ag, direction = "long",
              varying = c("Petal.Length", "Petal.Width"),
              v.names = "value", timevar = "measure",
              times = c("Petal.Length", "Petal.Width"),
              idvar = "Species")
nrow(ag)   #> 3
nrow(lg)   #> 6   one row per bar

# Step 3: one layer, fill does the clustering
ggplot(lg, aes(Species, value, fill = measure)) +
  geom_col(position = "dodge", width = 0.7) +
  labs(y = "Mean measurement (cm)", fill = NULL)
4.2

A summary needs its distribution

Mean and median as barstwo numbers, one chart0102020.09Mean19.20MedianThe bars differ by 0.89 mpg.The chart shows nothing else:no spread, no shape, no n.The distribution behind them32 cars, with both statistics marked102030mean 20.09median 19.20mpgThe mean sits above the median because theright tail is longer. sd is 6.03 across 32 cars.A bar chart of two summary numbers is a table with extra ink.Draw the distribution and mark the summaries on it.
Figure 4.2. The mean and median of mtcars mpg as two bars, and the same two numbers drawn on the distribution they came from.
Analysis
R code
  • Unit
    Left: one bar is one summary statistic. Right: one bar is a bin of 2 mpg, and the dashed lines are the two statistics.
  • Encoding
    Left: bar height is the value of the statistic. Right: horizontal position is mpg and bar height is how many of the 32 cars fell in that bin.
  • Scale
    Left runs 0 to 24 mpg. Right runs 9 to 35 mpg, covering the actual range.
  • Structure
    The mean is 20.09 and the median is 19.20, a difference of 0.89 mpg.
  • Groups and exceptions
    The right panel shows why they differ: a longer right tail pulls the mean above the median. The standard deviation is 6.03 across 32 cars, which is far larger than the gap between the two statistics.
  • Claim and limit
    Two bars differing by 0.89 is a table drawn with extra ink. The left panel has no spread, no shape and no sample size, so a reader cannot judge whether 0.89 is a lot. The right panel answers that in one look.
# the chart to avoid
ggplot(data.frame(x = c("Mean", "Median"),
                  y = c(mean(mtcars$mpg),
                        median(mtcars$mpg))),
       aes(x, y)) +
  geom_col(fill = "steelblue")

# the chart to draw instead
ggplot(mtcars, aes(mpg)) +
  geom_histogram(binwidth = 2, fill = "grey40") +
  geom_vline(xintercept = mean(mtcars$mpg),
             colour = "red", linetype = "dashed") +
  geom_vline(xintercept = median(mtcars$mpg),
             colour = "blue", linetype = "dashed") +
  labs(x = "Miles per gallon", y = "Cars")

mean(mtcars$mpg)     #> 20.091
median(mtcars$mpg)   #> 19.2
sd(mtcars$mpg)       #> 6.027

# when you do want summary bars, show the
# uncertainty on them
ggplot(mtcars, aes(factor(cyl), mpg)) +
  stat_summary(fun = mean, geom = "col",
               fill = "grey60") +
  stat_summary(fun.data = mean_se,
               geom = "errorbar", width = 0.2)
4.3

A limit is not a zoom

ylim(20, 40)rows outside the range are deleted2462025303540displhwy22.7coord_cartesian(ylim = c(20, 40))every row is kept, the view is zoomed2462025303540displhwy17.2Left: 81 of the 234 cars are gone before the smooth is fitted, so the curve is fitted to 153 rows.Right: all 234 rows are used and only the window changes.At displ = 5 the two fitted curves read 22.7 and 17.2. The difference is 5.5 mpg, produced by a changethat looks purely cosmetic. ggplot warns about the deleted rows once, in the console.
Figure 4.3. The same scatter and loess smooth, restricted to 20 to 40 mpg two different ways.
Analysis
R code
  • Unit
    One point is one car; the blue line is a loess fit over whatever rows survived.
  • Encoding
    Position gives engine size and highway economy. Both panels show the same window, 20 to 40 mpg.
  • Scale
    Identical visible ranges, which is exactly why the difference is hard to spot.
  • Structure
    Economy falls with engine size in both panels.
  • Groups and exceptions
    ylim keeps 153 of the 234 cars, deleting 81 before the smooth is fitted. coord_cartesian keeps all 234 and fits the smooth to every one of them.
  • Claim and limit
    At displ = 5 the two fitted curves read 22.7 and 17.2. A difference of 5.5 mpg was produced by a change that looks purely cosmetic. ggplot warns about the removed rows once, in the console, and never again.
# DELETES rows, then fits
ggplot(mpg, aes(displ, hwy)) +
  geom_point() +
  geom_smooth(method = "loess", formula = y ~ x) +
  ylim(20, 40)
#> Warning: Removed 81 rows containing non-finite
#> values (`stat_smooth()`).

# KEEPS rows, then zooms
ggplot(mpg, aes(displ, hwy)) +
  geom_point() +
  geom_smooth(method = "loess", formula = y ~ x) +
  coord_cartesian(ylim = c(20, 40))

# these two are the same thing:
ylim(20, 40)
scale_y_continuous(limits = c(20, 40))

# proof
a  <- ggplot(mpg, aes(displ, hwy)) + geom_point() +
        geom_smooth(method = "loess", formula = y ~ x)
b1 <- ggplot_build(a + ylim(20, 40))
b2 <- ggplot_build(a + coord_cartesian(ylim = c(20, 40)))
sum(!is.na(b1$data[[1]]$y))   #> 153
sum(!is.na(b2$data[[1]]$y))   #> 234

# Rule: to change what is COMPUTED, use a scale
#       limit. To change what is SEEN, use
#       coord_cartesian. Cropping a view is
#       almost always what you meant.
4.4

Telling the reader where to look

234567203040Engine displacement (litres)Highway economy (mpg)Five Corvettes:big engine, mid economyBigger engines use more fuel, with one clear exception234 cars, model years 1999 and 2008Source: US EPA fuel economy dataOne colour, one arrow, one sentence of text. The reader is told where to look and why.annotate() and a highlighted subset carry more of a presentation than any theme does.
Figure 4.4. The mpg scatter with one subset highlighted, an arrow, and a caption.
Analysis
R code
  • Unit
    One point is one car. Five of the 234 are drawn in red.
  • Encoding
    Position gives engine size and economy. Colour separates one subset from the rest, which are pushed back to grey.
  • Scale
    1.3 to 7.3 litres and 10 to 46 mpg, covering every car.
  • Structure
    Economy falls as engine size rises, which is the main trend and is now background.
  • Groups and exceptions
    The five highlighted cars are all Chevrolet Corvettes. They average 6.16 litres, larger than almost anything else, yet manage 24.8 mpg, which is better than most cars half their size.
  • Claim and limit
    Engine size predicts economy well, except for light sports cars. Five cars is a small exception and the chart does not establish why they differ. Weight would be the obvious next variable to check.
Worth marks next week One highlighted subset with one sentence of annotation does more for a presentation than any theme. The audience is following you at your pace, and the arrow tells them where to be.
two <- subset(mpg, class == "2seater")

ggplot(mpg, aes(displ, hwy)) +
  geom_point(colour = "grey70") +
  geom_point(data = two, colour = "#CC0000",
             size = 2.5) +
  annotate("text", x = 4.3, y = 34,
           label = "Five Corvettes",
           colour = "#CC0000") +
  annotate("segment", x = 4.7, xend = 5.9,
           y = 33, yend = 26, colour = "#CC0000",
           arrow = arrow(length = unit(2, "mm"))) +
  labs(title = "Bigger engines use more fuel, with one exception",
       subtitle = "234 cars, model years 1999 and 2008",
       x = "Engine displacement (litres)",
       y = "Highway economy (mpg)",
       caption = "Source: US EPA fuel economy data") +
  theme_minimal()

# the pattern: grey everything, then redraw the
# subset you care about on top of it.

# reference lines
geom_hline(yintercept = mean(mpg$hwy), linetype = "dotted")
geom_vline(xintercept = 4)
geom_abline(slope = -3.53, intercept = 35.7)

nrow(two)            #> 5
mean(two$displ)      #> 6.16
mean(two$hwy)        #> 24.8
4.5

From working to presentable

What R gives youthree lines of code246203040displhwydrv4frWhat you presentthe same three lines plus eight moreHeavier cars pay for it at the pump234 US models, 1999 and 2008203040246Engine displacement (litres)Highway mpgSource: US EPAFour wheelFront wheelRear wheelIdentical data, identical geoms. The right panel adds a title that states a finding, units on both axes,a readable legend, a source, and a palette that survives a projector.
Figure 4.5. The same data and the same geoms, before and after eight more lines.
Analysis
R code
  • Unit
    One point is one car in both panels, 234 each.
  • Encoding
    Identical: position for engine size and economy, colour for drive type.
  • Scale
    Identical ranges. The right panel changes the palette to three shades of one hue rather than three separate hues.
  • Structure
    The same falling cloud appears in both.
  • Groups and exceptions
    The left panel labels its axes displ and hwy and titles itself nothing. The right names the quantities, their units, the sample and the source.
  • Claim and limit
    The difference between the two is entirely presentation, and it is worth marks under two criteria. Nothing about the analysis improved. What improved is whether somebody who was not in the room can read it.
p <- ggplot(mpg, aes(displ, hwy, colour = drv)) +
       geom_point()

p +
  scale_colour_manual(
    values = c("4" = "#1B4965",
               "f" = "#5FA8D3",
               "r" = "#CAE9FF"),
    labels = c("Four wheel", "Front wheel",
               "Rear wheel")) +
  labs(title = "Heavier cars pay for it at the pump",
       subtitle = "234 US models, 1999 and 2008",
       x = "Engine displacement (litres)",
       y = "Highway economy (mpg)",
       colour = "Drive",
       caption = "Source: US EPA") +
  theme_minimal(base_size = 13) +
  theme(legend.position = "bottom",
        plot.title = element_text(face = "bold"))

# export at the size it will be shown
ggsave("engine-economy.png",
       width = 10, height = 5.6, dpi = 300)

# 10 by 5.6 inches is 16:9, which fills a slide.
# Text in ggplot is sized in points and does not
# scale, so saving small and enlarging in
# PowerPoint gives you coarse, oversized text.
4.Q

Knowledge check: Section 4

Q1. Two geom_bar() layers with different aes(y = ) and position = "dodge" draw one visible bar per group. Why?
dodge needs a width argument to take effect.
dodge separates groups within a single layer, so two layers just overlap.
stat = "identity" disables dodging.
Dodging is a position adjustment inside one layer. Two layers know nothing about each other. Reshape the data so the thing you want side by side is one column, then map it to fill in a single geom_col().
Q2. You add ylim(20, 40) to a plot with a smooth and the curve changes shape. What happened?
The curve was rescaled to the new axis range.
81 rows were deleted before the smooth was fitted, so it is now fitted to 153 cars.
loess switched to a different span.
A scale limit filters the data. At displ = 5 the two curves read 22.7 and 17.2, a difference of 5.5 mpg from what looks like a cosmetic change. coord_cartesian zooms without deleting anything.
Q3. You have twelve minutes to improve a chart before a presentation. What gives the most back?
Switching theme and tuning the palette.
A title that states the finding, units on both axes, and one highlighted subset with an annotation.
Increasing the resolution of the export.
A theme changes nothing that carries data. A title that states the finding tells the audience what they are looking at, and a highlighted subset tells them where to look. Both are marked under presentation quality and visualisation design.
Section 5

Assessment 3: the walkthrough

Business Decision Case Study. Forty per cent of the subject, and the presentation is next week.
5.1

What it is

ItemDetail
TitleBusiness Decision Case Study
TypeIndividual presentation and report
Weighting40 per cent, split 20 presentation and 20 report
Total marks40
Report length800 words, plus or minus 10 per cent
Presentation length As advised by your facilitator, based on class size
PresentationWeek 12, in class
SubmissionWeek 13, Tuesday 23:59 AEST, via MyKBS
Outcomes LO3, evaluate and apply visualisation techniques.
LO4, design visualisations to support decision-making.

The task in one sentence

Take a real business decision, find data that bears on it, visualise that data in ggplot2, and recommend what to do.

The word that carries the assessment Decision. Not a topic, not a dataset, not an interesting pattern. Somebody has to be choosing between options, and your charts have to help them choose.
One instruction students miss You must not simply reproduce a published graph. If an existing chart is relevant you may estimate the values from it and rebuild it yourself in R, and you should say that is what you did.
5.2

Where the forty marks are

40 marks, five criteria20 per cent presentation, 20 per cent report, 40 per cent of the subjectCase study selection and business relevancea genuine decision, specific and grounded8 marksData visualisation design and technical executionggplot2: scales, labels, themes, annotation8 marksAnalysis, interpretation, and recommendationspatterns read from the charts, then acted on8 marksPresentation and communication qualitystructure, pacing, slides that carry the message6 marksR code documentation and report qualitycommented code and an 800-word report10 marksThe largest single block is code and report, at 10 marks. It is also the one nobody sees you do.
Figure 5.1. The five marking criteria and their weights, from the published rubric.
Analysis
Detail
  • Unit
    One bar is one marking criterion out of 40 total marks.
  • Encoding
    Bar length gives the marks available, on a common scale.
  • Scale
    0 to 10 marks. The subject weighting is 40 per cent, so one mark is worth one per cent of the subject.
  • Structure
    Three criteria carry 8 marks each, one carries 6 and one carries 10.
  • Groups and exceptions
    Code documentation and report quality is the largest single block at 10 marks, which is a quarter of the assessment. Presentation and communication quality is the smallest at 6.
  • Claim and limit
    The largest block of marks is the one nobody watches you earn. It is also the only one you can finish calmly after the presentation is over, in the week between Week 12 and the Week 13 deadline.

What each band actually asks for

CriterionPass looks like Distinction looks like
Case study and business relevance
8 marks
Broadly relevant, but general and only partly connected to a decision Specific, grounded in a credible context, with a clear decision focus
Visualisation design and execution
8 marks
Basic and partly effective, with issues in formatting, labelling or readability Well designed and accurate, with effective labels, scales, themes and annotations
Analysis and recommendations
8 marks
Identifies some patterns; interpretation limited and recommendations thin Thoughtful and well supported, with practical and justified recommendations
Presentation and communication
6 marks
Main points get across, but clarity, structure or pacing suffer; slides basic Well organised, engaging, professionally delivered, slides support the message
Code and report quality
10 marks
Provided, but documentation, structure or referencing is basic or inconsistent Well documented and easy to follow; report clear, concise, structured, referenced
Read the difference between the two columns Almost every jump from pass to distinction is about deliberateness: labels chosen rather than inherited, a recommendation justified rather than stated, code commented rather than saved.
5.3

The five steps, and what each one produces

STEP 1

Choose the decision

A genuine business decision, supported by data that can be analysed visually. Use your own workplace, a previous employer, or an industry you know.

Produces: one sentence naming who decides what
STEP 2

Source and prepare the data

Import it into R. Clean, prepare and transform as needed. Annual reports, government datasets, industry reports and credible websites are all acceptable.

Produces: a data file and the code that reads it
STEP 3

Build the visualisations

ggplot2. Choose plot types, scales, labels, themes and annotations so the charts are accurate, readable and suited to the audience.

Produces: three or four charts, each answering one question
STEP 4

Analyse and recommend

Identify patterns, trends, comparisons and relationships. Explain the business problem and develop practical recommendations that follow from what you drew.

Produces: a recommendation per chart
STEP 5

Present and submit

Present in class in Week 12. Submit the report in Word, your slides, your R code as a .R file, and the data files, by Week 13 Tuesday 23:59.

Produces: four files in MyKBS
Step 1 is the one to finish today Everything else is work you can do. Choosing badly at step 1 caps your marks under the first criterion no matter how good the charts are.
5.4

Three dates

1Today, Week 11Choose the decision.Find the data.Bring both next week.nothing to submit2Week 12In-class presentation.Slides on screen,you at the front.20% of the subject3Week 13, Tue 23:59Report, slides, .R fileand data files,via MyKBS.20% of the subjectThe Week 13 submission is four things, not one. Any alternative file format may not be accepted,so the report is Word, the code is .R, and the data must load straight into that code.
Figure 5.2. What is due when, from the assessment brief.
Analysis
Detail
  • Unit
    One node is one milestone in the assessment.
  • Encoding
    Horizontal position gives time. The label under each node gives what that milestone is worth.
  • Scale
    Three weeks, from today to the submission deadline.
  • Structure
    The presentation comes first and the files come a week later.
  • Groups and exceptions
    Nothing is submitted today. The Week 13 submission is four separate files rather than one document.
  • Claim and limit
    The gap between Week 12 and Week 13 is deliberate. You can improve the report and tidy the code after you have presented, and after you have heard the questions your audience asked.

Working backwards from Week 13

ByYou should have
End of today A decision written as one sentence, and a candidate dataset you have actually opened
This weekend The data loading into R without manual edits, and one chart that says something
Two days before class All charts built and exported. Slides drafted. Rehearsed once against a clock
Week 12Present
Week 13, Tuesday 23:59 Report in Word, slides, code as .R, data files, all uploaded to MyKBS
Format, stated in the brief The report must be in Word. The code must be in .R format, not pasted into the report and not a notebook. The data must be in a format your code can read directly. Any alternative file formats may not be accepted.
The test for the data file Move your whole folder to a different computer, open the .R file, and run it top to bottom. If it fails, the marker will hit the same failure.
5.5

Choosing a case study that can actually be answered

The shape of a good one

Write it as a sentence with three parts: who has to decide, what the options are, and by when.

WorksWhy
"Should the cafe drop its Sunday trading hours next quarter?" Two options, a deadline, and sales data by day and hour would settle it
"Which two of our six product lines should get next year's marketing budget?" A ranking question with a fixed budget, answerable from revenue and margin
"Should the council put the new bus route through the north or the east corridor?" Two named alternatives, with public transport and census data available
"Is our December staffing level too high for the volume we actually get?" A number to compare against a pattern, with an obvious action either way

The shape of one that will cost you marks

StrugglesWhy
"An analysis of the Australian housing market" No decision, no decider, no options. This is a topic
"Exploring the mpg dataset" No business context at all, and we have used it for four weeks
"How can our company improve customer satisfaction?" A decision in principle, but the options are unbounded so no chart can close it
"Predicting next year's sales" A forecasting task, not a decision, and it needs methods this subject has not taught
The test Finish this sentence: "If the charts show X, they should do A; if the charts show Y, they should do B." If you cannot fill in A and B, there is no decision yet.
Scope it small One decision investigated properly scores higher than four investigated thinly. Eight minutes and 800 words do not hold much.
5.6

The charts that earn the marks

You have eleven weeks of techniques. These are the ones that answer business questions, with the week they came from.

If your question isDraw FromWatch for
Has it changed over time? Line chart, geom_line()Week 6 A truncated y axis, and a start date chosen to flatter the trend
Which category is largest? Bar chart, geom_col(), sortedWeeks 1, 10 Alphabetical ordering, and bars that do not start at zero
Do these two move together? Scatter with geom_smooth()Weeks 4, 10 Overplotting, and a pooled trend that reverses within groups
How is it spread? Histogram or violin with jitterWeeks 3, 11 A boxplot hiding two groups, and an inherited bin width
Does it hold in every segment? facet_wrap(), shared scalesWeeks 7, 11 Panels with too few observations to support a claim
How do two categories cross? geom_tile() on an aggregateWeeks 9, 11 Combinations that never appear, silently dropped
Is the comparison fair? Per-unit rates, not totalsWeek 9 The largest total and the best rate are often different groups
Three or four charts, not twelve Each one should answer a question you can state in a sentence. A chart you cannot justify in a sentence is one the marker will read as decoration, and decoration counts against design under the rubric.
5.7

A structure for the talk

Eight minutes, spent deliberatelyConfirm the length with your facilitator, then hold to it1 min1 min3.5 min1.5 min1 minThe decisionWho has to decide what, and by when.Name the decision in your first sentence.The dataWhere it came from, how big, what you cleaned.One slide. Include the row count.The evidenceThree or four charts, one finding each.This is where the marks are.The recommendationWhat you would do, and what it rests on.Tie each recommendation to a chart.Limits and questionsWhat the data cannot settle.Naming a limit earns marks.Three and a half of your eight minutes are charts. Everything else exists to frame them.Rehearse once against a clock. Most people overrun on the data slide.
Figure 5.3. A suggested division of an eight-minute presentation.
Analysis
Detail
  • Unit
    One block is one part of the talk, measured in minutes.
  • Encoding
    Block width gives the time allocated. Total width is the whole talk.
  • Scale
    Eight minutes. Confirm your actual length with your facilitator and rescale proportionally.
  • Structure
    The evidence section is the largest, at three and a half of the eight minutes.
  • Groups and exceptions
    The decision and the data together take two minutes. Limits and questions take one, and that minute earns marks under analysis.
  • Claim and limit
    Less than half the talk is charts, and everything else exists to frame them. This split is a starting point rather than a rule. The thing it is protecting against is a talk that spends four minutes on background and thirty seconds on findings.

Slide by slide

SlideContainsSays
1Title, your name, the decision "Ravenswood Cafe has to decide whether to keep trading on Sundays."
2The data: source, period, rows, what you cleaned "Point of sale exports, 14 months, 31,402 transactions, 212 voided rows removed."
3 to 6One chart each, one finding each "Sunday revenue is 38 per cent of Saturday and falls all year."
7The recommendation, tied to the charts "Close Sundays from June. Slides 4 and 5 are the reason."
8What the data cannot settle "This cannot tell us whether Sunday customers return midweek."
Three things that read as professional State the decision in your first sentence. Put the finding in the chart title rather than the variable names. Name one limitation before anyone asks.
Rehearse against a clock, once Most people overrun on the data slide and then rush the recommendation, which is the part carrying the marks.
5.8

The 800-word report

A budget that fits

SectionWordsJob
The decision100 Who decides, between what, by when, and why it matters
The data120 Source, period, size, and exactly what you cleaned or transformed
The visualisations320 One short paragraph per chart: what it shows, and what you read from it
Recommendations180 What to do, tied to a named figure, with the expected effect
Limitations80 What the data cannot settle, and what you would collect next
Total800 Plus or minus 10 per cent, so 720 to 880

Penalties may be applied for submissions that exceed the prescribed limit.

What 800 words cannot hold

  • A literature review. You are not writing one.
  • A description of what ggplot2 is. Your reader taught it to you.
  • Every chart you made. Include the three or four that carry the argument.
  • A recap of the whole industry. One sentence of context is enough.
Referencing KBS accepts any style provided it is used consistently. Reference your data source, and if you rebuilt a published chart from estimated values, reference the original and say that is what you did.
Each figure needs three things A number, a caption that states the finding rather than the variable names, and a sentence in the text that refers to it by that number. A figure nobody refers to reads as decoration.
5.9

The code, and the submission

Code that a marker can follow
Before you upload

Ten marks cover code and report together. The code half is the easiest place in this subject to pick up marks: it is entirely under your control and needs no new analysis.

# TECH3100 Assessment 3  |  Student name, number
# Ravenswood Cafe: should Sunday trading continue?
# Last run: R 4.3.3 on 2026-10-14
# 1. PACKAGES ------------------------------------------
library(ggplot2)

# 2. IMPORT --------------------------------------------
# Relative path, so the folder runs anywhere
sales <- read.csv("data/pos_export_2025.csv")
nrow(sales)   # 31402 transactions before cleaning

# 3. CLEAN ---------------------------------------------
# Voided sales carry a negative total and must not be
# counted as revenue (212 rows).
sales <- subset(sales, total > 0)
sales$weekday <- weekdays(as.Date(sales$date))

# 4. AGGREGATE (Week 9) --------------------------------
by_day <- aggregate(total ~ weekday, sales, sum)

# 5. FIGURE 1: revenue by weekday ----------------------
# Sorted by value: alphabetical order hides the
# ranking the decision turns on.
ggplot(by_day, aes(reorder(weekday, total), total)) +
  geom_col() + coord_flip() +
  labs(title = "Sunday takes less than half of Saturday",
       x = NULL, y = "Revenue ($)")
ggsave("figures/fig1-weekday.png", width = 10,
       height = 5.6, dpi = 300)
Comments answer "why", not "what" # subset the data tells a marker nothing. # Voided sales carry a negative total tells them you understood your data.

The submission checklist

FileCheck
1Report, .docx 720 to 880 words. Figures numbered and referred to by number. References consistent.
2Slides The deck you actually presented, with the charts at readable size.
3Code, .R Header block, numbered sections, comments explaining why. Not a notebook, not pasted into Word.
4Data files In a format your code reads directly. Relative paths, not C:/Users/...
5GenAI appendix, if used All prompts and responses, plus a reference in the KBS format.
The five-minute test that catches most failures Copy your submission folder somewhere new. Open the .R file in a clean R session. Run it from the first line to the last without touching anything. If it errors, fix it now rather than explaining it later.
Common breakages Absolute file paths. A package you installed months ago and never loaded in the script. A data file you edited by hand in Excel and did not re-export. A figure saved by clicking Export rather than by ggsave().
5.10

Generative AI: Level 2

What Level 2 permits

Use of generative AI is optional for this assessment. You may use it for research and content generation, provided it is appropriately referenced. You do not have to use it.

PermittedRequired if you do
Expanding your understanding of a technique A reference, in the same style as any other source
Idea generation in the research phase An appendix documenting the collaboration
Producing content that enhances the assessment, such as images Every prompt and every response used, in that appendix

Referencing guidance is on the Kaplan Library site, under referencing other sources.

Where students get caught

The appendix is not optional if you used it "All prompts and responses used for the assessment" means all of them, including the ones that did not work. An appendix with three tidy prompts after a week of use is incomplete.
Unapproved use during content generation The brief states that unapproved use during the content generation parts may result in penalties for academic misconduct, up to a mark of zero.
The practical reading If it helped you think, reference it and log it. If it wrote something that appears in your submission, reference it, log it, and be able to explain every line of it. You will be presenting this live, and you will be asked questions.
On generated code specifically Code you cannot explain is code you cannot defend in a Week 12 question. That risk is independent of the integrity policy.
5.11

What loses marks, criterion by criterion

CriterionWhat drops you into the lower bandsThe fix, which is usually small
Case study
8 marks
A topic rather than a decision. Nobody in the story has to choose anything. Name the decider and the two options in your first sentence.
Visualisation
8 marks
Default axis labels reading displ and hwy. Unsorted bars. A truncated axis. A title that names a variable not on the chart. labs() on every figure, reorder() on every bar chart, and read the title against the aes().
Analysis
8 marks
Describing the chart instead of reading it. "Revenue was highest in December" is a description; it is not yet a finding. For each chart write the sentence that starts "so we should".
Presentation
6 marks
Reading the slides. Overrunning. Charts too small to read from the back of the room. Rehearse once against a clock. Export at 10 by 5.6 inches and 300 dpi.
Code and report
10 marks
Uncommented code. Absolute file paths. A report over the word limit. Figures with no numbers and no references in the text. Header block, numbered sections, relative paths, and one full run in a clean session.
The pattern across all five rows Nothing in the right-hand column requires more analysis. Every one of them is a half-hour of care applied to work you have already done, and together they are the difference between a credit and a distinction.
Section 6

Where this leaves you

The end of the taught content, and the start of the week that matters most.
6.1

What today added

The techniques

  1. Four ways to show a distribution, and what each one throws away. Only 7 of 272 eruptions sit inside the middle of a boxplot that looked entirely ordinary.
  2. Bin width as a decision. The same 272 numbers show two groups at 12 bins and 22 at 80.
  3. Binned counts when a scatterplot saturates. 234 cars occupy 126 positions, so 108 marks were hidden.
  4. geom_tile(), which turns a Week 9 aggregate into a chart, and shows the 9 empty cells that a filtered table would have dropped.
  5. Faceting, shared scales by default, and facet_grid keeping the three empty panels that facet_wrap would remove.
  6. Scale limits against coordinate zoom. 81 rows deleted, and a fitted curve moving 5.5 mpg.
  7. Annotation, which is the cheapest thing you can add before a presentation.

The thread running through all eleven weeks

Every technique this subject has taught makes a choice on your behalf if you do not make it yourself.

Week 8Mean imputation chose to shrink your variance
Week 9An inner join chose which records to discard
Week 10Mapping a constant inside aes() chose a colour that was not the one you named
Week 1130 bins, shared scales, and a scale limit that deletes rows
This is what the rubric means by deliberate A distinction is not a chart with more features. It is a chart where you can say why each decision was made, including the ones you left at the default.
6.2

Before next week

TONIGHT

Write the sentence

"Who has to decide what, by when." If you cannot write it, you do not have a case study yet, and that is the first eight marks.

Five minutes
THIS WEEK

Get the data into R

Not into Excel. Into R, with read.csv() and a relative path, running from a script you can hand in.

This is where projects stall
THIS WEEK

Build one chart that says something

One. With a title stating the finding and both axes labelled. If you can build one you can build four.

Then write the "so we should" sentence
BEFORE CLASS

Rehearse once, against a clock

Out loud, standing, with the slides on screen. Confirm your length with me first. Most overruns happen on the data slide.

Worth six marks directly
BEFORE CLASS

Run the code clean

New session, new folder, top to bottom, no manual steps. Do it now rather than at 23:00 on the Tuesday of Week 13.

Part of the ten-mark criterion
BRING

Questions

About your data, your decision, or which chart to use. Bring the actual file. Ten minutes with it open beats an hour describing it.

Next week is your presentation

See you next week

TECH3100 Lesson 11 · ggplot part 2, and Assessment 3
Week 12: individual presentations, in class
Week 13: report, slides, .R code and data files, Tuesday 23:59 AEST via MyKBS

Press T for the table of contents. Arrow keys to navigate.

Contents

0.1Copyright notice
0.2Two halves today
0.3What you should be able to do by the end
0.4Setup and the datasets
1Showing a distribution
1.1Four views of one variable
1.2What the box threw away
1.3Bin width is a decision
1.4Which of the four to draw
1.QKnowledge check: Section 1
2Two variables at once
2.1Counting instead of drawing
2.2A Week 9 table, straight onto a chart
2.3Three labels that do not match their charts
2.QKnowledge check: Section 2
3Faceting
3.1One panel per group
3.2Shared scales against free ones
3.3A matrix of panels
3.4When to facet, and when not to
3.QKnowledge check: Section 3
4Finishing a figure
4.1The clustered bar chart, done properly
4.2A summary needs its distribution
4.3A limit is not a zoom
4.4Telling the reader where to look
4.5From working to presentable
4.QKnowledge check: Section 4
5Assessment 3: the walkthrough
5.1What it is
5.2Where the forty marks are
5.3The five steps, and what each one produces
5.4Three dates
5.5Choosing a case study that can actually be answered
5.6The charts that earn the marks
5.7A structure for the talk
5.8The 800-word report
5.9The code, and the submission
5.10Generative AI: Level 2
5.11What loses marks, criterion by criterion
6Where this leaves you
6.1What today added
6.2Before next week

Press T or Escape to close. Arrow keys to navigate.