TECH3100 · Data Visualisation in R

Data visualisation with ggplot

Part 1: the grammar of graphics
Lesson 10
Kaplan Business School Australia
All figures and numbers computed in R 4.3.3 with ggplot2 3.4.4
0.1

0.1 Copyright notice

Commonwealth of Australia · Copyright Regulations 1969

This material has been reproduced and communicated to you by or on behalf of Kaplan Business School pursuant to Part VB of the Copyright Act 1968 (the Act).

The material in this communication may be subject to copyright under the Act. Any further reproduction or communication of this material by you may be the subject of copyright protection under the Act.

Do not remove this notice.

0.2

0.2 Where this week sits

WeekTopic
1–4Variable types. Descriptive statistics. Histograms, boxplots, scatter plots, correlation.
5–7Trend lines. Time series. Multivariate data.
8Missing values: imputation and interpolation.
9Aggregates and joining tables.
10Data visualisation with ggplot, part 1.
11Data visualisation with ggplot, part 2.
12Individual presentations.
Picking up from last week Week 9 ended on the point that ggplot maps columns to things you can see. It cannot map something that is not a column. Every join and every aggregate you built last week existed to put the variable you want to show into a column of its own.
What does not change Everything from Weeks 1 to 9 still holds. Choosing the right chart, reading it honestly, reporting the sample size, naming what was dropped. ggplot changes how you draw, and changes nothing about what makes a chart correct.
0.3

0.3 What you should be able to do by the end

  1. Name the layers of a ggplot and say which three are compulsory.
  2. Write a plot as data, then an aesthetic mapping, then a geom, and add layers with +.
  3. Say what belongs inside aes() and what belongs outside it, and predict what happens when you get it wrong.
  4. Predict whether a mapping placed on the canvas or inside a geom will change the result.
  5. Explain why geom_bar() needs no y, and when to use stat = "identity" instead.
  6. Label, scale, theme and save a plot so that somebody else can read it.
The one that matters most Outcome 3. One pair of brackets decides whether a value is treated as data or as an instruction, and ggplot obeys both readings without complaint.
0.4

0.4 Setup and the datasets

Installing once

install.packages("ggplot2")   # once per machine
library(ggplot2)              # once per session

packageVersion("ggplot2")
#> '3.4.4'
Two names for the same thing The package is called ggplot2. The function is called ggplot, with no 2. library(ggplot) is an error that catches everybody once.

British and American spelling

ggplot accepts colour and color everywhere, and treats them as the same argument. Pick one and stay with it.

mpg, which ships with ggplot2

234 cars, 11 columns. Fuel economy from the US Environmental Protection Agency.

ColumnTypeRange or values
displnumeric1.6 to 7.0 litres
hwyinteger12 to 44 mpg
ctyintegercity mpg
drvcharacter4, f, r
classcharacter7 values
cylinteger4, 5, 6, 8

We also use airquality from Week 8, so the missing values come back.

1
Section 1

The grammar of graphics

A chart is not a picture you draw. It is a sentence you write, with a fixed set of parts. Learn the parts and every chart becomes the same job.

1.1

1.1 A chart in seven layers

The two letters in ggplot stand for grammar of graphics. A grammar is a small set of parts that combine in rules, and this is the whole set.

1. Datawhich data frameggplot(data = mpg)2. Aesthetic mappingwhich column goes on which channelaes(x = displ, y = hwy, colour = drv)3. Geometrywhich shape draws the marks+ geom_point()4. Statisticwhat the geom computes first+ geom_smooth()5. Scaleshow values become positions and colours+ scale_colour_brewer()6. Labelswhat the reader is told+ labs(title = ...)7. Themeeverything that carries no data+ theme_minimal()Every ggplot is these seven layers. Layers 1 to 3 are compulsory. The other four have defaults.
1.2

1.2 Anatomy of a call

Every ggplot you write has these pieces in this order, and each one has a job.

ggplot(data = mpg, aes(x = displ, y = hwy)) + geom_point(colour = "steelblue") + labs(x = "Engine size (L)", y = "Highway mpg") + theme_minimal()the data frameone row per mark on the pagethe mappingcolumn names, never quotedthe geometrythe shape that gets drawna fixed settinga value, outside aes()the labelswhat the reader is toldthe themeno data, appearance onlyRead it top to bottom.Each + adds one thing toan object you already have.Nothing is drawn until youprint the object.ggplot(mpg) creates an object with 0 layers. Adding geom_point() makes it 1 layer.Adding geom_smooth() makes it 2. The object is a recipe, and printing it runs the recipe.
1.3

1.3 What aes() actually is

An aesthetic mapping is a wiring diagram. It says which column travels down which visual channel. Week 7 called these channels; ggplot calls them aesthetics.

Columns in mpgdisplnumeric, 1.6 to 7.0hwyinteger, 12 to 44drvcharacter, 4, f, rclasscharacter, 7 valuescylinteger, 4, 5, 6, 8Visual channelsx positionread most accuratelyy positionread most accuratelycolournominal or continuousshapenominal, 6 values maximumsizecontinuous, judged roughlypanel (facet)nominal, any numberaes() is that arrow.It says which columntravels down which wire.Nothing about the arrow decides what the marks look like. That is the geom's job.
1.4

1.4 Building the object

A ggplot is an object. You can build it in pieces, keep it in a variable, and add to it later. Nothing appears on screen until you print it.

viz <- ggplot(data = mpg)
class(viz)              #> "gg" "ggplot"
length(viz$layers)      #> 0     an empty canvas

viz <- viz + aes(x = displ, y = hwy)
viz <- viz + geom_point()
length(viz$layers)      #> 1

viz <- viz + geom_smooth(method = "lm")
length(viz$layers)      #> 2

viz                     # this line draws it
print(viz)              # the same thing, explicitly
Why this matters Inside a script or a loop, typing the object name on its own does nothing. You have to call print(). That is the most common reason a plot fails to appear.

In practice you write it as one expression. The + goes at the end of the line, never at the start of the next one.

ggplot(data = mpg, aes(x = displ, y = hwy)) +
  geom_point(colour = "steelblue") +
  geom_smooth(method = "lm", formula = y ~ x) +
  labs(x = "Engine displacement (litres)",
       y = "Highway economy (mpg)") +
  theme_minimal()

# This does NOT work:
ggplot(mpg, aes(displ, hwy))
  + geom_point()
#> Error: Cannot use `+` with a single argument

# R finished the first line, so the second line
# starts a new expression with nothing on its left.
Positional arguments ggplot(mpg, aes(displ, hwy)) means the same as ggplot(data = mpg, aes(x = displ, y = hwy)), because data is the first argument and x and y are the first two aesthetics. Write the names out while you are learning.
1.5

1.5 Inside aes() and outside it

Inside aes()geom_point(aes(colour = "darkred"))246203040displhwycolourdarkredOutside aes()geom_point(colour = "darkred")246203040displhwyThe points on the left are not dark red. They are #F8766D, which is the first colour ggplothands to any group of size one. Dark red is #8B0000.Inside aes(), "darkred" is treated as data. ggplot invents a one-level variable, assigns it a colour,and builds a legend for it. Outside aes(), the string is an instruction and is obeyed.
Figure 1.5. The same word, "darkred", placed on each side of the aes() bracket. Both plots were produced in ggplot2 3.4.4 and the colours read back with ggplot_build().
  1. UnitOne point is one car. Both panels show the same 234 rows.
  2. EncodingPosition gives engine size and highway economy in both panels. The only difference between them is where the word "darkred" sits.
  3. ScaleIdentical axes, so any visible difference comes from the code.
  4. StructureBoth panels show the same falling relationship, because the data and the geoms are identical.
  5. Groups and exceptionsThe left panel is drawn in #F8766D, which is salmon. It also grew a legend with one entry. The right panel is drawn in #8B0000, which is dark red, and has no legend.
  6. Claim and limitInside aes(), a value is treated as data. Outside aes(), it is treated as an instruction. On the left ggplot invented a variable whose only value is the string "darkred", assigned it the first colour from its palette, and built a key to explain itself. Nothing warned you.
How to spot it A legend you did not ask for, with one entry, whose label is the value you typed. That legend is the symptom every time.
# WRONG: the string is treated as data
ggplot(airquality, aes(Ozone, Temp)) +
  geom_point(aes(colour = "darkred"))

# RIGHT: the string is treated as an instruction
ggplot(airquality, aes(Ozone, Temp)) +
  geom_point(colour = "darkred")

# Proof, read straight out of the plot object
p <- ggplot(airquality, aes(Ozone, Temp)) +
       geom_point(aes(colour = "darkred"))
unique(ggplot_build(p)$data[[1]]$colour)
#> "#F8766D"

col2rgb("darkred")
#>       [,1]
#> red    139     = #8B0000
#> green    0
#> blue     0

# The rule, in one line:
#   aes(x = column)    maps a column
#   argument = value   sets a constant
1.6

1.6 Where a mapping is placed

Mapping on the canvasggplot(mpg, aes(displ, hwy, colour = drv)) + geom_point() + geom_smooth(method = "lm")246203040displhwy3 fitted linesMapping on the geomggplot(mpg, aes(displ, hwy)) + geom_point(aes(colour = drv)) + geom_smooth(method = "lm")246203040displhwy1 fitted linedrv = 4drv = fdrv = rA mapping placed on the canvas is inherited by every layer. A mapping placed inside a geom stays there.Moving six characters changes the answer from three slopes to one. Neither version warns you.
Figure 1.6. The same aes(colour = drv), moved from the canvas into geom_point(). Both plots add the same geom_smooth() line afterwards.
  1. UnitOne point is one car; one line is a linear fit over whatever rows ggplot considers to be one group.
  2. EncodingPosition gives engine size and economy. Colour gives drive type in both panels, so the points look identical.
  3. ScaleIdentical axes and identical colours, taken from ggplot's three-colour default palette.
  4. StructureEconomy falls as engine size rises. Both panels agree about that.
  5. Groups and exceptionsThe left panel draws three fitted lines, one per drive type. The right panel draws one line over all 234 cars. The rear-wheel slope is −0.92 against −3.60 for front-wheel, so the three lines are saying something the single line hides.
  6. Claim and limitA mapping on the canvas is inherited by every layer; a mapping inside a geom stays in that geom. Moving six characters changes the answer from three slopes to one, and both versions run without a warning.
# colour on the CANVAS: every layer inherits it
ggplot(mpg, aes(displ, hwy, colour = drv)) +
  geom_point() +
  geom_smooth(method = "lm", formula = y ~ x)
#> 3 fitted lines

# colour on the GEOM: only that layer sees it
ggplot(mpg, aes(displ, hwy)) +
  geom_point(aes(colour = drv)) +
  geom_smooth(method = "lm", formula = y ~ x)
#> 1 fitted line

# the slopes the three lines are reporting
for (d in c("4", "f", "r")) {
  f <- lm(hwy ~ displ, subset(mpg, drv == d))
  cat(d, round(coef(f)[2], 2), "\n")
}
#> 4  -2.88
#> f  -3.60
#> r  -0.92
coef(lm(hwy ~ displ, mpg))[2]   #> -3.53

# Rule of thumb: put it on the canvas when every
# layer should use it, and in the geom when only
# one layer should.
1.7

1.7 The column's type chooses the scale

Month is numericaes(colour = Month)607590050100150TempOzoneMonth975Month made a factoraes(colour = factor(Month))607590050100150TempOzonefactor(Month)56789Left: a continuous scale, five shades of one blue, with a colour bar. Right: a discrete scale,five separate hues, with a key. Same column, same data, same aes() call.The column's type chose the scale. Month is stored as a number, so ggplot treated May toSeptember as a quantity. Wrapping it in factor() tells ggplot it is a label.Five months, one line of difference.
Figure 1.7. airquality coloured by Month. The only difference between the panels is factor().
  1. UnitOne point is one day. Both panels show the same 153 days.
  2. EncodingPosition gives temperature and ozone. Colour gives month in both panels.
  3. ScaleThe left panel uses a continuous scale running from #132B43 to #56B1F7. The right panel uses five separate hues from ggplot's discrete palette.
  4. StructureOzone rises with temperature in both panels, which is the finding from Week 8.
  5. Groups and exceptionsOn the left the legend is a colour bar with a numeric axis, which invites the reader to interpolate between months. On the right it is a key with five entries.
  6. Claim and limitggplot chose the scale from the column's storage type, not from what the column means. Month is stored as a number, so ggplot treated May to September as a quantity, and the colour bar implies a month 6.5 exists. factor() is how you say the numbers are labels.
# Month is stored as a number
class(airquality$Month)     #> "integer"

# so ggplot builds a CONTINUOUS colour scale
ggplot(airquality, aes(Temp, Ozone, colour = Month)) +
  geom_point()

# factor() says: these numbers are labels
ggplot(airquality,
       aes(Temp, Ozone, colour = factor(Month))) +
  geom_point()

# The same question comes up for x and y:
#   a numeric x gives a continuous axis
#   a factor  x gives one tick per level

# Fix it in the data rather than in every plot
aq <- airquality
aq$Month <- factor(aq$Month,
                   labels = c("May","Jun","Jul","Aug","Sep"))

# Now the legend reads in English and the order
# is the one you chose, not alphabetical.
1.8

1.8 The same chart, twice

Base R246203040displhwyggplot2246203040displhwydrv4frcols <- c("4"="#F8766D", "f"="#00BA38", "r"="#619CFF")plot(mpg$displ, mpg$hwy, pch = 19, col = cols[mpg$drv])legend("topright", names(cols), pch = 19, col = cols)ggplot(mpg, aes(displ, hwy, colour = drv)) + geom_point()Both draw the same 234 points.The legend, the palette and the scale are work you do once in base R and never in ggplot.
Figure 1.8. 234 cars coloured by drive type, drawn in base R and in ggplot2.
  1. UnitOne point is one car in both panels.
  2. EncodingPosition gives engine size and economy; colour gives drive type. The two panels use the same three hex colours.
  3. ScaleBase R defaults to a white panel with a box. ggplot defaults to a grey panel with white gridlines. Neither default carries information.
  4. StructureIdentical, because the data and the encoding are identical.
  5. Groups and exceptionsThe legend on the right was not requested. ggplot built it because a column was mapped to colour, and a mapped column always gets a key.
  6. Claim and limitEverything ggplot does for free in this figure is work you write by hand in base R: the palette, the lookup, the legend. The cost is that ggplot has opinions, and unlearning a default takes longer than setting one.
# base R: you build the palette and the legend
cols <- c("4" = "#F8766D",
          "f" = "#00BA38",
          "r" = "#619CFF")
plot(mpg$displ, mpg$hwy, pch = 19,
     col = cols[mpg$drv],
     xlab = "displ", ylab = "hwy")
legend("topright", legend = names(cols),
       pch = 19, col = cols, bty = "n")

# ggplot: you declare the mapping and stop
ggplot(mpg, aes(displ, hwy, colour = drv)) +
  geom_point()

# What ggplot did without being asked:
#   picked 3 evenly spaced hues
#   matched each drv value to one
#   drew a legend, titled it from the column name
#   placed it, sized it, spaced it

# In base R each of those is a line you maintain.
1.Q

1.Q Knowledge check: Section 1

Q1. geom_point(aes(colour = "blue")) runs without an error. What colour are the points?
Inside aes(), the string becomes a one-level variable. ggplot assigns that single level the first colour in its palette and adds a legend to explain it. The word you typed ends up as the legend label rather than as the colour.
Q2. ggplot(mpg, aes(displ, hwy, colour = drv)) + geom_point() + geom_smooth() draws how many smooth lines?
A mapping on the canvas is inherited. geom_smooth sees colour = drv, splits the data into three groups and fits each one separately. Moving colour inside geom_point() gives a single line.
Q3. aes(colour = Month) on airquality produces a colour bar rather than a key with five entries. Why?
ggplot reads the column's storage type, not its meaning. An integer column gets a continuous scale, which implies values between the months exist. factor(Month) tells ggplot the numbers are labels.
2
Section 2

Points, lines and smooths

The geoms you will use most, and the two things that go wrong on every scatterplot: marks hidden underneath other marks, and rows quietly removed before drawing.

2.1

2.1 geom_point and its aesthetics

# the minimum
ggplot(mpg, aes(displ, hwy)) + geom_point()

# mapped: the value comes from a column
geom_point(aes(colour = drv))     # 3 hues
geom_point(aes(shape  = drv))     # 3 shapes
geom_point(aes(size   = cyl))     # area scale
geom_point(aes(alpha  = cty))     # transparency

# set: the value is fixed for every point
geom_point(colour = "steelblue")
geom_point(size   = 3)
geom_point(shape  = 17)           # solid triangle
geom_point(alpha  = 0.3)

# both at once is fine
geom_point(aes(colour = drv), size = 3, alpha = .7)

Which aesthetic suits which variable

AestheticSuitsLimit
x, yanythingread most accurately
colournominal or continuous6 to 8 categories
shapenominal only6 by default, then ggplot refuses
sizecontinuous onlyjudged roughly, so use it for context
alphacontinuous, or densityweakest of the five
This is Week 7 The ranking is the one from the multivariate lesson: position first, then length, then angle, area, colour value, colour hue. ggplot gave the channels new names and changed nothing about how accurately people read them.

Mapping the same column to colour and shape at once is a reasonable choice for a printed handout, because it survives being photocopied.

2.2

2.2 Adding a smooth

60708090100050100150Temp (degrees F)Ozone (ppb)geom_smooth() default: loessmethod = "lm"grey band: the standard error of the fit, drawn by defaultggplot chose loess because n is under 1000.The straight line is Ozone = −147.00 + 2.43 × Temp. It explains 48.8 per cent of thevariance and misses the bend above 85 degrees.
Figure 2.2. airquality ozone against temperature, with the default smooth and a straight-line fit on the same axes.
  1. UnitOne point is one day, from the 116 days with a recorded ozone value.
  2. EncodingPosition gives temperature and ozone. The blue curve is the fitted value; the grey band is its standard error.
  3. ScaleOzone in parts per billion against temperature in degrees Fahrenheit. The vertical axis has been extended below zero so the whole band fits.
  4. StructureOzone rises with temperature, and the rise is steeper above 80 degrees than below it.
  5. Groups and exceptionsThe two fits disagree at both ends. The straight line sits above the data below 70 degrees and below it above 90. The grey band widens at both edges, where there are fewer points.
  6. Claim and limitThe relationship bends, and a straight line misses the bend. The straight line explains 48.8 per cent of the variance. The default smooth is loess, which is a local average rather than a model, so it cannot be reported as an equation and it should not be extrapolated.
What the straight line says $$\widehat{\text{Ozone}} = -147.00 + 2.43 \times \text{Temp}$$

Every extra degree is associated with 2.43 more parts per billion, on average, across this range. The intercept has no meaning here, because no day in the data is near zero degrees.

ggplot(airquality, aes(Temp, Ozone)) +
  geom_point() +
  geom_smooth()
#> `geom_smooth()` using method = 'loess'
#>  and formula = 'y ~ x'

# ggplot picks loess when n < 1000 and gam above it.
# Say which you used, because the default changes
# with the size of your data.

# a straight line instead
ggplot(airquality, aes(Temp, Ozone)) +
  geom_point() +
  geom_smooth(method = "lm", formula = y ~ x)

# turn the band off when you do not want to
# discuss it
geom_smooth(method = "lm", formula = y ~ x,
            se = FALSE)

# the same fit as a number
f <- lm(Ozone ~ Temp, data = airquality)
round(coef(f), 2)
#> (Intercept)        Temp
#>     -147.00        2.43
round(summary(f)$r.squared, 4)   #> 0.4877
2.3

2.3 Marks hidden under other marks

As drawngeom_point()246203040displhwyTransparencygeom_point(alpha = 0.25)246203040displhwyJittergeom_jitter(width = .1, height = .6)246203040displhwympg has 234 rows and only 126 distinct (displ, hwy) pairs.108 points sit exactly underneath another point in the left panel.Nearly half the data is invisible in a plot that looks complete.Transparency shows where the density is. Jitter shows how many marks are really there.
Figure 2.3. The same 234 cars drawn three ways.
  1. UnitOne mark is one car in all three panels.
  2. EncodingPosition gives engine size and economy. The middle panel adds transparency; the right panel adds a small random offset to each point.
  3. ScaleIdentical axes across all three, so the differences are the drawing, not the data.
  4. StructureAll three show economy falling as engine size rises.
  5. Groups and exceptionsmpg has 234 rows and only 126 distinct pairs of displ and hwy. 108 marks in the left panel sit exactly underneath another mark.
  6. Claim and limitNearly half the data is invisible in a scatterplot that looks complete. Transparency shows where the density is; jitter shows how many marks are really there. Jitter moves the points, so the reader is no longer looking at exact values, and that has to be said in the caption.
# how bad is it?
nrow(mpg)                                #> 234
nrow(unique(mpg[, c("displ", "hwy")]))   #> 126
# 108 points are sitting on top of another

# 1. transparency: overlap becomes darkness
ggplot(mpg, aes(displ, hwy)) +
  geom_point(alpha = 0.25)

# 2. jitter: a small random offset
ggplot(mpg, aes(displ, hwy)) +
  geom_jitter(width = 0.1, height = 0.6)

# geom_jitter(...) is short for
# geom_point(position = position_jitter(...))

# 3. count: one mark per pair, sized by how many
ggplot(mpg, aes(displ, hwy)) + geom_count()

# 4. for very large n, bin instead of plotting
ggplot(diamonds, aes(carat, price)) +
  geom_hex(bins = 40)

# Try transparency first. Jitter only when the
# exact values do not matter to the reader.
2.4

2.4 Rows removed before drawing

Warning message:Removed 37 rows containing missing values (`geom_point()`).153 rows of airquality, one block per day37 dropped116 drawnggplot removes incomplete rows for you and tells you how many. The chart still appears.Every statistic on that chart now rests on 116 days, not 153. Week 8 applies to the plot as wellas to the table.
Figure 2.4. Plotting airquality ozone. The red blocks are the days ggplot removed.
  1. UnitOne block is one day of the 153 in the dataset.
  2. EncodingRed marks a day that was dropped; pale blue marks a day that was drawn.
  3. ScaleAll 153 days in calendar order, reading left to right and top to bottom.
  4. StructureThe dropped days are the ones with no ozone reading, which is Week 8's missingness pattern.
  5. Groups and exceptions37 of 153 days were removed, which is 24.2 per cent. The block in the second row is the June run.
  6. Claim and limitThe chart is drawn from 116 days, not 153, and the only evidence is a warning in the console. Anyone reading the figure later sees a complete-looking scatterplot. The n belongs in the caption.
ggplot(airquality, aes(Temp, Ozone)) +
  geom_point()
#> Warning message:
#> Removed 37 rows containing missing values
#> (`geom_point()`).

sum(is.na(airquality$Ozone))     #> 37
nrow(airquality)                 #> 153

# Decide what to do, then say what you did:

# (a) keep the warning and report the n
labs(caption = "116 of 153 days had an ozone reading")

# (b) drop them on purpose, so the count is yours
aq <- subset(airquality, !is.na(Ozone))
nrow(aq)                          #> 116

# (c) show where they are
aq2 <- airquality
aq2$missing <- is.na(aq2$Ozone)
ggplot(aq2, aes(Temp, Ozone, colour = missing)) +
  geom_point()

# Silencing the warning with na.rm = TRUE hides
# the only signal you were given. Do not.
2.5

2.5 The geoms you will use this semester

GeomDraws NeedsWeek it replaces
geom_point()one mark per rowx and y Week 4 scatter plots
geom_line()points joined in x orderx and y Week 6 time series
geom_smooth()a fitted curve and its bandx and y Week 5 trend lines
geom_bar()one bar per category, countedx only Week 1 bar charts
geom_col()one bar per row, at the height you givex and y Week 9 aggregate tables
geom_histogram()counts in binsx only Week 3 histograms
geom_boxplot()five-number summary per groupx and y Week 3 boxplots
geom_tile()a filled rectangle per cellx, y and fill Week 7 heat maps
The pattern Every chart from the first nine weeks has a geom. You are not learning new charts this week. You are learning one way to write all of them.
2.Q

2.Q Knowledge check: Section 2

Q1. A scatterplot of 234 cars looks sparse, and nrow(unique(...)) returns 126. What does that tell you?
The rows are all there and all drawn. 108 of them land on coordinates already occupied, so they are invisible. The rows themselves differ in other columns; only the two plotted columns repeat.
Q2. geom_smooth() prints a message naming loess. Why should you state the method in your caption?
The default is chosen from the size of your data. Code that produced loess on a sample can produce gam on the full dataset, and the two curves differ. Naming the method makes the figure reproducible.
Q3. Your plot prints "Removed 37 rows containing missing values". What is the correct response?
na.rm = TRUE removes the message and changes nothing else, which leaves you with a chart of 116 days that looks like a chart of 153. The number belongs in the caption either way.
3
Section 3

Bars and counts

Bar charts are where ggplot does the most work on your behalf, and where beginners lose the most time. One geom does two different jobs.

3.1

3.1 geom_bar computes its own y

What you give itmpg: 234 rows, no count column anywheremanufacturermodeldisplclassaudia41.8compactaudia42.0compactchevroletc15005.3suvdodgecaravan3.3minivan…………stat_count()runs before drawingand makes a new columnWhat it computes7 rowsclasscountsuv62compact47midsize41subcompac35pickup33minivan112seater5020406052seater47compact41midsize11minivan33pickup35subcompa62suvcountgeom_bar()aes(x = class)There is no y inthe aes() call.geom_bar() computes its own y. That is why the aes() call names only x, and why the y axis says count.The 7 numbers on the right exist nowhere in mpg. The plot made them.
Figure 3.1. mpg has no count column. geom_bar makes one before it draws anything.
  1. UnitOne bar is one value of class. Seven bars for seven classes.
  2. EncodingHorizontal position gives the class. Bar height gives the number of rows in mpg with that class.
  3. ScaleThe vertical axis runs from zero and is labelled count, which is a name ggplot invented.
  4. StructureThe seven counts sum to 234, which is every row of mpg.
  5. Groups and exceptionssuv is the largest at 62 and 2seater the smallest at 5. Neither number appears anywhere in the mpg data frame.
  6. Claim and limitgeom_bar runs a statistic before it draws. That statistic is stat_count, and it is why the aes() call names only x. The bar heights are computed by the plot, so they cannot be checked against your data unless you compute them separately.
# no y anywhere in this call
ggplot(mpg, aes(x = class)) + geom_bar()

# what happened underneath
p <- ggplot(mpg, aes(x = class)) + geom_bar()
ggplot_build(p)$data[[1]][, c("x", "y", "count")]
#>   x  y count
#> 1 1  5     5
#> 2 2 47    47
#> 3 3 41    41
#> 4 4 11    11
#> 5 5 33    33
#> 6 6 35    35
#> 7 7 62    62

sum(table(mpg$class))    #> 234

# the same numbers, computed yourself
table(mpg$class)

# geom_histogram does the same thing for a
# continuous x: it bins, then counts
ggplot(mpg, aes(x = hwy)) +
  geom_histogram(bins = 20)

# Always set bins or binwidth. The default of 30
# is a placeholder and ggplot says so.
3.2

3.2 When you already have the numbers

Week 9 output, used directly as Week 10 inputclasshwycompact28.30subcompact28.14midsize27.292seater24.80minivan22.36suv18.13pickup16.88aggregate(hwy ~ class, mpg, mean)010203028.3compact28.1subcompact27.3midsize24.82seater22.4minivan18.1suv16.9pickupmean hwyThe bars are the numbers in the table. No counting happened.stat = "identity" tells geom_bar to use the y you supplied instead of computing one.Two different jobs, one geom. Count rows, or draw numbers you already have.Every aggregate table from Week 9 goes straight into the second form.
Figure 3.2. A Week 9 aggregate table plotted directly, with no counting.
  1. UnitOne bar is one class. The seven rows of the aggregate table become seven bars.
  2. EncodingBar length gives the mean highway economy for that class. The bars are horizontal because the class names are long.
  3. ScaleMiles per gallon, starting at zero, which is correct for a bar chart because bar length is being read as a quantity.
  4. StructureEconomy falls from compact and subcompact down to pickup.
  5. Groups and exceptionscompact averages 28.30 and pickup 16.88, a difference of about eleven and a half miles per gallon.
  6. Claim and limitThe bars are the numbers you supplied, so nothing was computed by the plot. Each bar is a mean over between 5 and 62 cars, and the chart does not show those counts. A mean over 5 cars and a mean over 62 sit side by side looking equally solid.
The rule Counting rows: geom_bar(). Drawing numbers you already have: geom_col(). If you find yourself writing stat = "identity", you wanted geom_col().
# Week 9 gives you the table
ag <- aggregate(hwy ~ class, data = mpg, FUN = mean)
ag
#>        class      hwy
#> 1    2seater 24.80000
#> 2    compact 28.29787
#> ...

# two ways to draw it, both identical
ggplot(ag, aes(x = class, y = hwy)) +
  geom_bar(stat = "identity")

ggplot(ag, aes(x = class, y = hwy)) +
  geom_col()

# geom_col() IS geom_bar(stat = "identity").
# Prefer geom_col: the name says what it does.

# long labels: flip the chart, do not rotate text
ggplot(ag, aes(x = reorder(class, hwy), y = hwy)) +
  geom_col() +
  coord_flip() +
  labs(x = NULL, y = "Mean highway mpg")

# Carry the count through so it can be reported
ag$n <- as.vector(table(mpg$class))
3.3

3.3 Where the bars sit

position = "stack"the default02040602secommidminpicsubsuvposition = "dodge"side by side020402secommidminpicsubsuvposition = "fill"proportions00.512secommidminpicsubsuvdrv = 4drv = fdrv = rTwelve bars in all three panels. The tallest reaches 62, then 51, then exactly 1.The data never changed. Only the rule for where each bar starts changed.Stack answers ‘how many in total’. Dodge answers ‘how do the groups compare’.Fill answers ‘what share’, and throws the totals away to do it.
Figure 3.3. Car class by drive type, drawn with the three position adjustments.
  1. UnitOne rectangle is one combination of class and drive type. There are twelve non-empty combinations, and all three panels draw all twelve.
  2. EncodingHorizontal position gives class, fill gives drive type, and the vertical extent gives the count. Only the rule for where each rectangle starts changes.
  3. ScaleStack reaches 62, dodge reaches 51, fill reaches exactly 1. The third panel is a proportion, so its axis is unitless.
  4. StructureEvery class is dominated by one or two drive types. No class uses all three evenly.
  5. Groups and exceptionssuv is the tallest bar under stack at 62, and its 4wd segment is the tallest single bar under dodge at 51. Under fill it looks the same height as every other class.
  6. Claim and limitThe three panels answer three different questions from identical data. Fill discards the totals to show shares, so a class with 5 cars and a class with 62 look equally important. Say which question you are answering.
# stack: the default. Totals are readable,
# individual segments above the first are not.
ggplot(mpg, aes(class, fill = drv)) +
  geom_bar(position = "stack")

# dodge: side by side. Segments are comparable,
# totals are gone.
ggplot(mpg, aes(class, fill = drv)) +
  geom_bar(position = "dodge")

# fill: proportions. Shares are comparable,
# totals are gone entirely.
ggplot(mpg, aes(class, fill = drv)) +
  geom_bar(position = "fill") +
  labs(y = "Share of class")

# the same three counts, as a table
table(mpg$class, mpg$drv)

# Which to choose:
#   "how many altogether"      -> stack
#   "how do the groups compare"-> dodge
#   "what share of each"       -> fill, and
#      print the totals somewhere else
3.4

3.4 Putting the bars in an order

Default orderaes(x = class)020406052seater47compact41midsize11minivan33pickup35subcompa62suvOrdered by countaes(x = reorder(class, -table(class)[class]))020406062suv47compact41midsize35subcompa33pickup11minivan52seaterLeft: alphabetical, because class is a character column and R sorts characters.Right: the same seven bars, in an order that answers a question.A reader can find the largest category instantly on the right and has to search for it on the left.
Figure 3.4. The same seven counts, in the default order and in an order that answers a question.
  1. UnitOne bar is one car class, in both panels.
  2. EncodingBar height gives the count. The only difference between the panels is the order along the horizontal axis.
  3. ScaleIdentical vertical scales, so heights are directly comparable across panels.
  4. StructureThe same seven numbers appear in both panels: 5, 47, 41, 11, 33, 35 and 62.
  5. Groups and exceptionsIn the left panel the largest bar sits sixth from the left, because suv comes late in the alphabet. In the right panel it is first.
  6. Claim and limitAlphabetical order is an accident of how the column was stored. A reader can find the largest class instantly on the right and has to search for it on the left. Ordering is a choice, and the default is somebody else's.
# default: R sorts character columns alphabetically
ggplot(mpg, aes(x = class)) + geom_bar()

# reorder by how many rows are in each level
ggplot(mpg, aes(x = reorder(class, class, length))) +
  geom_bar()

levels(reorder(mpg$class, mpg$class, length))
#> "2seater" "minivan" "pickup" "subcompact"
#> "midsize" "compact" "suv"          ascending

# descending: negate the sorting value
ag <- aggregate(hwy ~ class, mpg, mean)
ggplot(ag, aes(reorder(class, -hwy), hwy)) +
  geom_col()

# an order you choose by hand
mpg$class <- factor(mpg$class,
  levels = c("2seater", "subcompact", "compact",
             "midsize", "minivan", "suv", "pickup"))

# Order by size for a ranking. Keep a natural
# order (Mon to Sun, May to Sep) when one exists.
# Alphabetical needs a reason.
3.5

3.5 A bar chart checklist

CHECK 01

Which job

Are you counting rows, or drawing numbers you already computed? The first is geom_bar(), the second is geom_col().

No aes y means counting
CHECK 02

Zero on the axis

Bar length is read as a quantity, so the axis has to start at zero. ggplot does this by default and you can override it. Do not.

Truncating a bar axis misleads
CHECK 03

Order

Alphabetical only when the reader will look a category up by name. Otherwise order by size, or by a natural sequence.

reorder() is one function call
CHECK 04

Position

Stack for totals, dodge for comparison, fill for shares. Fill throws the totals away, so print them somewhere else.

Same data, three questions
CHECK 05

Labels that fit

Long category names go on the vertical axis with coord_flip(). Rotating text to 45 degrees makes the reader tilt their head.

coord_flip() before angled text
CHECK 06

The count behind each bar

When a bar is a mean, the number of rows behind it belongs on the chart or in the caption. This is the Week 9 habit, applied to a picture.

n = 5 and n = 62 look identical
3.Q

3.Q Knowledge check: Section 3

Q1. ggplot(ag, aes(class, mean_hwy)) + geom_bar() gives an error about y. Why?
geom_bar defaults to stat = "count", which produces y itself. Giving it a y as well is a contradiction. Use geom_col(), which is geom_bar(stat = "identity") under a clearer name.
Q2. You change position = "stack" to position = "fill" and the tallest bar drops from 62 to 1. What happened to the data?
Fill converts each bar to shares summing to one. The twelve rectangles are still there. The totals have been removed from the picture, which is the trade you are making.
Q3. Your bar chart of car classes puts the biggest bar sixth from the left. What caused that?
Alphabetical order comes from how R sorts characters, and it has nothing to do with your question. reorder() fixes it in one call, and the fix belongs in the plot rather than in the reader's head.
4
Section 4

Making it readable

Everything so far puts data on the page. This section is about the person who has to read it without you standing next to them.

4.1

4.1 Labels

Defaults246203040displhwydrv4frAfter labs()Bigger engines use more fuel234 cars, model years 1999 and 2008246203040Engine displacement (litres)Highway economy (mpg)Drive4frSource: US EPA fuel economy dataIdentical data, identical geoms. The right panel adds five strings and nothing else.Default labels are column names, which are written for you and not for your reader.
Figure 4.1. The same plot, before and after one labs() call.
  1. UnitOne point is one car in both panels. 234 points each.
  2. EncodingIdentical: position for engine size and economy, colour for drive type.
  3. ScaleIdentical axes and identical colours. Nothing about the data changed.
  4. StructureEconomy falls as engine size rises, in both panels.
  5. Groups and exceptionsThe left panel labels its axes displ and hwy, which are the names of the columns. The right panel names the quantities, the units and the source.
  6. Claim and limitDefault labels are written for the person who wrote the code. Nobody outside this room knows what hwy means. Five strings is the whole difference between a working file and a figure you can send.
Five strings, in order of value The title, because it states the finding. The two axis labels, because they carry the units. The caption, because it carries the source and the n. The legend title last, because the reader can often infer it.
ggplot(mpg, aes(displ, hwy, colour = drv)) +
  geom_point() +
  labs(
    title    = "Bigger engines use more fuel",
    subtitle = "234 cars, model years 1999 and 2008",
    x        = "Engine displacement (litres)",
    y        = "Highway economy (mpg)",
    colour   = "Drive",
    caption  = "Source: US EPA fuel economy data"
  )

# every aesthetic can be renamed the same way
labs(fill = "Drive", size = "Cylinders",
     shape = "Transmission")

# remove a label rather than blanking it
labs(x = NULL)      # no space reserved
labs(x = "")        # space reserved, nothing in it

# Write the title as a SENTENCE that states the
# finding, not as a description of the chart.
#   weak : "Displacement against highway mpg"
#   good : "Bigger engines use more fuel"
4.2

4.2 Scales

# axis limits
+ scale_y_continuous(limits = c(0, 50))
+ ylim(0, 50)                    # the short form

# limits DROP rows outside the range and warn.
# To zoom without dropping, use coord_cartesian:
+ coord_cartesian(ylim = c(20, 40))

# where the ticks go
+ scale_x_continuous(breaks = seq(2, 7, by = 1))

# how the numbers are written
+ scale_y_continuous(
    labels = function(v) paste0(v, " mpg"))

# colours: a named palette
+ scale_colour_brewer(palette = "Dark2")

# colours: exactly the ones you want
+ scale_colour_manual(
    values = c("4" = "#0072B2",
               "f" = "#D55E00",
               "r" = "#009E73"))

# a continuous colour scale between two ends
+ scale_colour_gradient(low  = "#EEEEEE",
                        high = "#CC0000")

The naming pattern

Every scale function is scale_, then the aesthetic, then the type. Once you know the pattern you can guess the name.

scale_x_continuousnumeric x axis
scale_y_log10log vertical axis
scale_colour_manualcolours you list
scale_fill_brewerfill, from a named palette
scale_size_areasize, with area proportional
limits and coord_cartesian are different ylim(20, 40) deletes every row outside that range before the geoms run, so a smooth fitted afterwards is fitted to the survivors. coord_cartesian(ylim = c(20, 40)) keeps every row and zooms at the end. On a plot with geom_smooth() the two give visibly different curves.

Week 7's rule still applies: a colour scale for an ordered variable should be one hue getting lighter or darker, and separate hues belong to categories.

4.3

4.3 Themes

theme_grey()2462040theme_bw()2462040theme_minimal()2462040theme_classic()2462040theme_grey()grey92 panelwhite gridthe defaulttheme_bw()white panelgrey92 gridprint and slidestheme_minimal()no panel fillgrey92 gridclean, keeps gridtheme_classic()white panelno gridaxis lines onlyA theme changes nothing that carries data. Same points, same axes, four sets of decoration.
Figure 4.3. One plot under four built-in themes. The panel and grid colours were read out of the theme objects.
  1. UnitOne point is one car, in all four panels.
  2. EncodingIdentical in all four. A theme controls nothing that carries data.
  3. ScaleIdentical axes throughout.
  4. StructureThe same falling cloud appears four times.
  5. Groups and exceptionstheme_grey fills the panel grey92 and draws white gridlines. theme_bw reverses that. theme_minimal removes the panel fill and keeps the grid. theme_classic removes the grid and draws two axis lines.
  6. Claim and limitA theme changes decoration only. The grey default is designed for a screen and can look heavy on a printed page or a projector. Pick one theme and use it for every figure in a report, because a reader notices inconsistency faster than they notice grey.
# the four you will actually use
+ theme_grey()      # the default
+ theme_bw()        # white panel, grey grid
+ theme_minimal()   # no panel fill, grid kept
+ theme_classic()   # axis lines, no grid

# set one for the whole session
theme_set(theme_minimal())

# change individual pieces
+ theme(
    plot.title   = element_text(face = "bold",
                                size = 14),
    axis.text    = element_text(colour = "black",
                                size = 10),
    legend.position = "bottom",
    panel.grid.minor = element_blank()
  )

# the element_ functions, by what they style
#   element_text()  words
#   element_line()  gridlines, axis lines, ticks
#   element_rect()  backgrounds and borders
#   element_blank() remove it entirely

# theme() must come AFTER a theme_*() call, or
# the theme_*() will overwrite your changes.
4.4

4.4 Legends

Where a legend comes from

You never ask for a legend. ggplot draws one whenever a column is mapped to a non-position aesthetic, and titles it with the column's name.

# a legend appears, titled "drv"
aes(colour = drv)

# rename it
+ labs(colour = "Drive type")

# move it
+ theme(legend.position = "bottom")
+ theme(legend.position = "none")    # remove

# change what the key entries say
+ scale_colour_discrete(
    labels = c("4" = "Four wheel",
               "f" = "Front wheel",
               "r" = "Rear wheel"))

# order the entries
+ scale_colour_discrete(
    breaks = c("f", "4", "r"))

When to remove it

  • The same variable is already on an axis, so the legend repeats it.
  • There are two or three groups and you can label them directly on the chart, which is closer to the data and easier to read.
  • The colour is decoration and carries nothing, which means it should have been set outside aes() in the first place.
A legend is a lookup table Every glance at a legend is a trip away from the data and back. Direct labels beat a legend whenever the chart has room for them.
A legend you did not expect Go back to slide 1.5. A legend with one entry, labelled with something you typed, means a constant ended up inside aes().
4.5

4.5 Saving the figure

# save the plot in a variable, then save the file
p <- ggplot(mpg, aes(displ, hwy)) + geom_point()

ggsave("engine-economy.png", plot = p,
       width = 8, height = 5, units = "in",
       dpi = 300)

# with no plot = argument it saves the LAST plot
# you printed, which is rarely what you meant
ggsave("something.png")

# the format comes from the extension
ggsave("fig.pdf",  plot = p, width = 8, height = 5)
ggsave("fig.svg",  plot = p, width = 8, height = 5)
ggsave("fig.png",  plot = p, width = 8, height = 5)

# defaults worth knowing
#   units = "in"
#   dpi   = 300
#   width and height come from the current
#   plotting window if you leave them out

Why the size argument matters more than it looks

Text in a ggplot is sized in points, and points do not scale with the figure. Save the same plot at 8 by 5 inches and at 4 by 2.5 inches, and the second one has text that is twice as large relative to the panel.

The practical rule Save at the size the figure will be printed or displayed at. Shrinking a large export in Word shrinks the text with it, and enlarging a small one makes the text coarse.

Which format

.pngslides, Word, email. Fixed resolution.
.pdfreports and printing. Scales without blurring.
.svgthe web, or editing in Illustrator afterwards.

For your Week 12 presentation, export at the size of the slide and check the axis text is readable from the back of the room before you submit.

4.Q

4.Q Knowledge check: Section 4

Q1. You add ylim(20, 40) to a plot with geom_smooth() and the fitted curve changes shape. Why?
Scale limits filter the data. Every row outside the range is dropped and ggplot warns about it, and the smooth is then fitted to the survivors. coord_cartesian(ylim = ...) zooms at the end and leaves the fit alone.
Q2. Your theme() call has no effect on the plot. What is the likely cause?
Layers are applied in order, and theme_minimal() replaces the whole theme. Put your theme() call after the theme_*() call, so your specific changes land on top of the general one.
Q3. ggsave("fig.png") saves a plot you did not want. Why?
With no plot argument, ggsave uses last_plot(). In a script that draws several figures, that is whichever one printed most recently. Naming the plot explicitly removes the guesswork.
5
Section 5

Putting it together

A short workflow, the mistakes that cost the most time, and what carries over into Part 2 next week.

5.1

5.1 A workflow that works every time

StepDo thisWhy
1Get the variable you want to show into a column ggplot can only map columns. Row names, and variables spread across column headers, have to be joined or reshaped in first. That is Weeks 8 and 9.
2Check the type of every column you plan to map str() takes a second and tells you whether you will get a continuous scale or a discrete one.
3Write the canvas: data and the aesthetic mapping Decide which variable deserves position, which is the accurate channel, before you decide anything about appearance.
4Add one geom and look at it Most problems show up here. Overplotting, dropped rows, a legend you did not want.
5Read the console Removed rows, the smoothing method, the bin count. ggplot tells you these once and never again.
6Add labels Title as a sentence that states the finding. Units on both axes. n and source in the caption.
7Theme and save at the final size One theme across the whole report, and an export sized for where it will be read.
Steps 4 and 5 are the ones people skip Adding six layers and then looking at the result makes every problem harder to locate. Add one, look, add the next.
5.2

5.2 Five errors to avoid

The errorWhat you see The fix
A constant inside aes() Points in the wrong colour and a legend with one entry, labelled with the word you typed. Move it outside the aes() bracket. Mapped values come from columns; set values are constants.
A mapping in the wrong place One fitted line where you expected three, or three where you expected one. Canvas-level mappings are inherited by every layer. Decide which layers should see it.
A number that should be a label A colour bar, or an axis with ticks at 5.5 and 6.5 for a variable that only takes whole values. factor(), preferably in the data frame rather than inside every aes() call.
Ignoring the console A chart that looks complete, drawn from 116 of 153 rows. Read every message once. Decide what to do about it, then report the n you actually used.
Overplotting A scatterplot that looks sparse when nearly half the marks are hidden. Compare nrow() against nrow(unique(...)) on the two plotted columns. Then alpha, jitter or bin.
What they share None of these produces an error. All five produce a chart. That is why the checking has to be a habit rather than a reaction.
5.3

5.3 Summary

The grammar

  • Data, then an aesthetic mapping, then a geom. Those three are compulsory. Scales, labels and themes have defaults.
  • A ggplot is an object. + adds to it, and printing draws it.
  • Inside aes() is a mapping from a column. Outside it is a constant. Getting that wrong gives you a colour you did not ask for and a legend you did not want.
  • A mapping on the canvas is inherited by every layer. Inside a geom it stays there, and the difference was three fitted lines against one.
  • The column's storage type chooses the scale, so factor() is how you say a number is a label.

The geoms

  • geom_bar() counts rows and computes its own y. geom_col() draws the numbers you supply.
  • Stack, dodge and fill answer three different questions from twelve identical rectangles.
  • Order the bars on purpose. Alphabetical is an accident of storage.
  • 234 rows and 126 distinct pairs meant 108 invisible points.
  • 37 rows were removed before drawing, and the only warning was one line in the console.

The reader

A title that states the finding, units on both axes, the n and the source in the caption, one theme throughout, exported at the size it will be read.

5.4

5.4 Before next week

Practise

  • Take three figures you made in Weeks 3, 4 and 6 with base R and rewrite them in ggplot. Keep the same data and the same message. Note which one took longest and why.
  • Run the aes() trap yourself. Draw one plot with a colour inside aes() and one with it outside, and read both colours back with ggplot_build(p)$data[[1]]$colour.
  • Take a Week 9 aggregate table and draw it with geom_col(), ordered by size, flipped, with the count behind each bar in the caption.
  • Plot airquality ozone against temperature and write down, in one sentence, how many days the figure actually rests on.

Week 11: ggplot part 2

Next week extends the same grammar rather than replacing it. Facets, which split one plot into a grid of panels. Coordinate systems. Combining several plots into one figure. Annotating a chart with text and reference lines.

The link forward Faceting is the Week 7 idea of small multiples, written as one more layer. Everything you learned there about shared axes applies, and ggplot shares them by default.
Bring One of your rewritten figures, saved as a png at the size you would put it on a slide, and one sentence on what the ggplot version made easier or harder.
TECH3100 · Lesson 10
← → navigate · T contents

Contents

Press T or Escape to close