This material has been reproduced and communicated to you by or on behalf of Kaplan Business School pursuant to Part VB of the Copyright Act 1968 (the Act).
The material in this communication may be subject to copyright under the Act. Any further reproduction or communication of this material by you may be the subject of copyright protection under the Act.
Do not remove this notice.
| Week | Topic |
|---|---|
| 1–4 | Variable types. Descriptive statistics. Histograms, boxplots, scatter plots, correlation. |
| 5–7 | Trend lines. Time series. Multivariate data. |
| 8 | Missing values: imputation and interpolation. |
| 9 | Aggregates and joining tables. |
| 10 | Data visualisation with ggplot, part 1. |
| 11 | Data visualisation with ggplot, part 2. |
| 12 | Individual presentations. |
+.aes() and what belongs outside it, and predict what
happens when you get it wrong.geom_bar() needs no y, and when to use stat = "identity"
instead.install.packages("ggplot2") # once per machine
library(ggplot2) # once per session
packageVersion("ggplot2")
#> '3.4.4'
library(ggplot) is an error that catches everybody once.ggplot accepts colour and color everywhere, and treats them as the
same argument. Pick one and stay with it.
234 cars, 11 columns. Fuel economy from the US Environmental Protection Agency.
| Column | Type | Range or values |
|---|---|---|
| displ | numeric | 1.6 to 7.0 litres |
| hwy | integer | 12 to 44 mpg |
| cty | integer | city mpg |
| drv | character | 4, f, r |
| class | character | 7 values |
| cyl | integer | 4, 5, 6, 8 |
We also use airquality from Week 8, so the missing values come
back.
A chart is not a picture you draw. It is a sentence you write, with a fixed set of parts. Learn the parts and every chart becomes the same job.
The two letters in ggplot stand for grammar of graphics. A grammar is a small set of parts that combine in rules, and this is the whole set.
Every ggplot you write has these pieces in this order, and each one has a job.
An aesthetic mapping is a wiring diagram. It says which column travels down which visual channel. Week 7 called these channels; ggplot calls them aesthetics.
A ggplot is an object. You can build it in pieces, keep it in a variable, and add to it later. Nothing appears on screen until you print it.
viz <- ggplot(data = mpg)
class(viz) #> "gg" "ggplot"
length(viz$layers) #> 0 an empty canvas
viz <- viz + aes(x = displ, y = hwy)
viz <- viz + geom_point()
length(viz$layers) #> 1
viz <- viz + geom_smooth(method = "lm")
length(viz$layers) #> 2
viz # this line draws it
print(viz) # the same thing, explicitly
print(). That is the most common reason a plot fails to appear.In practice you write it as one expression. The + goes at the end of
the line, never at the start of the next one.
ggplot(data = mpg, aes(x = displ, y = hwy)) +
geom_point(colour = "steelblue") +
geom_smooth(method = "lm", formula = y ~ x) +
labs(x = "Engine displacement (litres)",
y = "Highway economy (mpg)") +
theme_minimal()
# This does NOT work:
ggplot(mpg, aes(displ, hwy))
+ geom_point()
#> Error: Cannot use `+` with a single argument
# R finished the first line, so the second line
# starts a new expression with nothing on its left.
ggplot(mpg, aes(displ, hwy)) means the same as
ggplot(data = mpg, aes(x = displ, y = hwy)), because data is the first argument and x
and y are the first two aesthetics. Write the names out while you are learning.# WRONG: the string is treated as data
ggplot(airquality, aes(Ozone, Temp)) +
geom_point(aes(colour = "darkred"))
# RIGHT: the string is treated as an instruction
ggplot(airquality, aes(Ozone, Temp)) +
geom_point(colour = "darkred")
# Proof, read straight out of the plot object
p <- ggplot(airquality, aes(Ozone, Temp)) +
geom_point(aes(colour = "darkred"))
unique(ggplot_build(p)$data[[1]]$colour)
#> "#F8766D"
col2rgb("darkred")
#> [,1]
#> red 139 = #8B0000
#> green 0
#> blue 0
# The rule, in one line:
# aes(x = column) maps a column
# argument = value sets a constant# colour on the CANVAS: every layer inherits it
ggplot(mpg, aes(displ, hwy, colour = drv)) +
geom_point() +
geom_smooth(method = "lm", formula = y ~ x)
#> 3 fitted lines
# colour on the GEOM: only that layer sees it
ggplot(mpg, aes(displ, hwy)) +
geom_point(aes(colour = drv)) +
geom_smooth(method = "lm", formula = y ~ x)
#> 1 fitted line
# the slopes the three lines are reporting
for (d in c("4", "f", "r")) {
f <- lm(hwy ~ displ, subset(mpg, drv == d))
cat(d, round(coef(f)[2], 2), "\n")
}
#> 4 -2.88
#> f -3.60
#> r -0.92
coef(lm(hwy ~ displ, mpg))[2] #> -3.53
# Rule of thumb: put it on the canvas when every
# layer should use it, and in the geom when only
# one layer should.# Month is stored as a number
class(airquality$Month) #> "integer"
# so ggplot builds a CONTINUOUS colour scale
ggplot(airquality, aes(Temp, Ozone, colour = Month)) +
geom_point()
# factor() says: these numbers are labels
ggplot(airquality,
aes(Temp, Ozone, colour = factor(Month))) +
geom_point()
# The same question comes up for x and y:
# a numeric x gives a continuous axis
# a factor x gives one tick per level
# Fix it in the data rather than in every plot
aq <- airquality
aq$Month <- factor(aq$Month,
labels = c("May","Jun","Jul","Aug","Sep"))
# Now the legend reads in English and the order
# is the one you chose, not alphabetical.# base R: you build the palette and the legend
cols <- c("4" = "#F8766D",
"f" = "#00BA38",
"r" = "#619CFF")
plot(mpg$displ, mpg$hwy, pch = 19,
col = cols[mpg$drv],
xlab = "displ", ylab = "hwy")
legend("topright", legend = names(cols),
pch = 19, col = cols, bty = "n")
# ggplot: you declare the mapping and stop
ggplot(mpg, aes(displ, hwy, colour = drv)) +
geom_point()
# What ggplot did without being asked:
# picked 3 evenly spaced hues
# matched each drv value to one
# drew a legend, titled it from the column name
# placed it, sized it, spaced it
# In base R each of those is a line you maintain.geom_point(aes(colour = "blue")) runs without an error. What colour are the points?ggplot(mpg, aes(displ, hwy, colour = drv)) + geom_point() + geom_smooth() draws how many smooth lines?aes(colour = Month) on airquality produces a colour bar rather than a key with five entries. Why?The geoms you will use most, and the two things that go wrong on every scatterplot: marks hidden underneath other marks, and rows quietly removed before drawing.
# the minimum
ggplot(mpg, aes(displ, hwy)) + geom_point()
# mapped: the value comes from a column
geom_point(aes(colour = drv)) # 3 hues
geom_point(aes(shape = drv)) # 3 shapes
geom_point(aes(size = cyl)) # area scale
geom_point(aes(alpha = cty)) # transparency
# set: the value is fixed for every point
geom_point(colour = "steelblue")
geom_point(size = 3)
geom_point(shape = 17) # solid triangle
geom_point(alpha = 0.3)
# both at once is fine
geom_point(aes(colour = drv), size = 3, alpha = .7)
| Aesthetic | Suits | Limit |
|---|---|---|
| x, y | anything | read most accurately |
| colour | nominal or continuous | 6 to 8 categories |
| shape | nominal only | 6 by default, then ggplot refuses |
| size | continuous only | judged roughly, so use it for context |
| alpha | continuous, or density | weakest of the five |
Mapping the same column to colour and shape at once is a reasonable choice for a printed handout, because it survives being photocopied.
Every extra degree is associated with 2.43 more parts per billion, on average, across this range. The intercept has no meaning here, because no day in the data is near zero degrees.
ggplot(airquality, aes(Temp, Ozone)) +
geom_point() +
geom_smooth()
#> `geom_smooth()` using method = 'loess'
#> and formula = 'y ~ x'
# ggplot picks loess when n < 1000 and gam above it.
# Say which you used, because the default changes
# with the size of your data.
# a straight line instead
ggplot(airquality, aes(Temp, Ozone)) +
geom_point() +
geom_smooth(method = "lm", formula = y ~ x)
# turn the band off when you do not want to
# discuss it
geom_smooth(method = "lm", formula = y ~ x,
se = FALSE)
# the same fit as a number
f <- lm(Ozone ~ Temp, data = airquality)
round(coef(f), 2)
#> (Intercept) Temp
#> -147.00 2.43
round(summary(f)$r.squared, 4) #> 0.4877# how bad is it?
nrow(mpg) #> 234
nrow(unique(mpg[, c("displ", "hwy")])) #> 126
# 108 points are sitting on top of another
# 1. transparency: overlap becomes darkness
ggplot(mpg, aes(displ, hwy)) +
geom_point(alpha = 0.25)
# 2. jitter: a small random offset
ggplot(mpg, aes(displ, hwy)) +
geom_jitter(width = 0.1, height = 0.6)
# geom_jitter(...) is short for
# geom_point(position = position_jitter(...))
# 3. count: one mark per pair, sized by how many
ggplot(mpg, aes(displ, hwy)) + geom_count()
# 4. for very large n, bin instead of plotting
ggplot(diamonds, aes(carat, price)) +
geom_hex(bins = 40)
# Try transparency first. Jitter only when the
# exact values do not matter to the reader.ggplot(airquality, aes(Temp, Ozone)) +
geom_point()
#> Warning message:
#> Removed 37 rows containing missing values
#> (`geom_point()`).
sum(is.na(airquality$Ozone)) #> 37
nrow(airquality) #> 153
# Decide what to do, then say what you did:
# (a) keep the warning and report the n
labs(caption = "116 of 153 days had an ozone reading")
# (b) drop them on purpose, so the count is yours
aq <- subset(airquality, !is.na(Ozone))
nrow(aq) #> 116
# (c) show where they are
aq2 <- airquality
aq2$missing <- is.na(aq2$Ozone)
ggplot(aq2, aes(Temp, Ozone, colour = missing)) +
geom_point()
# Silencing the warning with na.rm = TRUE hides
# the only signal you were given. Do not.| Geom | Draws | Needs | Week it replaces |
|---|---|---|---|
| geom_point() | one mark per row | x and y | Week 4 scatter plots |
| geom_line() | points joined in x order | x and y | Week 6 time series |
| geom_smooth() | a fitted curve and its band | x and y | Week 5 trend lines |
| geom_bar() | one bar per category, counted | x only | Week 1 bar charts |
| geom_col() | one bar per row, at the height you give | x and y | Week 9 aggregate tables |
| geom_histogram() | counts in bins | x only | Week 3 histograms |
| geom_boxplot() | five-number summary per group | x and y | Week 3 boxplots |
| geom_tile() | a filled rectangle per cell | x, y and fill | Week 7 heat maps |
nrow(unique(...)) returns 126. What does that tell you?geom_smooth() prints a message naming loess. Why should you state the method in your caption?Bar charts are where ggplot does the most work on your behalf, and where beginners lose the most time. One geom does two different jobs.
# no y anywhere in this call
ggplot(mpg, aes(x = class)) + geom_bar()
# what happened underneath
p <- ggplot(mpg, aes(x = class)) + geom_bar()
ggplot_build(p)$data[[1]][, c("x", "y", "count")]
#> x y count
#> 1 1 5 5
#> 2 2 47 47
#> 3 3 41 41
#> 4 4 11 11
#> 5 5 33 33
#> 6 6 35 35
#> 7 7 62 62
sum(table(mpg$class)) #> 234
# the same numbers, computed yourself
table(mpg$class)
# geom_histogram does the same thing for a
# continuous x: it bins, then counts
ggplot(mpg, aes(x = hwy)) +
geom_histogram(bins = 20)
# Always set bins or binwidth. The default of 30
# is a placeholder and ggplot says so.geom_bar(). Drawing numbers you already have:
geom_col(). If you find yourself writing stat = "identity", you wanted
geom_col().# Week 9 gives you the table
ag <- aggregate(hwy ~ class, data = mpg, FUN = mean)
ag
#> class hwy
#> 1 2seater 24.80000
#> 2 compact 28.29787
#> ...
# two ways to draw it, both identical
ggplot(ag, aes(x = class, y = hwy)) +
geom_bar(stat = "identity")
ggplot(ag, aes(x = class, y = hwy)) +
geom_col()
# geom_col() IS geom_bar(stat = "identity").
# Prefer geom_col: the name says what it does.
# long labels: flip the chart, do not rotate text
ggplot(ag, aes(x = reorder(class, hwy), y = hwy)) +
geom_col() +
coord_flip() +
labs(x = NULL, y = "Mean highway mpg")
# Carry the count through so it can be reported
ag$n <- as.vector(table(mpg$class))# stack: the default. Totals are readable,
# individual segments above the first are not.
ggplot(mpg, aes(class, fill = drv)) +
geom_bar(position = "stack")
# dodge: side by side. Segments are comparable,
# totals are gone.
ggplot(mpg, aes(class, fill = drv)) +
geom_bar(position = "dodge")
# fill: proportions. Shares are comparable,
# totals are gone entirely.
ggplot(mpg, aes(class, fill = drv)) +
geom_bar(position = "fill") +
labs(y = "Share of class")
# the same three counts, as a table
table(mpg$class, mpg$drv)
# Which to choose:
# "how many altogether" -> stack
# "how do the groups compare"-> dodge
# "what share of each" -> fill, and
# print the totals somewhere else# default: R sorts character columns alphabetically
ggplot(mpg, aes(x = class)) + geom_bar()
# reorder by how many rows are in each level
ggplot(mpg, aes(x = reorder(class, class, length))) +
geom_bar()
levels(reorder(mpg$class, mpg$class, length))
#> "2seater" "minivan" "pickup" "subcompact"
#> "midsize" "compact" "suv" ascending
# descending: negate the sorting value
ag <- aggregate(hwy ~ class, mpg, mean)
ggplot(ag, aes(reorder(class, -hwy), hwy)) +
geom_col()
# an order you choose by hand
mpg$class <- factor(mpg$class,
levels = c("2seater", "subcompact", "compact",
"midsize", "minivan", "suv", "pickup"))
# Order by size for a ranking. Keep a natural
# order (Mon to Sun, May to Sep) when one exists.
# Alphabetical needs a reason.Are you counting rows, or drawing numbers you already computed? The first is
geom_bar(), the second is geom_col().
Bar length is read as a quantity, so the axis has to start at zero. ggplot does this by default and you can override it. Do not.
Alphabetical only when the reader will look a category up by name. Otherwise order by size, or by a natural sequence.
Stack for totals, dodge for comparison, fill for shares. Fill throws the totals away, so print them somewhere else.
Long category names go on the vertical axis with coord_flip(). Rotating text
to 45 degrees makes the reader tilt their head.
When a bar is a mean, the number of rows behind it belongs on the chart or in the caption. This is the Week 9 habit, applied to a picture.
ggplot(ag, aes(class, mean_hwy)) + geom_bar() gives an error about y. Why?position = "stack" to position = "fill" and the tallest bar drops from 62 to 1. What happened to the data?Everything so far puts data on the page. This section is about the person who has to read it without you standing next to them.
ggplot(mpg, aes(displ, hwy, colour = drv)) +
geom_point() +
labs(
title = "Bigger engines use more fuel",
subtitle = "234 cars, model years 1999 and 2008",
x = "Engine displacement (litres)",
y = "Highway economy (mpg)",
colour = "Drive",
caption = "Source: US EPA fuel economy data"
)
# every aesthetic can be renamed the same way
labs(fill = "Drive", size = "Cylinders",
shape = "Transmission")
# remove a label rather than blanking it
labs(x = NULL) # no space reserved
labs(x = "") # space reserved, nothing in it
# Write the title as a SENTENCE that states the
# finding, not as a description of the chart.
# weak : "Displacement against highway mpg"
# good : "Bigger engines use more fuel"# axis limits
+ scale_y_continuous(limits = c(0, 50))
+ ylim(0, 50) # the short form
# limits DROP rows outside the range and warn.
# To zoom without dropping, use coord_cartesian:
+ coord_cartesian(ylim = c(20, 40))
# where the ticks go
+ scale_x_continuous(breaks = seq(2, 7, by = 1))
# how the numbers are written
+ scale_y_continuous(
labels = function(v) paste0(v, " mpg"))
# colours: a named palette
+ scale_colour_brewer(palette = "Dark2")
# colours: exactly the ones you want
+ scale_colour_manual(
values = c("4" = "#0072B2",
"f" = "#D55E00",
"r" = "#009E73"))
# a continuous colour scale between two ends
+ scale_colour_gradient(low = "#EEEEEE",
high = "#CC0000")
Every scale function is scale_, then the aesthetic, then the type. Once you
know the pattern you can guess the name.
| scale_x_continuous | numeric x axis |
| scale_y_log10 | log vertical axis |
| scale_colour_manual | colours you list |
| scale_fill_brewer | fill, from a named palette |
| scale_size_area | size, with area proportional |
ylim(20, 40) deletes every row outside that range before the geoms run, so a
smooth fitted afterwards is fitted to the survivors.
coord_cartesian(ylim = c(20, 40)) keeps every row and zooms at the end. On a plot
with geom_smooth() the two give visibly different curves.Week 7's rule still applies: a colour scale for an ordered variable should be one hue getting lighter or darker, and separate hues belong to categories.
# the four you will actually use
+ theme_grey() # the default
+ theme_bw() # white panel, grey grid
+ theme_minimal() # no panel fill, grid kept
+ theme_classic() # axis lines, no grid
# set one for the whole session
theme_set(theme_minimal())
# change individual pieces
+ theme(
plot.title = element_text(face = "bold",
size = 14),
axis.text = element_text(colour = "black",
size = 10),
legend.position = "bottom",
panel.grid.minor = element_blank()
)
# the element_ functions, by what they style
# element_text() words
# element_line() gridlines, axis lines, ticks
# element_rect() backgrounds and borders
# element_blank() remove it entirely
# theme() must come AFTER a theme_*() call, or
# the theme_*() will overwrite your changes.You never ask for a legend. ggplot draws one whenever a column is mapped to a non-position aesthetic, and titles it with the column's name.
# a legend appears, titled "drv"
aes(colour = drv)
# rename it
+ labs(colour = "Drive type")
# move it
+ theme(legend.position = "bottom")
+ theme(legend.position = "none") # remove
# change what the key entries say
+ scale_colour_discrete(
labels = c("4" = "Four wheel",
"f" = "Front wheel",
"r" = "Rear wheel"))
# order the entries
+ scale_colour_discrete(
breaks = c("f", "4", "r"))
aes() in the first place.aes().# save the plot in a variable, then save the file
p <- ggplot(mpg, aes(displ, hwy)) + geom_point()
ggsave("engine-economy.png", plot = p,
width = 8, height = 5, units = "in",
dpi = 300)
# with no plot = argument it saves the LAST plot
# you printed, which is rarely what you meant
ggsave("something.png")
# the format comes from the extension
ggsave("fig.pdf", plot = p, width = 8, height = 5)
ggsave("fig.svg", plot = p, width = 8, height = 5)
ggsave("fig.png", plot = p, width = 8, height = 5)
# defaults worth knowing
# units = "in"
# dpi = 300
# width and height come from the current
# plotting window if you leave them out
Text in a ggplot is sized in points, and points do not scale with the figure. Save the same plot at 8 by 5 inches and at 4 by 2.5 inches, and the second one has text that is twice as large relative to the panel.
| .png | slides, Word, email. Fixed resolution. |
| reports and printing. Scales without blurring. | |
| .svg | the web, or editing in Illustrator afterwards. |
For your Week 12 presentation, export at the size of the slide and check the axis text is readable from the back of the room before you submit.
ylim(20, 40) to a plot with geom_smooth() and the fitted curve changes shape. Why?theme() call has no effect on the plot. What is the likely cause?ggsave("fig.png") saves a plot you did not want. Why?A short workflow, the mistakes that cost the most time, and what carries over into Part 2 next week.
| Step | Do this | Why |
|---|---|---|
| 1 | Get the variable you want to show into a column | ggplot can only map columns. Row names, and variables spread across column headers, have to be joined or reshaped in first. That is Weeks 8 and 9. |
| 2 | Check the type of every column you plan to map | str() takes a second and tells you whether you will get a continuous scale
or a discrete one. |
| 3 | Write the canvas: data and the aesthetic mapping | Decide which variable deserves position, which is the accurate channel, before you decide anything about appearance. |
| 4 | Add one geom and look at it | Most problems show up here. Overplotting, dropped rows, a legend you did not want. |
| 5 | Read the console | Removed rows, the smoothing method, the bin count. ggplot tells you these once and never again. |
| 6 | Add labels | Title as a sentence that states the finding. Units on both axes. n and source in the caption. |
| 7 | Theme and save at the final size | One theme across the whole report, and an export sized for where it will be read. |
| The error | What you see | The fix |
|---|---|---|
| A constant inside aes() | Points in the wrong colour and a legend with one entry, labelled with the word you typed. | Move it outside the aes() bracket. Mapped values come from columns; set values are constants. |
| A mapping in the wrong place | One fitted line where you expected three, or three where you expected one. | Canvas-level mappings are inherited by every layer. Decide which layers should see it. |
| A number that should be a label | A colour bar, or an axis with ticks at 5.5 and 6.5 for a variable that only takes whole values. | factor(), preferably in the data frame rather than inside every aes()
call. |
| Ignoring the console | A chart that looks complete, drawn from 116 of 153 rows. | Read every message once. Decide what to do about it, then report the n you actually used. |
| Overplotting | A scatterplot that looks sparse when nearly half the marks are hidden. | Compare nrow() against nrow(unique(...)) on the two plotted
columns. Then alpha, jitter or bin. |
+ adds to it, and printing draws it.aes() is a mapping from a column. Outside it is a constant. Getting
that wrong gives you a colour you did not ask for and a legend you did not want.factor() is how you say a
number is a label.geom_bar() counts rows and computes its own y. geom_col() draws
the numbers you supply.A title that states the finding, units on both axes, the n and the source in the caption, one theme throughout, exported at the size it will be read.
aes() and one
with it outside, and read both colours back with
ggplot_build(p)$data[[1]]$colour.geom_col(), ordered by size,
flipped, with the count behind each bar in the caption.airquality ozone against temperature and write down, in one sentence,
how many days the figure actually rests on.Next week extends the same grammar rather than replacing it. Facets, which split one plot into a grid of panels. Coordinate systems. Combining several plots into one figure. Annotating a chart with text and reference lines.
Press T or Escape to close