This material has been reproduced and communicated to you by or on behalf of Kaplan Business School pursuant to Part VB of the Copyright Act 1968 (the Act).
The material in this communication may be subject to copyright under the Act. Any further reproduction or communication of this material by you may be the subject of copyright protection under the Act.
Do not remove this notice.
| Week | Topic |
|---|---|
| 1–4 | Variable types. Descriptive statistics. Histograms, boxplots, scatter plots, correlation. |
| 5–6 | Group presentations. Trend lines. Time series. |
| 7 | Visualisation for multivariate data. |
| 8 | Missing values: imputation and interpolation. |
| 9 | Aggregates and joining tables. |
| 10–11 | Data visualisation with ggplot. |
| 12 | Individual presentations. |
tapply, aggregate,
table and ave, and read the dplyr equivalents.54 rows. Number of warp breaks on a loom, under two wool types and three tension levels, nine observations in every combination.
wb <- warpbreaks
str(wb)
#> 'data.frame': 54 obs. of 3 variables:
#> $ breaks : num 26 30 54 25 70 52 51 26 67 18 ...
#> $ wool : Factor w/ 2 levels "A","B"
#> $ tension: Factor w/ 3 levels "L","M","H"
table(wb$wool, wb$tension)
#> L M H
#> A 9 9 9
#> B 9 9 9
A balanced design, which makes every marginal mean a clean average of the cells beneath it.
orders, customers and products, built by hand so that every row is printable and every join result can be checked by eye.
50 states of crime data, joined to state.name, state.region,
state.division and state.area, all of which ship with R.
One pattern with three steps. Split the rows into groups, apply a function to each group, combine the answers. Every command in this section is that pattern with different arguments.
Splitting warpbreaks on tension gives three groups of 18. Applying
mean to each gives three numbers. Combining them gives a table with three rows.
The same split can produce a table with one row per group, or the original table with one extra column. Which you want depends on the question.
tapply(wb$breaks, wb$tension, mean)
#> L M H
#> 36.38889 26.38889 21.66667
aggregate(breaks ~ tension, data = wb, FUN = mean)
#> tension breaks
#> 1 L 36.38889
#> 2 M 26.38889
#> 3 H 21.66667
# tapply returns a named vector
# aggregate returns a data frameaggregate(breaks ~ wool + tension, wb, mean)
tapply(wb$breaks, list(wb$wool, wb$tension), mean)
#> L M H
#> A 44.55556 24.0000 24.55556
#> B 28.22222 28.7778 18.77778
# counts, no aggregation function needed
table(wb$wool, wb$tension)
# sums laid out as a two-way table
xtabs(breaks ~ wool + tension, data = wb)aggregate when you want a data frame you will join or plot next. tapply
when you want a quick number at the console. table when you only need counts.library(dplyr)
wb %>%
group_by(tension) %>%
summarize(mean_breaks = mean(breaks))
wb %>%
group_by(wool, tension) %>%
summarize(mean_breaks = mean(breaks),
.groups = "drop")summarize names the output column; aggregate reuses the input
name..groups = "drop". Forgetting it changes what the next verb in the chain does.n() counts rows inside summarize. In base R that is
length() or table().NA by default.# every cell, with three statistics at once
cell <- aggregate(breaks ~ wool + tension, wb,
function(v) c(n = length(v),
mean = mean(v),
sd = sd(v)))
# aggregate returns a matrix column, so flatten it
cell <- do.call(data.frame, cell)
names(cell) <- c("wool", "tension", "n", "mean", "sd")
cell
#> wool tension n mean sd
#> 1 A L 9 44.55556 18.097644
#> 2 B L 9 28.22222 9.858724
#> 3 A M 9 24.00000 8.660254
#> 4 B M 9 28.77778 9.431036
#> 5 A H 9 24.55556 10.272671
#> 6 B H 9 18.77778 4.893305
# the same means as a two-way layout
round(tapply(wb$breaks,
list(wb$wool, wb$tension), mean), 2)
# Report the count and the spread alongside every
# mean. A mean on its own hides both.# the marginal means, one variable at a time
tapply(wb$breaks, wb$wool, mean)
#> A B
#> 31.03704 25.25926
# the cell means, both variables at once
round(tapply(wb$breaks,
list(wb$wool, wb$tension), mean), 2)
#> L M H
#> A 44.56 24.00 24.56
#> B 28.22 28.78 18.78
# the picture that shows it
means <- tapply(wb$breaks,
list(wb$wool, wb$tension), mean)
matplot(t(means), type = "b", pch = 19, lty = 1,
lwd = 2, col = c("#0072B2", "#D55E00"),
xaxt = "n", xlab = "tension",
ylab = "mean breaks")
axis(1, 1:3, colnames(means))
legend("topright", c("wool A", "wool B"), bty = "n",
col = c("#0072B2", "#D55E00"), lty = 1, lwd = 2)
# Before you report a grouped mean, group by the
# other variables too and check the ordering holds.| Function | Returns |
|---|---|
| mean, median | Centre. Median when the group is skewed. |
| sd, var, IQR | Spread. Report one of these next to every mean. |
| min, max, range | Extremes. |
| quantile | Any percentile. Returns several numbers at once. |
| length | Rows in the group. The dplyr name is n(). |
| function(v) length(unique(v)) | Distinct values. The dplyr name is n_distinct(). |
| sum | Totals, and counts of a logical column. |
# several statistics in one call
aggregate(breaks ~ tension, wb,
function(v) c(n = length(v),
mean = mean(v),
sd = sd(v),
iqr = IQR(v)))
# quantiles
aggregate(breaks ~ tension, wb,
function(v) quantile(v, c(.25, .5, .75)))
# distinct values, base R
length(unique(wb$tension)) #> 3
# counting a condition
aggregate(cbind(high = breaks > 30) ~ tension,
wb, sum)
Week 8's rules apply inside every group, and the two main functions behave differently.
aq <- airquality
# tapply passes the NAs straight to mean()
tapply(aq$Ozone, aq$Month, mean)
#> 5 6 7 8 9
#> NA NA NA NA NA
tapply(aq$Ozone, aq$Month, mean, na.rm = TRUE)
#> 5 6 7 8 9
#> 23.61538 29.44444 59.11538 59.96154 31.44828
# aggregate with a formula DROPS incomplete rows
# before it groups, silently
aggregate(Ozone ~ Month, aq, mean)
aggregate(Ozone ~ Month, aq, length)
#> Month Ozone
#> 1 5 26
#> 2 6 9
#> 3 7 26
#> 4 8 26
#> 5 9 29
Every month has 30 or 31 days. The counts above are 26, 9, 26, 26 and 29. June contributed nine days to its mean, because 21 of its Ozone readings are missing.
Nothing in the output says so. The June mean of 29.44 looks like the other four, and it rests on a third as much data.
aggregate(cbind(mean = Ozone) ~ Month, aq, mean)
aggregate(cbind(n = Ozone) ~ Month, aq, length)The formula interface uses na.action = na.omit by default. Set
na.action = na.pass and supply na.rm = TRUE yourself if you want
control over which rows are dropped.
# the test is computed once per group
tm <- tapply(wb$breaks, wb$tension, mean)
tm
#> L M H
#> 36.38889 26.38889 21.66667
keep <- names(tm)[tm > 25]
keep #> "L" "M"
sub <- wb[wb$tension %in% keep, ]
nrow(sub) #> 36
nrow(wb) #> 54
# the dplyr form does it in one chain
# wb %>% group_by(tension) %>%
# filter(mean(breaks) > 25)
# Contrast with a ROW filter, which tests each
# row on its own value:
nrow(wb[wb$breaks > 25, ]) #> 30
# Different rows, different count, different
# question. Say which one you meant.ave() lets you compute.# ave() computes a group statistic and returns it
# aligned to the original rows: same length as the
# input, one value per row
wb$cellmean <- ave(wb$breaks, wb$wool, wb$tension)
wb$dev <- wb$breaks - wb$cellmean
nrow(wb) #> 54
round(max(wb$dev), 2) #> 25.44
# within every cell the deviations sum to zero
tapply(wb$dev, list(wb$wool, wb$tension), sum)
#> L M H
#> A 0 0 0
#> B 0 0 0
# ave() takes any function, not just the mean
wb$rank <- ave(wb$breaks, wb$tension,
FUN = function(v) rank(-v))
# the dplyr form
# wb %>% group_by(wool, tension) %>%
# mutate(dev = breaks - mean(breaks))aggregate(Ozone ~ Month, airquality, length) returns 26, 9, 26, 26 and 29 for the five months, each of which has 30 or 31 days. Why?na.action = na.omit removes incomplete rows silently. June contributes 9 days to its mean because 21 of its Ozone readings are missing, and the printed mean gives no hint of it.summarise and mutate applied to the same grouping?Data arrives split across tables so that each fact is stored once. A join is how you put the pieces back together for one question, and the choice of join is the choice of which unmatched rows to keep.
Storing a customer's name once per order means storing it thousands of times, and changing it in thousands of places. Splitting the table fixes that, and creates the need to join.
customers <- data.frame(
customer_id = c("C1","C2","C3","C4","C5"),
name = c("Ng","Patel","Okafor","Silva","Tran"),
city = c("Brisbane","Perth","Sydney",
"Cairns","Hobart"),
stringsAsFactors = FALSE)
orders <- data.frame(
order_id = c("O1","O2","O3","O4","O5","O6","O7"),
customer_id = c("C1","C1","C2","C3","C3","C3","C9"),
product_id = c("P1","P2","P2","P3","P4","P1","P2"),
qty = c(1, 2, 1, 5, 1, 3, 2),
stringsAsFactors = FALSE)
# is the key unique on each side?
anyDuplicated(customers$customer_id) #> 0
anyDuplicated(orders$customer_id) #> 2
# which keys sit on only one side?
setdiff(orders$customer_id,
customers$customer_id) #> "C9"
setdiff(customers$customer_id,
orders$customer_id) #> "C4" "C5"
# Run these three lines before every join.inner <- merge(orders, customers, by = "customer_id")
left <- merge(orders, customers, by = "customer_id",
all.x = TRUE)
right <- merge(orders, customers, by = "customer_id",
all.y = TRUE)
full <- merge(orders, customers, by = "customer_id",
all = TRUE)
c(inner = nrow(inner), left = nrow(left),
right = nrow(right), full = nrow(full))
#> inner left right full
#> 6 7 8 9
# all.x keeps every row of the FIRST argument
# all.y keeps every row of the SECOND
# all keeps every row of both
# Which one is right depends on the question:
# "orders with their customer details" -> left
# "orders that we can attribute" -> inner
# "every customer, ordering or not" -> right
# "a complete audit of both tables" -> fullmerge(x, y,
by = "customer_id", # shared key name
all.x = FALSE, # keep unmatched x?
all.y = FALSE, # keep unmatched y?
suffixes = c(".x", ".y"))
# keys with different names
merge(orders, cust2,
by.x = "customer_id", by.y = "id")
# more than one key column
merge(a, b, by = c("year", "state"))
# no 'by' at all: merge uses EVERY shared column
# name, which is rarely what you meantsort = FALSE. The row order
of your inputs is not preserved.by argument, merge joins on every shared column name at
once. Name the key explicitly, every time.library(dplyr)
inner_join(orders, customers, by = "customer_id")
left_join (orders, customers, by = "customer_id")
right_join(orders, customers, by = "customer_id")
full_join (orders, customers, by = "customer_id")
# keys with different names
orders %>%
inner_join(cust2, by = c("customer_id" = "id"))
# suffixes
orders %>%
inner_join(customers, by = "customer_id",
suffix = c("_order", "_customer"))
# the two that return no new columns
semi_join(orders, customers, by = "customer_id")
anti_join(orders, customers, by = "customer_id")by prints a message naming the columns it chose, rather than
proceeding in silence.semi_join and anti_join have no base R function of their own. In
base R they are %in% on a single column.for (nm in c("inner", "left", "right", "full")) {
z <- get(nm)
cat(sprintf("%-6s rows %d NA cells %d\n",
nm, nrow(z), sum(is.na(z))))
}
#> inner rows 6 NA cells 0
#> left rows 7 NA cells 2
#> right rows 8 NA cells 6
#> full rows 9 NA cells 8
# the habit worth forming
before <- nrow(orders)
after <- nrow(left)
stopifnot(after == before) # a left join on a
# unique right key
# cannot change the
# row count
# stopifnot() turns an assumption into an error
# instead of a wrong number further downstream.full <- merge(orders, customers,
by = "customer_id", all = TRUE)
full[!complete.cases(full), ]
#> customer_id order_id product_id qty name city
#> 7 C4 <NA> <NA> NA Silva Cairns
#> 8 C5 <NA> <NA> NA Tran Hobart
#> 9 C9 O7 P2 2 <NA> <NA>
sum(is.na(full)) #> 8
sum(!complete.cases(full)) #> 3
# Which side failed to supply the row?
full$missing_customer <- is.na(full$name)
full$missing_order <- is.na(full$order_id)
# Two questions, two different answers:
# "how many orders can we attribute?" -> 6
# "how many customers have we served?" -> 3
# Neither is nrow(full).# customers stores it as 'id', orders as
# 'customer_id'
cust2 <- customers
names(cust2)[1] <- "id"
# option 1: rename first
names(cust2)[1] <- "customer_id"
merge(orders, cust2, by = "customer_id")
# option 2: name both sides in the call
merge(orders, cust2,
by.x = "customer_id", by.y = "id")
#> 6 rows, key column named customer_id
# dplyr
# inner_join(orders, cust2,
# by = c("customer_id" = "id"))
a <- data.frame(k = 1:3, v = c(10, 20, 30))
b <- data.frame(k = 2:4, v = c(200, 300, 400))
merge(a, b, by = "k")
#> k v.x v.y
#> 1 2 20 200
#> 2 3 30 300
merge(a, b, by = "k",
suffixes = c("_left", "_right"))
#> k v_left v_right
#> 1 2 20 200
#> 2 3 30 300
.x and .y, which say nothing about where
each column came from. Three joins later you will be reading v.x.y and guessing.
Name the suffixes after the tables.# join twice, on two different keys
oc <- merge(orders, customers, by = "customer_id")
ocp <- merge(oc, products, by = "product_id")
nrow(orders) #> 7
nrow(oc) #> 6 O7 lost: customer C9 unknown
nrow(ocp) #> 6 every product_id matched
ocp$revenue <- ocp$qty * ocp$price
ocp[order(ocp$order_id),
c("order_id","name","item","qty","price","revenue")]
#> order_id name item qty price revenue
#> O1 Ng Keyboard 1 89 89
#> O2 Ng Monitor 2 340 680
#> O3 Patel Monitor 1 340 340
#> O4 Okafor Cable 5 15 75
#> O5 Okafor Dock 1 210 210
#> O6 Okafor Keyboard 3 89 267
aggregate(revenue ~ name, ocp, sum)
#> name revenue
#> Ng 769
#> Okafor 552
#> Patel 340
7 orders became 6 at the first join and stayed 6 at the second. That single lost row is order O7.
ocp <- orders |>
merge(customers, by = "customer_id") |>
merge(products, by = "product_id")merge(a, b) is called with no by argument. What does R join on?date as well as the key you intended, the join silently requires both to match, and your result quietly shrinks.Joins fail quietly. They return a data frame either way, and the number of rows in it is the only warning you get.
visits <- data.frame(
customer_id = c("C1","C1","C3","C3","C3"),
visit = c("V1","V2","V3","V4","V5"),
stringsAsFactors = FALSE)
fan <- merge(orders, visits, by = "customer_id")
nrow(orders) #> 7
nrow(visits) #> 5
nrow(fan) #> 13
table(fan$customer_id)
#> C1 C3
#> 4 9
# 2 x 2 = 4 and 3 x 3 = 9
# Detect it before it happens
anyDuplicated(orders$customer_id) > 0 #> TRUE
anyDuplicated(visits$customer_id) > 0 #> TRUE
# Duplicates on BOTH sides is the dangerous case.
# If you only need one row per customer, aggregate
# first and join the summary
vc <- aggregate(visit ~ customer_id, visits, length)
nrow(merge(orders, vc, by = "customer_id")) #> 6ocp <- merge(merge(orders, customers,
by = "customer_id"),
products, by = "product_id")
ocp$revenue <- ocp$qty * ocp$price
sum(ocp$revenue) #> 1661
# what the orders table actually contains
p <- products$price[match(orders$product_id,
products$product_id)]
sum(orders$qty * p) #> 2341
# the gap
1 - 1661 / 2341 #> 0.2904
# the check that would have caught it
nrow(orders) #> 7
nrow(ocp) #> 6
stopifnot(nrow(ocp) == nrow(orders))
#> Error: nrow(ocp) == nrow(orders) is not TRUE
# A left join keeps the row and makes the problem
# visible as an NA instead of hiding it
ocl <- merge(orders, customers,
by = "customer_id", all.x = TRUE)
sum(is.na(ocl$name)) #> 1# semi: order rows that have a customer
sum(orders$customer_id %in%
customers$customer_id) #> 6
# anti: order rows that do not
sum(!orders$customer_id %in%
customers$customer_id) #> 1
orders[!orders$customer_id %in%
customers$customer_id, ]
#> order_id customer_id product_id qty
#> 7 O7 C9 P2 2
# anti, the other way: customers who never ordered
customers[!customers$customer_id %in%
orders$customer_id, ]
#> customer_id name city
#> 4 C4 Silva Cairns
#> 5 C5 Tran Hobart
# the same three with set functions
length(intersect(orders$customer_id,
customers$customer_id)) #> 3 keys
setdiff(orders$customer_id,
customers$customer_id) #> "C9"
setdiff(customers$customer_id,
orders$customer_id) #> "C4" "C5"check_join <- function(x, y, by) {
cat("left rows :", nrow(x), "\n")
cat("right rows :", nrow(y), "\n")
cat("dup key in left :",
anyDuplicated(x[[by]]) > 0, "\n")
cat("dup key in right :",
anyDuplicated(y[[by]]) > 0, "\n")
cat("keys in left only:",
length(setdiff(x[[by]], y[[by]])), "\n")
cat("keys in right only:",
length(setdiff(y[[by]], x[[by]])), "\n")
invisible(NULL)
}
check_join(orders, customers, "customer_id")
#> left rows : 7
#> right rows : 5
#> dup key in left : TRUE
#> dup key in right : FALSE
#> keys in left only: 1
#> keys in right only: 2
| dup in right FALSE | A left join cannot change the row count. Assert it with
stopifnot. |
| dup in both TRUE | Fan-out is coming. Aggregate one side first, or accept that the result is a list of pairings and not a list of orders. |
| keys in left only | Rows an inner join will delete. Here, 1. |
| keys in right only | Rows a right or full join will add, filled with NA. Here, 2. |
before <- nrow(orders)
res <- merge(orders, customers,
by = "customer_id", all.x = TRUE)
stopifnot(nrow(res) == before)
An assertion that fails is a good afternoon. A wrong total that never fails is a bad quarter.A join matches on a key. rbind stacks tables that already have the same
columns, which is what you want when the tables are two months of the same thing rather than
two facts about the same entity.
may <- airquality[airquality$Month == 5, ]
june <- airquality[airquality$Month == 6, ]
both <- rbind(may, june)
nrow(may); nrow(june); nrow(both)
#> 31
#> 30
#> 61
# rbind demands identical column names in the same
# order, and errors if they differ. That strictness
# is the feature.
# many tables at once
parts <- split(airquality, airquality$Month)
nrow(do.call(rbind, parts)) #> 153
cbind(a, b) # glues columns side by side
# and matches NOTHING
cbind pairs row 1 with row 1 and row 2 with row 2, whatever those rows contain. If
either table has been sorted, filtered or aggregated since you last looked, the pairing is
wrong and the result still prints. Use a join whenever a key exists.| merge | Same entities, different facts, linked by a key. |
| rbind | Different entities, same facts, same columns. |
| cbind | Only when the rows are already aligned and you can prove it. |
anyDuplicated() on both key columns first.cbind a safe way to combine two data frames?Join, then group, then summarise, then draw. Four steps, four row counts, and one figure at the end of it.
USArrests holds crime rates per state and no geography. R's
state.* vectors hold geography and no crime rates. Neither answers a question about
regions until they are joined.
# USArrests keeps the state in the row names, not
# in a column. A key has to be a column.
arrests <- data.frame(state = rownames(USArrests),
USArrests,
row.names = NULL,
stringsAsFactors = FALSE)
head(arrests, 3)
#> state Murder Assault UrbanPop Rape
#> 1 Alabama 13.2 236 58 21.2
#> 2 Alaska 10.0 263 48 44.5
#> 3 Arizona 8.1 294 80 31.0
# the four state vectors are parallel, so a data
# frame is one call
meta <- data.frame(
state = state.name,
region = as.character(state.region),
division = as.character(state.division),
area = state.area,
stringsAsFactors = FALSE)
check_join(arrests, meta, "state")
#> left rows : 50
#> right rows : 50
#> dup key in left : FALSE
#> dup key in right : FALSE
#> keys in left only: 0
#> keys in right only: 0
sj <- merge(arrests, meta, by = "state")
nrow(sj) #> 50
USArrests and mtcars both store
their identifier there, and a row name is not something you can join on. Every analysis that
starts from one of these datasets starts with that line.setdiff() in both directions before you blame the join.ag <- aggregate(cbind(Murder, Assault,
UrbanPop, Rape) ~ region,
sj, mean)
ag[, 2:5] <- round(ag[, 2:5], 3)
ag$n <- as.vector(table(sj$region)[ag$region])
ag
#> region Murder Assault UrbanPop Rape n
#> North Central 5.700 120.333 64.417 18.442 12
#> Northeast 4.700 126.667 70.556 13.778 9
#> South 11.706 220.000 59.438 21.163 16
#> West 7.031 187.231 70.615 29.054 13
sum(ag$n) #> 50 every state accounted for
barplot(setNames(ag$Murder, ag$region),
col = "#CC0000", las = 1,
ylab = "Mean murder rate per 100,000")
# cbind(...) ~ region aggregates four columns in
# one call. Without it you would write four
# separate aggregate() calls and bind the results.# a region table covering only the lower 48
cont <- meta[!meta$state %in%
c("Alaska", "Hawaii"),
c("state", "region")]
nrow(cont) #> 48
inn <- merge(arrests, cont, by = "state")
lef <- merge(arrests, cont, by = "state",
all.x = TRUE)
nrow(inn) #> 48
nrow(lef) #> 50
sum(is.na(lef$region)) #> 2
lef$state[is.na(lef$region)]
#> "Alaska" "Hawaii"
# what the two lost states did to the answer
mean(inn$Murder) #> 7.7938
mean(arrests$Murder) #> 7.7880
Dropping Alaska and Hawaii moved the overall mean murder rate by 0.006, which is nothing. Their rates, 10.0 and 5.3, sit either side of the average and cancel.
A left join makes the loss visible as two NAs. An inner join makes it invisible. Prefer the version that leaves evidence.
USArrests by region have to start with data.frame(state = rownames(USArrests), USArrests)?check_join reports no duplicate keys and no keys on only one side. What follows?Both halves of this lesson end in the same place. State what you grouped on, state which rows you kept, and give the count.
| What you want | Use | Because |
|---|---|---|
| One row per group, for a table or a chart | aggregate, or group_by + summarize | The original rows are no longer needed once you have the group statistic. |
| Every original row, plus its group's statistic | ave, or group_by + mutate | Centring, ranking within group, and share-of-group all need the original rows. |
| Every order, with customer details attached | left join, orders on the left | Orders is the table you cannot afford to lose rows from. |
| Only the orders you can fully attribute | inner join, with the dropped count reported | The choice is defensible; hiding how many rows it cost is not. |
| Every customer, including those who never ordered | right join, or left join with customers first | Customers with no orders are the finding, not an inconvenience. |
| A full audit of what matches and what does not | full join, then complete.cases | Both kinds of unmatched row appear in one result. |
| Two months of the same table, stacked | rbind | There is no key to match on; the rows are new entities, not new facts. |
| The error | Why it happens | What to do instead |
|---|---|---|
| Reporting a group mean with no count | The output of aggregate looks complete, and the June column of
airquality rests on 9 days rather than 30 without saying so. |
Return length alongside every statistic, in the same table. |
| Grouping on one variable when two matter | The marginal mean is the natural first thing to compute, and it can report an ordering that reverses inside the cells. | Group by the second variable too and check the ordering survives. Draw the interaction plot. |
| Joining without checking the key | A join returns a data frame whether it lost 30 per cent of your rows or multiplied them by four. | Run the six lines of check_join, then assert the expected row count with
stopifnot. |
| Using cbind where a join belongs | Two frames have the same number of rows, so gluing them looks safe. | Join on the key. If there is no key, establish why the rows are aligned before you rely on it. |
aggregate returns a data frame, tapply a vector,
table counts, ave broadcasts a group statistic back to every
row.ChickWeight, grouping on
Diet and a banded Time. Report n, mean and sd for every cell, and
check whether any ordering reverses.check_join on arrests against a version of
meta from which you have deleted five states at random. Predict all four row
counts before you run the joins, then run them.state.area. Say which of the two answers your claim needs.Every chart in the next two weeks takes a data frame in the shape this lesson produces: one row per observation, one column per variable, with the grouping variable present as a column rather than implied by the layout.
ChickWeight cell table, and one sentence on whether the diet ordering held at
every time band.Press T or Escape to close