Playground: Chapter 4

Data Visualization

This page practises the ideas of Chapter 4 in four parts. Practise the chapter works through the book’s exercises, with starter code, hints, and solutions. Go further goes beyond the book, with new questions and some of R’s built-in datasets. Check your understanding asks short questions whose answers open with a click. Do it yourself gives open tasks with an empty code space and no answers: you decide which graph answers the question, as you will in your own thesis.

The first run on this page takes a little longer, because your browser downloads the ggplot2 and dplyr packages. In the exercises, replace each ______ with your own code and press Run Code.

Practise the chapter

These are the exercises at the end of Chapter 4, with the same numbers.

Exercise 1: Choosing a binwidth

Draw a histogram of study_hours in the first semester. Try binwidths of 1, 5, and 10, and explain which shows the shape best.

NoteHint

The binwidth is the width of each bar in hours. Run the code three times with different values and compare the pictures.

TipSolution
ggplot(first_sem, aes(x = study_hours)) +
  geom_histogram(binwidth = 5, colour = "white")

A binwidth of 5 hours shows the shape best: most students study between about 15 and 40 hours a week, the peak is in the mid-twenties, and a long tail stretches to the right, up to 67 hours. With a binwidth of 1, the bars are jagged and the shape drowns in noise; with 10, only a few bars remain, and the tail almost disappears. There is no correct binwidth, so try several before choosing. (R warns that 8 rows were removed: those students have no study hours recorded.)

Exercise 2: Employment in order, in one colour

Draw a bar chart of how many students have each level of employment. Make the bars a single colour of your choice, and order the levels from no job to a full-time job.

NoteHint

The levels are "None", "Part-time job", and "Full-time job", written in the order you want inside c(). A fixed colour is set outside aes(), as a colour name in quotation marks such as "steelblue".

TipSolution
students |>
  mutate(employment = factor(employment,
                             levels = c("None", "Part-time job", "Full-time job"))) |>
  ggplot(aes(x = employment)) +
  geom_bar(fill = "steelblue") +
  labs(x = "Paid work", y = "Number of students")

307 students have no job, 196 a part-time job, and 97 a full-time job. Employment is an ordered variable, so the bars should follow its order, not the alphabet, which would put “Full-time job” first. Because the colour is set outside aes(), all bars share it and no legend appears.

Exercise 3: GPA by faculty

Draw box plots of first-semester GPA for each faculty, with clear axis labels.

NoteHint

The geom for box plots is named after them. A label says what was measured, in words a reader outside the study understands, with units where there are any.

TipSolution
ggplot(first_sem, aes(x = faculty, y = gpa)) +
  geom_boxplot() +
  labs(x = "Faculty", y = "GPA in semester 1 (0 to 4)")

The five boxes are almost level, with medians between 3.06 and 3.14, and the boxes overlap almost completely: GPA hardly differs between faculties. Each box holds the middle half of the students, and the points beyond the whiskers are students with unusually high or low GPAs.

Exercise 4: Caffeine and sleep

Draw a scatter plot of caffeine against sleep in the first semester, with a trend line, and describe what it suggests.

NoteHint

alpha runs from 0 (invisible) to 1 (solid); with 600 points, a value around 0.3 to 0.5 shows where they pile up. The method for a straight trend line is a linear model, written as text.

TipSolution
ggplot(first_sem, aes(x = caffeine_mg, y = sleep_hours)) +
  geom_point(alpha = 0.4) +
  geom_smooth(method = "lm") +
  labs(x = "Caffeine per day (mg)", y = "Hours of sleep per night")

The line falls clearly: students who take more caffeine sleep less, by roughly half an hour for every extra 100 mg, and the points follow the line more closely than in most relationships in the study. The graph cannot say which causes which: caffeine may shorten sleep, or students who sleep little may drink more coffee to get through the day. Chapter 8 returns to this relationship.

Exercise 5: Saving a graph

Take any graph from this chapter and save it as a PNG file, 16 cm wide and 10 cm high, at 300 dpi.

Saving a file needs your own computer, so this exercise is in the download project.

p <- ggplot(first_sem, aes(x = sleep_hours)) +
  geom_histogram(binwidth = 0.5, fill = "steelblue", colour = "white")
ggsave("sleep.png", plot = p, width = 16, height = 10, units = "cm", dpi = 300)

The file appears in the project folder. Setting the size in centimetres, rather than dragging a window, means the text in the graph has a known size on the printed page, and 300 dpi keeps it sharp.

Exercise 6: The misleading graph, redrawn

Redraw the book’s misleading bar chart of wellbeing by faculty as a bar chart whose axis starts at zero, with the faculties in order. Then compare the three versions in the chapter (the misleading one, the redesigned strip plot, and yours) and explain which you would put in a thesis, and why.

NoteHint

reorder(faculty, x) orders the faculties by the values of x. A bar’s length stands for its value only if the axis starts at the bottom of the scale.

TipSolution
ggplot(faculty_means, aes(x = reorder(faculty, wellbeing), y = wellbeing)) +
  geom_col(fill = "grey50") +
  coord_cartesian(ylim = c(0, 100)) +
  labs(x = NULL, y = "Average wellbeing in semester 1 (0 to 100)")

The five bars are now almost the same height, between 58.9 and 61.9, a difference of 3 points on a 100-point scale, and that is the honest picture of the averages. The misleading version made the same 3 points look like a several-fold difference. Of the three, the redesigned strip plot is best for a thesis: like your bar chart, it does not exaggerate, but it also shows every student, and so shows that the differences within each faculty are far larger than the differences between them. A bar chart of averages can never show that.

Go further

These exercises go beyond the book.

Exercise 7: The same numbers three ways

Show the average wellbeing of each faculty as a pie chart, a bar chart, and a dot plot, and decide which makes the comparison easiest.

NoteHint

A bar chart of values already in the data uses the geom for columns, not geom_bar(), which counts rows. A dot plot is points.

TipSolution
ggplot(faculty_means, aes(x = reorder(faculty, wellbeing), y = wellbeing)) +
  geom_col(fill = "grey50") +
  labs(x = NULL, y = "Average wellbeing")

ggplot(faculty_means, aes(x = wellbeing, y = reorder(faculty, wellbeing))) +
  geom_point(size = 3) +
  labs(x = "Average wellbeing", y = NULL)

In the pie chart, the five slices look equal, and a pie of averages makes no sense anyway: slices are parts of a whole, and averages do not add up to anything. The bar chart shows the order, but the differences look tiny, because the bars start at zero. The dot plot makes the comparison easiest: the dots are positions on a common scale, the axis can zoom in without distorting anything (dots, unlike bars, do not represent values by length), and the two groups of faculties, around 59 and around 62, stand out at once.

Exercise 8: A line chart, zoomed and not

Draw the chapter’s line chart of average wellbeing by semester and workshop group twice: once as ggplot2 draws it, and once with the y axis from 0 to 100. Write a one-sentence caption for each.

NoteHint

The two numbers are the ends of the wellbeing scale.

TipSolution
p + coord_cartesian(ylim = c(0, 100))

Zoomed in, the axis runs only from about 58 to 66, and the invited group’s rise in semester 2 looks dramatic. On the full scale, the same rise, about 5 points, looks modest. Both are honest if the caption says what the reader sees. For example: “Average wellbeing by semester and workshop group; note that the y axis shows only part of the 0-to-100 scale”, and “Average wellbeing on the full 0-to-100 scale: the invited group’s rise in semester 2 is about 5 points.” Lines, unlike bars, may be zoomed, because they represent values by position; the reader only needs to be told.

Exercise 9: Anscombe’s quartet yourself

Calculate the mean and standard deviation of y and the correlation of x and y in each of Anscombe’s four datasets, then plot the four in separate panels.

NoteHint

The function for a correlation is a short form of its name. The function that splits a graph into panels, one for each value of a variable, wraps them into rows.

TipSolution
anscombe_long |>
  summarise(mean_y = mean(y), sd_y = sd(y), r = cor(x, y), .by = set)

ggplot(anscombe_long, aes(x = x, y = y)) +
  geom_point() +
  geom_smooth(method = "lm", se = FALSE) +
  facet_wrap(~ set)

All four sets have a mean of 7.50, a standard deviation of 2.03, and a correlation of 0.82, to two decimal places, and the same trend line. The panels show four completely different stories: a noisy straight line (I), a smooth curve (II), a perfect line with one outlier (III), and no relationship at all except for one extreme point (IV). Only the first is described well by the numbers. Always plot the data before summarising it.

Exercise 10: Exercise and wellbeing

Graph the relationship between the number of days a week students exercise and their wellbeing in the first semester. Exercise days take only the values 0 to 7, so ordinary points would pile up in eight columns: spread them a little sideways, make them transparent, and add a trend line.

NoteHint

The geom that adds a little random noise to each point’s position was used in the book’s redesigned graph. height = 0 keeps the wellbeing values exact and spreads the points only sideways.

TipSolution
ggplot(first_sem, aes(x = exercise_days, y = wellbeing)) +
  geom_jitter(width = 0.2, height = 0, alpha = 0.3) +
  geom_smooth(method = "lm") +
  labs(x = "Days of exercise per week", y = "Wellbeing in semester 1 (0 to 100)")

The trend line rises gently: students who exercise more days report somewhat higher wellbeing, but the points spread widely at every number of days, so the relationship is weak (the correlation, from Chapter 6, is about 0.2). Few students exercise six or seven days, so the right end of the line rests on little data. Jittering changed only where the points are drawn, not the data; with height = 0, every wellbeing value is still exact. (R warns about 2 removed rows: students without an exercise value.)

Check your understanding

Answer each question in your own words first, then click to see a model answer.

1. Why do researchers prefer a bar chart or dot plot to a pie chart?

Because people judge positions and lengths on a common scale far more accurately than angles and areas. In a bar chart or dot plot, the values sit side by side on the same axis, so small differences and the order of the groups are visible at once; in a pie chart, similar slices look alike, and ranking them needs a legend.

2. When is a y axis that does not start at zero acceptable?

When the values are shown by position, as in a line chart, scatter plot, or dot plot, and the reader is told that the axis is zoomed. It is never acceptable for bars, because a bar shows its value by its length: with a shortened axis, a bar twice as long no longer means twice as much.

3. What is the difference between mapping a colour and setting a colour?

Mapping, inside aes(), links colour to a variable, so each group gets its own colour and ggplot2 adds a legend. Setting, outside aes(), gives every element the same fixed colour. Putting a fixed colour such as "steelblue" inside aes() is a common mistake: ggplot2 treats it as data with one value, colours it pink, and adds a pointless legend.

4. Which kind of palette suits categories, and which suits quantities?

Categories need a qualitative palette: colours that are clearly different but equally prominent, since no category is more important than another. Quantities need a sequential palette, running from light to dark, or a diverging palette, running in two directions away from a meaningful midpoint such as zero. A rainbow palette suits neither, because its bright bands suggest boundaries that are not in the data.

5. What is an exploratory graph for, and how does it differ from an explanatory graph?

An exploratory graph is for the researcher: quick, made in large numbers, and used to find patterns, problems, and surprises in the data. An explanatory graph is for readers: one carefully designed figure that shows one message clearly, with labels, units, an honest axis, a suitable palette, and a caption. Most graphs made during a project are exploratory; only a few become explanatory figures in the thesis.

Do it yourself

These tasks have no starter code and no answers. Each code space below is ready to run: write your own code, as you would for a thesis. The data (students, semesters, and first_sem) and the ggplot2 and dplyr packages are already loaded, and R’s built-in datasets are always available.

1. Choose a question about two variables of the study that the chapter did not graph. Choose the graph with the chapter’s table of graphs by kinds of variables, make an explanatory version ready for a thesis, and write its caption.

2. The code below draws a deliberately poor graph. Improve it, and write down each change you made and why. (This is the only code space on the page with code in it: the task is to change it.)

3. R’s built-in ChickWeight data records the weight of 50 chicks on four diets, measured every few days from birth to day 21. Make a figure for a report that shows how weight grows over time on each diet, and write its caption. Think about what one line should represent, and whether panels or colours compare the diets better.

4. ggplot2’s mpg data records the fuel economy of 234 cars. Make three quick exploratory graphs of the relationship between engine size (displ) and motorway fuel economy (hwy), then turn the most informative one into an explanatory graph.

5. In the download project, save your figure from Task 3 as a PDF file 16 cm wide and 10 cm high, and as a PNG at 300 dpi, and compare the two files when you zoom in.

Work on your own computer

NoteDownload the Chapter 4 project

The project contains the data and all four parts of this page as an R script: the exercises with blanks, the questions to check your understanding (answers in solutions.R), and the open tasks, each with space to write your code.

  • Download chapter04.zip, unzip it, and double-click chapter04.Rproj.
  • Or type this one line in RStudio’s Console:
usethis::use_course("https://polla-fattah.github.io/data2thesis_r/playground/chapter04.zip")