Data Visualization
Data Visualization
A graph is an instrument and an argument
Chapter 4
Polla Fattah
By the end of today you can
- explain why graphs are needed alongside numerical summaries;
- tell exploratory graphs from explanatory ones;
- explain why position and length are read more accurately than angle, area, or colour;
- choose a graph from the question and the variables;
- build graphs with ggplot2 from data, mappings, and layers;
- judge whether a graph is honest, and improve a misleading one;
- export figures at the size and resolution a thesis needs.
Summaries can hide almost anything
anscombe_long |>
summarise(mean_x = mean(x), mean_y = mean(y),
sd_y = sd(y), correlation = cor(x, y), .by = set)| Set | mean x | mean y | sd y | correlation |
|---|---|---|---|---|
| 1 to 4 | 9 | 7.5 | 2.03 | 0.82 |
Anscombe’s four datasets share every summary.
Four identical summaries, four different stories
A line, a curve, an outlier, and one point making the whole trend.
Look first
Anyone who analysed these data without looking would draw the same, wrong, conclusion from all four.
Many errors in data, and in reasoning about data, are found only by looking.
Exploring and explaining
| Exploratory graph | Explanatory graph | |
|---|---|---|
| for | the researcher | the reader |
| purpose | understand the data | show a known finding |
| number | many, quickly | a few, carefully |
| polish | none needed | clear message and labels |
A thesis contains a few explanatory graphs, chosen from many exploratory ones.
A graph turns numbers into visual properties
flowchart LR
A["Position on a<br>common scale"] --> B["Length"] --> C["Angle<br>and slope"] --> D["Area"] --> E["Colour<br>shade"]
Cleveland and McGill’s experiments: judged most accurately on the left, least on the right.
Put the most important comparison into position or length.
The same five numbers, two ways
In the pie the slices look equal. In the bars the decline from A to E is obvious.
Two more rules of thumb
- put values to be compared next to each other on the same scale, within one panel;
- use colour to separate a few categories, not to carry the main message.
Prefer bar charts and dot plots to pie charts, and scatter plots to bubble charts.
Choosing a graph
| Question | Variables | Graph |
|---|---|---|
| how are values distributed? | one numeric | histogram, density |
| how many in each category? | one categorical | bar chart |
| do groups differ? | numeric by categorical | box plot, points |
| how are two variables related? | two numeric | scatter plot |
| how does it change over time? | numeric over time | line chart |
| how do many variables relate? | several numeric | heat map |
The grammar of graphics
flowchart LR
D["Data<br>a data frame"] --> M["Mappings<br>aes(): variable → property"] --> G["Geoms<br>points, bars, lines"]
Every ggplot2 graph combines these three parts, joined with +.
Data and mappings give an empty canvas
A geom draws the data
One row per student
Each student’s first-semester record, with their background information.
Most graphs are simplest with one row per case.
Histograms show a distribution
The binwidth changes the picture
Too narrow is noisy, too wide hides the shape. Try a few widths.
Bar charts count categories
Box plots compare groups
Reading a box plot
| Part | Shows |
|---|---|
| box | the middle half of the values |
| line inside | the median |
| whiskers | the typical range |
| points beyond | unusual values |
Whether the study-mode difference is larger than chance is a question for Chapter 8.
Scatter plots show relationships
Line charts show change over time
Mapping and setting
aes(fill = workshop) # mapped: depends on the data
geom_histogram(fill = "steelblue") # set: the same everywhereInside aes() |
Outside aes() |
|
|---|---|---|
| meaning | depends on a variable | fixed for the whole layer |
| legend | added | none |
The most common ggplot2 mistake
fill and colour
| Aesthetic | Controls |
|---|---|
fill |
the inside of a shape: a bar, a box |
colour |
the outline, or points and lines |
A white colour on histogram bars separates them clearly.
From exploratory to explanatory
flowchart LR
A["Labels"] --> B["Scales and<br>colours"] --> C["Theme"] --> D["Annotations"]
Each addition is a design decision with a reason behind it.
Labels say what was measured
p <- ggplot(wellbeing_by_sem,
aes(x = semester, y = mean_wellbeing, colour = workshop)) +
geom_line(linewidth = 1) +
geom_point(size = 2.5) +
labs(title = "Wellbeing over the two years",
subtitle = "Average score per semester, by workshop group",
x = "Semester", y = "Average wellbeing (0 to 100)",
colour = "Workshop")Units in the axis labels, the message in the title. Storing the graph as p lets you add to it.
Scales and colour-blind-friendly palettes
Whole semesters only: there is no semester 2.5.
Viridis palettes work for colour vision deficiency, about one man in twelve, and print in greyscale.
Colour should never be the only cue
In a printed figure, groups should also differ by:
- shape or line type;
- labels placed directly on the lines.
Manual colours: scale_colour_manual(values = c("Invited" = "#1b9e77", "Not invited" = "#7570b3")).
Themes control everything that is not data
Annotations give a reference
A graph can be accurate and still mislead
Most misleading graphs are not deliberate. They come from defaults nobody questioned.
| Choice | The risk |
|---|---|
| axis range | lengths no longer mean their values |
| aspect ratio | the same line looks steep or flat |
| clutter | decoration competes with the data |
| colour | bright bands suggest boundaries |
Bar axes start at zero
A bar represents its value by its length.
If the axis starts elsewhere, a bar twice as long no longer means twice as much.
Points and lines use position, so their axes may be zoomed, as long as the reader is told.
Zooming a line chart
The wellbeing line chart’s axis runs only from about 58 to 66, which makes a 5-point gap look large.
Either say so in the caption, or show the full scale and let readers judge the size.
Aspect ratio, clutter, and colour
- a graph somewhat wider than tall; the same shape for graphs that will be compared;
- no heavy grid lines, repeated legends, or three-dimensional effects;
- qualitative palettes for categories;
- sequential palettes for quantities in one direction;
- diverging palettes around a meaningful midpoint such as zero.
A rainbow palette is poor for all of these.
A misleading graph
The Education bar looks several times taller than Social Sciences. The averages differ by about 3 points out of 100.
What is wrong with it
- the axis starts at 58, so the bar lengths mean nothing;
- the colours repeat the axis labels, and the legend repeats them again;
- the faculties are alphabetical, not in a meaningful order;
- the axis titles are variable names;
- only averages are shown, hiding how much students differ.
The same data, redesigned
Differences between faculties are small compared with the differences within them.
How the redesign was built
ggplot(first_sem, aes(x = wellbeing, y = reorder(faculty, wellbeing))) +
geom_jitter(height = 0.15, alpha = 0.25, colour = "grey45") +
stat_summary(fun = mean, geom = "point", size = 3.5, colour = "#b2182b") +
labs(x = "Wellbeing in semester 1 (0 to 100)", y = NULL)| Function | Job |
|---|---|
reorder() |
order faculties by their average |
geom_jitter() |
spread points so they do not hide each other |
stat_summary() |
calculate and draw each mean |
Facets: small multiples
Density plots compare distributions
Heat maps show a table of numbers
correlations |>
as.data.frame() |>
mutate(var1 = rownames(correlations)) |>
pivot_longer(-var1, names_to = "var2", values_to = "r") |>
ggplot(aes(x = var1, y = var2, fill = r)) +
geom_tile() +
geom_text(aes(label = round(r, 2))) +
scale_fill_gradient2(limits = c(-1, 1))The matrix becomes a long table with one row per pair (Chapter 3).
Reading the heat map
A diverging palette, white at zero. The numbers matter: colour alone is read imprecisely.
Exporting figures
ggsave("figures/wellbeing-lines.png", plot = p,
width = 16, height = 10, units = "cm", dpi = 300)
ggsave("figures/wellbeing-lines.pdf", plot = p,
width = 16, height = 10, units = "cm")The file type follows from the file name.
Save at the printed size
| Decision | Rule |
|---|---|
| width | about 8 cm for one journal column, 16 cm for a full page |
| resolution | at least 300 dpi for PNG |
| format | PDF or SVG when accepted: sharp at any size |
| text | check it at the final size; adjust base_size |
The figure code lives in the script: a supervisor’s request is one changed line.
In your field: economics and business
Practical lab: the Chapter 4 playground
Work through the playground exercises in your browser, with hints and solutions.
Every exercise also runs in RStudio, from the downloadable chapter project.
Practical exercises 1–3: the basic graphs
- A histogram of study hours with binwidths 1, 5, and 10: which shows the shape best?
- A bar chart of employment in one colour, ordered from no job to full-time.
- Box plots of GPA by faculty, with clear labels.
Practical exercises 4–6: relationships and honesty
- Caffeine against sleep, with a trend line: what does it suggest?
- Save a graph as a PNG, 16 by 10 cm, at 300 dpi.
- Redraw the misleading graph from zero, and choose one of the three for a thesis.
Try this yourself
Pick one variable from your own research.
- make three exploratory graphs of it in five minutes;
- choose the one that answers your question;
- turn it into an explanatory graph: labels, scale, theme, one annotation;
- save it at 16 cm wide.
Then show it to someone for five seconds and ask what they saw.
Troubleshooting guide (Part 1)
| Symptom | Likely cause |
|---|---|
| salmon colour and a strange legend | a fixed colour inside aes() |
| points hide each other | overplotting; use alpha or geom_jitter() |
| categories in alphabetical order | factor levels not set, or no reorder() |
| “semester 2.5” on the axis | a numeric axis without breaks |
Troubleshooting guide (Part 2)
| Symptom | Likely cause |
|---|---|
| small differences look dramatic | a zoomed axis on bars |
| text unreadable in the thesis | saved too large, then shrunk |
| groups vanish in print | colour was the only cue |
| a blurry figure | a PNG below 300 dpi |
Completion checklist
Misconceptions to leave behind (Part 1)
| Misconception | Better mental model |
|---|---|
| summary statistics tell the whole story | Anscombe’s sets share every summary |
| more elements make a graph more informative | decoration competes with the data |
| pie charts are good for proportions | bars or dots are almost always clearer |
Misconceptions to leave behind (Part 2)
| Misconception | Better mental model |
|---|---|
| zooming an axis is always wrong | wrong for bars, acceptable for lines if stated |
| colour makes a graph clearer | colour separates a few categories only |
| a graph is decoration at the end | graphs are one of the main instruments |
The chapter in one sentence
Look at the data first, then choose a graph that puts the key comparison in position or length and tells the truth about its size.
Next: Chapter 5
The next chapter moves from a topic to a plan:
- what data analysis is for;
- from a topic to a research question and a hypothesis;
- variables, measurement, reliability, and validity;
- study designs and randomisation;
- an analysis plan written before the data is seen.
Questions
Think of a graph that convinced you of something.
Would it still convince you with the axis starting at zero?