Lecture slides

Descriptive Statistics and Exploratory Data Analysis

Descriptive Statistics and Exploratory Data Analysis

Say what you found, in the plainest sense

Chapter 6

Polla Fattah

By the end of today you can

  • explain the place of exploratory analysis in research;
  • choose a summary that suits the level and shape of a variable;
  • summarise centre and spread, and explain the standard deviation;
  • describe shape, and check it against the normal distribution;
  • find unusual values, and decide what to do with them;
  • describe how much data is missing, and why;
  • interpret correlations, and why correlation is not causation;
  • describe a sample in a thesis.

Description is not a formality

  • the shape of a variable decides which tests are appropriate;
  • an unusual value can dominate an analysis;
  • missing data can quietly change who the results are about.

Descriptive statistics are the first results in almost every thesis.

Three kinds of analysis

Kind Does Chapters
exploratory gets to know the data; proves nothing 6
inferential draws conclusions about a population 7 to 10
predictive predicts new cases 11 to 16

Exploration comes first

flowchart TB
  subgraph R1[" "]
    direction LR
    A[Question] --> B[Import] --> C[Clean]
  end
  subgraph R2[" "]
    direction LR
    D[Explore] --> E[Test or predict] --> F[Report]
  end
  R1 --> R2

Return to exploration whenever a later result surprises you.

Three measures of the centre

Measure Definition Useful for
mean the average: sum divided by count symmetrical numbers
median the middle value when sorted skewed numbers, ordinal data
mode the most common value categories

\[ \bar{x} = \frac{1}{n} \sum_{i=1}^{n} x_i \]

One heavy user moves the mean

caffeine <- c(0, 120, 150, 180, 900)
mean(caffeine)
#> [1] 270
median(caffeine)
#> [1] 150

The mean is higher than four of the five students.

The mean is sensitive to extreme values; the median is not.

Mean and median in the study

first_sem |>
  summarise(mean_sleep      = mean(sleep_hours, na.rm = TRUE),
            median_sleep    = median(sleep_hours, na.rm = TRUE),
            mean_caffeine   = mean(caffeine_mg, na.rm = TRUE),
            median_caffeine = median(caffeine_mg, na.rm = TRUE))
Mean Median
sleep (hours) 6.48 6.5
caffeine (mg) 189 160

Why caffeine’s mean is higher

A minority of heavy users pull the mean (solid) above the median (dashed).

Choosing a summary

Variable Centre Spread Examples
nominal mode, percentages none faculty, gender
ordinal median, percentages IQR financial worry, one item
numeric, symmetrical mean SD sleep, wellbeing
numeric, skewed median IQR caffeine, age

The level of measurement comes first, then the shape.

Same average, different stories

group_a <- c(6.0, 6.5, 6.5, 7.0, 6.5)
group_b <- c(4.0, 8.5, 5.0, 9.0, 6.0)

Both average 6.5 hours.

In group A everyone sleeps about the same; group B ranges from 4 to 9 hours.

A mean without a measure of spread hides half the story.

Deviations from the mean

Each copper line is a deviation, \(x_i - \bar{x}\). The standard deviation summarises them.

The standard deviation

\[ s = \sqrt{\frac{1}{n-1} \sum_{i=1}^{n} (x_i - \bar{x})^2} \]

  • deviations add to zero, so they are squared first;
  • the average squared deviation is the variance;
  • its square root returns to the original units.

Dividing by \(n - 1\) gives a better estimate from a sample.

The SD is roughly the typical distance

sd(group_a)                          #> 0.35
sd(group_b)                          #> 2.18
mean(abs(group_b - mean(group_b)))   #> 1.80

Group B’s SD is close to the average length of the copper lines.

It is a little larger, because squaring weights the longest lines more.

The interquartile range

quantile(first_sem$caffeine_mg, c(0.25, 0.5, 0.75), na.rm = TRUE)
#> 25% 50% 75%
#> 105 160 240
IQR(first_sem$caffeine_mg, na.rm = TRUE)
#> [1] 135

The spread of the middle half: \(Q_3 - Q_1\). It ignores the extremes, like the median.

A box plot’s box is exactly the IQR.

Pairs that belong together

Centre Spread
mean standard deviation
median interquartile range

Sleep: SD ≈ 1 hours, so a typical student sleeps within about an hour of 6.5.

The shape of a distribution

Shape Meaning Mean and median
symmetrical two sides mirror each other about equal
skewed right a long tail of high values mean above median
skewed left a long tail of low values mean below median
bimodal two peaks often two groups mixed

Sleep is symmetrical, caffeine is not

scales = "free" gives each panel its own axes: hours and milligrams differ greatly.

Skewness puts a number on it

library(psych)
skew(first_sem$sleep_hours, na.rm = TRUE)   #> -0.10
skew(first_sem$caffeine_mg, na.rm = TRUE)   #>  1.74
skew(students$age, na.rm = TRUE)            #>  1.34

0 is symmetrical, positive a tail to the right. Between −1 and +1 is mild.

Most graduate students are in their twenties, fewer are older.

The normal distribution

Symmetrical and bell-shaped, described completely by its mean and SD.

Many tests assume data, or averages of data, follow it approximately (Chapter 7).

Within Share of values
1 SD of the mean about 68%
2 SD about 95%
3 SD about 99.7%

Checking sleep against the rule

z <- (sleep - mean(sleep)) / sd(sleep)
c(within_1_sd = mean(abs(z) < 1),
  within_2_sd = mean(abs(z) < 2),
  within_3_sd = mean(abs(z) < 3))
#> 0.707  0.958  0.995

Very close to 68-95-99.7.

z-scores give each value’s distance from the mean in standard deviations.

Q-Q plots compare with the normal

Sleep sits on the line. Caffeine bends away at the top: its long right tail.

Outliers are information first

Impossible values were set to NA in Chapter 3. Outliers are possible, just unusual.

An unusual value may be:

  • an error that cleaning missed;
  • a participant who misunderstood a question;
  • a real case the research should pay attention to.

Find them and look at them before anything else.

Two rules for flagging

Rule Flags values Suits
box plot rule beyond 1.5 × IQR from the quartiles skewed variables
z-score rule more than 3 SD from the mean roughly normal variables

The box plot rule for caffeine

q1 <- quantile(first_sem$caffeine_mg, 0.25, na.rm = TRUE)
q3 <- quantile(first_sem$caffeine_mg, 0.75, na.rm = TRUE)
upper_fence <- q3 + 1.5 * (q3 - q1)
sum(first_sem$caffeine_mg > upper_fence, na.rm = TRUE)
#> [1] 35

35 students take more than the upper fence of 443 mg.

For sleep, roughly normal, the z-score rule flags 3 students.

The most extreme cases

first_sem |>
  filter(sleep_hours <= 4.5, caffeine_mg >= 600) |>
  select(student_id, sleep_hours, caffeine_mg, study_hours, wellbeing)

Nine students sleep 3.5 to 4.3 hours, take 610 to 900 mg of caffeine, and study up to 67 hours a week.

Nothing impossible: a real and worrying group, exactly what the thesis is about.

Never delete a value just because it is unusual

  1. Check whether it is an error (Chapter 3).
  2. If it is real, keep it and report it.
  3. Choose summaries that are not thrown off by it, such as the median.
  4. Check whether conclusions change without it: a sensitivity analysis.

Removing inconvenient values for a cleaner result is research misconduct.

Count the gaps first

semesters |> summarise(across(everything(), ~ sum(is.na(.x))))
students  |> summarise(across(everything(), ~ sum(is.na(.x))))

About 1% of each semester measurement is missing.

38 students did not answer the question about money worries.

Why data is missing matters more than how much

Kind The gaps depend on Consequence
missing completely at random nothing unbiased, a little less precise
missing at random something measured correctable by including it
missing not at random the missing value itself the answers left are not typical

Students most worried about money may be the least willing to say so.

The biggest gap: students who left

left_ids <- setdiff(students$student_id,
                    semesters$student_id[semesters$semester == 3])
first_sem |>
  mutate(left_study = student_id %in% left_ids) |>
  summarise(students = n(), wellbeing = mean(wellbeing),
            considering_dropout = mean(considering_dropout == "Yes"),
            .by = left_study)
Students Wellbeing Considered dropout
stayed 563 60.7 12%
left 37 57.4 59%

What attrition does to the later semesters

The students still present in semesters 3 and 4 were, on average, doing better.

A simple average of the later semesters overestimates how well students were doing.

  • report how many left and how they differed;
  • mixed models (Chapter 10) use every observation each student gave;
  • multiple imputation exists for large gaps.

The correlation coefficient

\(r\) measures how closely two numeric variables follow a straight line, from −1 to +1.

\(r\) (either sign) Often called
about 0.1 small
about 0.3 medium
about 0.5 large

Rough guides, not rules. A curved relationship can have \(r\) near zero.

A correlation matrix

first_sem |>
  select(stress, burnout, support, wellbeing, gpa) |>
  cor(use = "pairwise.complete.obs") |> round(2)
burnout support wellbeing GPA
stress 0.65 −0.33 −0.53 −0.31
support −0.24 1 0.36 0.36

pairwise.complete.obs uses every student with both values for each pair.

A scatter plot matrix

pairs() draws one small scatter plot for each pair of variables.

Correlation is not causation

first_sem |>
  select(caffeine_mg, sleep_hours, gpa) |>
  cor(use = "pairwise.complete.obs") |> round(2)

flowchart LR
    S["Less sleep"] --> C["More caffeine"]
    S --> G["Lower GPA"]
    C -. "correlation −0.18" .- G

Caffeine may only look harmful because heavy users sleep less: a confounder (Chapter 8).

Describing the sample

first_sem |>
  summarise(students      = n(),
            female_pct    = 100 * mean(gender == "Female"),
            phd_pct       = 100 * mean(programme == "PhD"),
            part_time_pct = 100 * mean(study_mode == "Part-time"),
            age_mean      = mean(age, na.rm = TRUE),
            age_sd        = sd(age, na.rm = TRUE)) |>
  round(1)

Who took part: the first table in the methods chapter.

Every summary needs its spread and its n

A mean of 6.4 hours says little until the reader knows:

  • whether students differ by minutes or by hours;
  • whether it comes from 20 students or 600.

With missing values, the n differs from variable to variable, so report it for each.

A description table for the main variables

describe_var <- function(x) {
  c(n = sum(!is.na(x)),
    mean = mean(x, na.rm = TRUE), sd = sd(x, na.rm = TRUE),
    median = median(x, na.rm = TRUE), iqr = IQR(x, na.rm = TRUE),
    missing = sum(is.na(x)))
}
sapply(first_sem[, c("sleep_hours", "study_hours", "caffeine_mg",
                     "gpa", "wellbeing")], describe_var) |> t() |> round(2)

sapply() applies the function to each column; t() makes each variable a row.

Writing it up

The sample consisted of 600 graduate students (52% female; 29% PhD students), aged 23 to 52 (M = 29.7, SD = 4.8). Students slept 6.5 hours per night on average (SD = 1.0), and 65% slept less than 7 hours.

Every number comes from the code.

An exploration checklist

  1. How many values are missing, and why?
  2. What is the typical value, and how spread out are the values?
  3. What is the shape: symmetrical or skewed, one peak or two?
  4. Are there unusual values, and are they errors or real?
  5. How does it relate to the other variables?

In your field: earth sciences

Old Faithful: eruptions last about 2 or about 4.5 minutes. The mean, 3.5, describes almost none of them. Report the two groups separately.

Practical lab: the Chapter 6 playground

Work through the playground exercises in your browser, with hints and solutions.

Every exercise also runs in RStudio, from the downloadable chapter project.

Practical exercises 1–3: summaries and shape

  1. Mean, median, SD, and IQR of study hours: which pair would you report?
  2. A Q-Q plot of wellbeing: does it look normal?
  3. Count unusually high study hours with the box plot rule, and look at those students.

Practical exercises 4–6: gaps and relationships

  1. Compare the sleep of students who left with those who stayed.
  2. Classify the support–satisfaction correlation, and explain why it is not causal.
  3. Choose and justify a summary for every variable in students.

Try this yourself

Take three variables from your own data.

  • report centre, spread, n, and missing for each;
  • draw each one’s shape and name it;
  • flag unusual values and decide whether they are errors;
  • write two sentences describing your sample.

Troubleshooting guide (Part 1)

Symptom Likely cause
the mean seems too high a skewed variable; report the median
a mean with nothing beside it spread and n left out
many outliers in a skewed variable the z-score rule used on skewed data
a mean describes no real case a bimodal variable

Troubleshooting guide (Part 2)

Symptom Likely cause
later semesters look better students who were struggling left
r near zero for a clear pattern a curved relationship
a strong correlation read as cause a confounder not considered
different n for each variable missing values; report n per variable

Completion checklist

Misconceptions to leave behind (Part 1)

Misconception Better mental model
the mean is the typical value only for roughly symmetrical variables
a mean on its own is a result it needs its spread and n
outliers should be removed real values are reported and handled

Misconceptions to leave behind (Part 2)

Misconception Better mental model
a few missing values do not matter why they are missing matters more
r = 0 means no relationship no straight-line relationship
correlation shows cause a confounder can create it

The chapter in one sentence

Describe every variable by its centre, spread, shape, gaps, and unusual values before you test anything.

Next: Chapter 7

The next chapter moves from describing the sample to inference:

  • sampling distributions and standard errors;
  • confidence intervals and the bootstrap;
  • the logic of a test, built by shuffling;
  • t-tests, effect sizes, and chi-square tests.

Questions

Which variable in your own data is most likely to be skewed?

Who is most likely to be missing from your data, and why?