Playground: Chapter 6

Descriptive Statistics and Exploratory Data Analysis

This page practises the ideas of Chapter 6 in four parts. Practise the chapter works through the book’s exercises, with starter code, hints, and solutions. Go further goes beyond the book, with new questions and some of R’s built-in datasets. Check your understanding asks short questions whose answers open with a click. Do it yourself gives open tasks with an empty code space and no answers: you decide how to describe the data, as you will in your own thesis.

In the exercises, replace each ______ with your own code and press Run Code. Graphs on this page use base R’s quick plotting functions, such as hist() and qqnorm(), which are ideal for exploring.

Practise the chapter

These are the exercises at the end of Chapter 6, with the same numbers.

Exercise 1: Four summaries of study hours

Calculate the mean, median, SD, and IQR of first-semester study_hours. Decide whether the variable is closer to symmetrical or skewed, and which pair of summaries you would report.

NoteHint

Each summary has a function of its own, named after it; the IQR’s name is written in capital letters.

TipSolution
first_sem |>
  summarise(
    mean   = mean(study_hours, na.rm = TRUE),
    median = median(study_hours, na.rm = TRUE),
    sd     = sd(study_hours, na.rm = TRUE),
    iqr    = IQR(study_hours, na.rm = TRUE)
  )

The mean (27.6 hours) is a little above the median (26): study hours are mildly skewed to the right, with a tail of students who study a great deal. The skew is small, so the mean and SD (27.6 and 13.1) are acceptable; the median and IQR (26 and 19) would be the safer pair if the tail were longer. Whichever pair you choose, report the pair: a centre without a spread says little.

Exercise 2: Is wellbeing normal?

Draw a Q-Q plot of first-semester wellbeing, and judge whether it looks normal.

NoteHint

Base R’s Q-Q plot against the normal distribution has a short name: “Q-Q” and “normal” run together. qqline() adds the line the points would follow if the data were exactly normal.

TipSolution
qqnorm(first_sem$wellbeing)
qqline(first_sem$wellbeing)

The points follow the line closely, with only small wiggles at the two ends: wellbeing is close to normal. Real data is never exactly normal; the question is whether it is close enough for methods that assume normality, and here it clearly is. A skewed variable, such as caffeine, would bend away from the line at one end.

Exercise 3: Unusually high study hours

Using the box plot rule, count the students with unusually high study_hours in the first semester, and look at their other values.

NoteHint

The box plot rule flags values more than one and a half IQRs above the third quartile.

TipSolution
fence <- q3 + 1.5 * IQR(x, na.rm = TRUE)

The upper fence is 65.5 hours, and one student is above it: S0312, a part-time student with no job, studies 67 hours a week and sleeps 3.9 hours a night. The values are extreme but consistent with each other (little sleep fits very long study hours), and none is impossible, so this is a real, unusual student rather than an error. The right decision is to keep the case, and perhaps check whether the results change without it.

Exercise 4: Sleep of students who left

Compare the first-semester sleep_hours of students who later left the study with those who stayed, and say whether they differ in sleep as they did in wellbeing.

NoteHint

The students who left are those in the first list but not in the second: the set difference of the two lists.

TipSolution
left_ids <- setdiff(students$student_id,
                    semesters$student_id[semesters$semester == 3])

The 37 students who left slept almost exactly as long in their first semester as the 563 who stayed (6.4 and 6.5 hours), but their wellbeing was more than 3 points lower (57.4 against 60.7). The missing data is related to wellbeing but not to sleep. Attrition must therefore be checked variable by variable: analyses of wellbeing in later semesters describe a group that has lost some of its least happy members, while analyses of sleep are much less affected.

Exercise 5: Support and satisfaction

Find the correlation between support and satisfaction, classify it with Cohen’s guidelines (0.1 small, 0.3 medium, 0.5 large), and explain why it does not show that support causes satisfaction.

NoteHint

A scale score is the average of its items for each student, that is, the mean of each row. The correlation function has a short name.

TipSolution
scores <- questionnaire |>
  mutate(
    support      = rowMeans(pick(support_1:support_6), na.rm = TRUE),
    satisfaction = rowMeans(pick(satisfaction_1:satisfaction_4), na.rm = TRUE)
  )
cor(scores$support, scores$satisfaction, use = "complete.obs")

The correlation is 0.41, a medium correlation by Cohen’s guidelines (above 0.3, but below the 0.5 of a large one): students who feel better supported by their supervisor are more satisfied with their programme. It does not show that support causes satisfaction. Satisfied students may view their supervisors more kindly (the causation may run the other way), and a third variable, such as how well the research is going, may raise both.

Exercise 6: A summary for every variable

For each variable in students, choose the summary you would report in a thesis, using the chapter’s table of summaries by level of measurement and shape, and justify each choice in one sentence. Write your answer first, then open the model answer.

age is ratio but skewed to the right (a few much older students), so the median and IQR, or the mean and SD with a note on the skew. gender, faculty, programme, study_mode, has_children, lives_away, workshop, and considering_dropout are nominal, so counts and percentages; there is no average faculty. employment and financial_worry are ordinal, so percentages in each category and, for financial worry, the median; their steps are not equal, so a mean would suggest a precision the data does not have. workshop_sessions is a count with only seven values, so the median and IQR, or the percentage attending each number of sessions. student_id and supervisor_id are identifiers and are not summarised, although the number of students per supervisor is worth reporting. Each choice follows from the level of measurement first, and then from the shape.

Go further

These exercises go beyond the book.

Exercise 7: The standard deviation by hand

Five students sleep 4, 6, 7, 8, and 10 hours. Build their standard deviation step by step: the deviations from the mean, their squares, the variance, and its square root. Then compare with sd().

NoteHint

The sum of squared deviations is divided by one less than the number of values, length(sleep_five) - 1. The standard deviation is the square root of the variance.

TipSolution
variance <- sum(deviations^2) / (length(sleep_five) - 1)
sqrt(variance)
sd(sleep_five)

The deviations are −3, −1, 0, 1, and 3; they always add up to zero, which is why they are squared before averaging. The squares add up to 20, the variance is 20 / 4 = 5, and the standard deviation is √5 = 2.24 hours, exactly what sd() gives. In words: a typical student’s sleep is about 2.2 hours away from the average of 7. Dividing by n − 1 rather than n corrects for the fact that deviations from the sample mean are slightly smaller than deviations from the unknown population mean.

Exercise 8: Mean and median of age

Compare the mean and the median age of the students, and use a histogram to explain the difference.

NoteHint

The median is the middle value, and its function is named after it. In the histogram, look at which side has the longer tail.

TipSolution
mean(students$age, na.rm = TRUE)
median(students$age, na.rm = TRUE)
hist(students$age, breaks = 20, main = "", xlab = "Age (years)")

The mean age is 29.7 and the median 29. The histogram shows why: most students are in their mid-to-late twenties, but a tail of older students stretches to 52, and those few large values pull the mean up while the median, the middle student, stays put. The difference is small here, but it is the typical pattern of a right-skewed variable, and the direction of the gap between mean and median tells you the direction of the skew.

Exercise 9: Where are the gaps?

Count the missing values in every column of semesters, separately for each semester, and describe the pattern.

NoteHint

across(everything(), ...) applies the same calculation to every column; the function that answers “is this value missing?” with TRUE or FALSE goes inside sum(), which counts the TRUE values.

TipSolution
semesters |>
  group_by(semester) |>
  summarise(across(everything(), ~ sum(is.na(.x)))) |>
  as.data.frame()

The gaps are few and scattered: at most 8 missing values in any column in any semester, in the self-reported variables (sleep, study hours, exercise, caffeine). GPA (with a single exception), supervisor meetings, and wellbeing are complete. The pattern looks like occasional skipped questions, which is consistent with data missing completely at random. The larger gap in this data is invisible in this table: the 37 students with no rows at all in semesters 3 and 4. Counting NAs never shows missing rows.

Exercise 10: When one mean misleads

R’s built-in faithful data records 272 waiting times, in minutes, between eruptions of the Old Faithful geyser. Calculate the mean waiting time, draw a histogram, and then split the waits into short and long at a cut-off you read from the histogram.

NoteHint

The histogram has two peaks; the cut-off belongs in the valley between them.

TipSolution
faithful |>
  mutate(type = if_else(waiting < 67, "short", "long")) |>
  group_by(type) |>
  summarise(eruptions = n(), mean_wait = mean(waiting))

The mean wait is 70.9 minutes, but the histogram has two peaks, around 55 and 80 minutes, with a valley in between exactly where the mean falls: only 20 of the 272 waits are within 3 minutes of it. The mean describes almost no actual wait. Split at the valley (any cut-off between about 65 and 70 works), there are 99 short waits averaging 54.6 minutes and 173 long waits averaging 80.2. A bimodal distribution usually means two kinds of case mixed together, here two kinds of eruption, and it should be described as two groups, not by one centre.

Check your understanding

Answer each question in your own words first, then click to see a model answer.

1. Which summary suits an ordinal variable, such as the five answers to a question about money worries?

The percentage in each category, and the median as the centre. An ordinal variable has a meaningful order, so the middle answer makes sense, but its steps are not necessarily equal, so the mean, which treats “not at all” to “a little” as the same distance as “very” to “extremely”, suggests a precision the data does not have.

2. A report says that sleep has a mean of 6.5 hours and an SD of 1.1 hours. What does the SD mean in words?

That a typical student’s sleep is about 1.1 hours away from the average of 6.5 hours. If sleep is roughly normal, about two-thirds of students sleep between 5.4 and 7.6 hours, and about 95% between 4.3 and 8.7 hours.

3. Why is an outlier not simply deleted?

Because an unusual value is usually a real case, and deleting real cases makes the data look tidier and less varied than the world it describes. An outlier should be checked: an impossible value is an error and is set to missing, but a possible one stays. If it strongly affects a result, the honest approach is to report the analysis with and without it.

4. What is the difference between data missing at random and data missing not at random?

Missing at random means the gaps depend on something that was measured, such as study mode, so an analysis that includes that variable can allow for them. Missing not at random means the gaps depend on the missing value itself, for example when the students most worried about money are the least willing to answer the question about money. That is the hardest case, because the students who answered are no longer typical, and nothing in the data shows it.

5. Two variables have a correlation of 0. Does that mean they are unrelated?

No. The correlation coefficient measures only straight-line relationships. A curved relationship, such as wellbeing rising with sleep up to 8 hours and falling after, can give a correlation near zero although the two variables are strongly related. This is one more reason to plot the data before summarising it.

Do it yourself

These tasks have no starter code and no answers. Each code space below is empty and ready to run: write your own code, as you would for a thesis. The data (students, semesters, questionnaire, and first_sem) and the dplyr package are already loaded, and R’s built-in datasets are always available.

1. Produce a complete “Table 1” describing the sample: the number of students, and for each baseline variable the percentages, or the mean and SD, or the median and IQR, as its level and shape require, with the number of missing values. Then write the paragraph that describes the sample, with every number calculated by code.

2. Describe a variable the chapter did not describe, supervisor_meetings in the first semester: its centre, spread, and shape, and any unusual values. Decide which summaries to report, and write two sentences about it.

3. Describe R’s built-in airquality data for a report: the centre, spread, and shape of ozone and temperature, the missing data in each variable, and the relationship between the two.

4. R’s built-in precip data gives the average yearly rainfall, in inches, of 70 US cities. Describe its distribution, name the cities with unusual values, and decide whether the mean or the median better describes a typical city.

Work on your own computer

NoteDownload the Chapter 6 project

The project contains the data and all four parts of this page as an R script: the exercises with blanks, the questions to check your understanding (answers in solutions.R), and the open tasks, each with space to write your code.

  • Download chapter06.zip, unzip it, and double-click chapter06.Rproj.
  • Or type this one line in RStudio’s Console:
usethis::use_course("https://polla-fattah.github.io/data2thesis_r/playground/chapter06.zip")