Playground: Chapter 19

Putting It All Together

This page practises the ideas of Chapter 19 in four parts. Practise the chapter works through the book’s exercises, with starter code, hints, and solutions. Go further goes beyond the book, with new questions about reporting and choosing methods. Check your understanding asks short questions whose answers open with a click. Do it yourself gives open tasks with no answers, ending with a capstone: a research question of your own, carried from hypothesis to limitation.

The main practice for this chapter is the complete thesis project in the download project: from the raw survey export, through a cleaning script and an analysis script, to a results chapter written in Quarto. This page practises the pieces that make a results chapter, which run in the browser: helper functions for reporting, the sample table, sentences with numbers taken from models, and the chain of reasoning. In the exercises, replace each ______ with your own code and press Run Code.

In the setup, study has one row per student with the background variables, the stress and support scores, and the first-semester records, as in the chapter, and the chapter’s helpers mean_sd(), percent(), format_p(), and describe() are defined.

Practise the chapter

These are the exercises at the end of Chapter 19, with the same numbers.

Exercise 1: Run the whole project

Download the Chapter 19 project, run R/01-clean-data.R and R/02-analysis.R, and render results.qmd. Then change one cleaning rule (for example, treat ages above 70 as impossible), render again, and note which numbers change.

This exercise needs the project on your own computer, so it is done in the download project; its exercises.txt walks through the steps.

In R/01-clean-data.R, the rule for impossible ages becomes, for example:

age = if_else(age < 18 | age > 70, NA, age)

After running both scripts and rendering again, only the numbers that involve age change: the mean age in the sample table, and the number of students with a known age. The models do not use age, so their coefficients, intervals, and p-values, and every sentence built from them, stay exactly the same. That is the point of the exercise: in a project where every number comes from code, a changed decision flows through to every result it affects, and to nothing else, without anyone retyping a number.

Exercise 2: Students with children in Table 1

Add a row to the sample table for the share of students with children, and build the table by programme.

NoteHint

The share of a yes-or-no condition, formatted as a percentage, is exactly what one of the chapter’s two helper functions does.

TipSolution
Value = percent(d$has_children == "Yes")

20% of Master’s students and 40% of PhD students have children, 26% overall. Because describe() builds a column for any group of students, one new line in one function adds the row to every column of the table: the helper-function habit at work. The difference between the programmes is worth a sentence in the thesis, since children were related to wellbeing and dropout in earlier chapters; it is also a reminder that differences between programmes may partly reflect differences in the students’ lives.

Exercise 3: A sentence for the effect of support on GPA

Write a reporting sentence, with the numbers inserted by code, for the effect of support on GPA in the chapter’s regression model.

NoteHint

The chapter’s helper that writes a p-value in the usual style, including “p < .001” for very small values, is defined in the setup.

TipSolution
format_p(p)

The sentence reads: “Each one-point increase in supervisor support was associated with a GPA 0.11 points higher (95% CI 0.08 to 0.14, p < .001), holding sleep, study hours, and stress constant.” Every number comes from the model, rounded as a journal would print it; the words “associated with” and “holding constant” say exactly what an observational regression can support. In the download project’s results.qmd, the same numbers go into the text as inline code.

Exercise 4: A third panel for the thesis figure

Add a third panel to the thesis figure: wellbeing by faculty as a box plot.

This exercise needs the patchwork package and the project’s figure code, so it is done in the download project.

In R/02-analysis.R, a third panel is made like the other two, and added with +:

panel_c <- study |>
  ggplot(aes(x = faculty, y = wellbeing)) +
  geom_boxplot() +
  labs(x = NULL, y = "Wellbeing in semester 1 (0-100)")

figure_1 <- (panel_a + panel_b + panel_c) +
  plot_annotation(tag_levels = "A") &
  theme_minimal(base_size = 12)

patchwork labels the new panel C automatically. Three panels side by side are narrow, so the figure may need to be wider when saved, or arranged as (panel_a + panel_b) / panel_c, with the box plot below the other two, and the caption must now describe all three panels.

Exercise 5: A method for your own question

Choose a research question from your own field. Using the chapter’s guide to choosing a method, decide which method you would use, and which chapter of this book you would reread first. Write your answer first, then open the model answer.

An example from public health: “Does a text-message reminder programme reduce missed clinic appointments, among patients followed for a year?” The goal is to explain or compare: does the programme make a difference? The outcome, missing an appointment, is yes or no, so the guide points to logistic regression (Chapters 7 and 8). But each patient has many appointments over the year, so the observations are repeated, and a logistic mixed-effects model with a random intercept for each patient is needed (Chapter 10). If clinics, not patients, were randomised to the programme, the clinic would be a second level. The chapter to reread first is Chapter 10, then Chapter 5 on design, to check whether the programme was assigned at random, which decides whether the conclusion can be causal.

Exercise 6: The chain of reasoning for RQ2

Complete the chain of reasoning for RQ2, whether students sleep less than 7 hours, using the results of Chapter 7. Run the test, then fill in the chain.

NoteHint

The hypothesis compares the average sleep with the recommended 7 hours. For the empty rows, the chapter’s chain for RQ3, RQ5, and RQ9 shows the kind of text each step needs.

TipSolution
sleep_test <- t.test(study$sleep_hours, mu = 7, alternative = "less")

Average first-semester sleep is 6.48 hours (95% CI 6.40 to 6.57, n = 594), and the one-sided test gives p < .001; 65% of students sleep less than 7 hours. The completed chain:

Step RQ2
Question Do students sleep less than 7 hours?
Hypothesis Average sleep is below 7 hours (\(H_0\): it is 7 hours)
Design Observational, first semester
Variables Self-reported hours of sleep per night
Analysis One-sample t-test against 7, one-sided, with a confidence interval (Chapter 7)
Result Mean 6.48 hours (95% CI 6.40 to 6.57), p < .001
Conclusion Graduate students sleep about half an hour less than the recommended 7 hours; the null hypothesis is rejected
Limitation Sleep is self-reported and may be over- or underestimated; one university; the first semester only

The one-sided test was justified because the hypothesis named the direction before the data was seen (Chapter 7). The conclusion is descriptive, not causal, which matches the question: it says how much students sleep, not why.

Go further

These exercises go beyond the book.

Exercise 7: A helper for confidence intervals

Write a helper function that formats an estimate with its 95% confidence interval, and use it in three sentences: the sleep coefficient for GPA, the difference in wellbeing between full-time and part-time students, and the odds ratio for stress in the dropout model.

NoteHint

The last number in the string is the upper end of the interval, the third argument of the function.

TipSolution
paste0(sprintf(f, estimate), " (95% CI ", sprintf(f, low), " to ", sprintf(f, high), ")")

The three sentences report a GPA higher by 0.11 (95% CI 0.08 to 0.13) points per hour of sleep, wellbeing higher by 3.8 (95% CI 1.7 to 5.9) points for full-time students, and odds of considering dropout multiplied by 3.6 (95% CI 2.3 to 5.8) per point of stress. One helper formats every interval in the same way, with the number of decimal places chosen for each quantity, which is how a thesis keeps its results consistent. The last sentence uses paste0(), which adds no spaces of its own, so that the full stop follows the bracket directly.

Exercise 8: The chain of reasoning for RQ4

Complete the chain of reasoning for RQ4, whether stress differs between faculties, with a one-way ANOVA, \(\eta^2\), and Tukey’s test.

NoteHint

Keep the pairs of faculties whose adjusted p-value is below the usual threshold.

TipSolution
tukey[tukey[, "p adj"] < 0.05, , drop = FALSE]
Step RQ4
Question Does stress differ between faculties?
Hypothesis Average stress differs between at least two faculties (\(H_0\): all five averages are equal)
Design Observational, first-semester questionnaire
Variables Stress score (1 to 5, six items, stress_4 reversed), faculty
Analysis One-way ANOVA with \(\eta^2\), Tukey’s test (Chapter 8)
Result F(4, 595) = 4.15, p = .003, \(\eta^2\) = 0.03; Health Sciences (3.37) higher than Humanities (3.03), the only significant pair
Conclusion Stress differs between faculties, but faculty explains only 3% of the differences between students
Limitation Students chose their faculty, so the difference may reflect who studies there rather than the faculty itself; one university

The result is significant, and small: knowing a student’s faculty says little about their stress. A conclusion that reports only “stress differed significantly between faculties” would overstate it.

Exercise 9: Table 1 by faculty

Use the same describe() function to build the sample table with one column for each faculty.

NoteHint

Inside the loop, f is the name of one faculty; keep the students of that faculty.

TipSolution
table_1[[f]] <- describe(filter(study, faculty == f))$Value

One function and a loop produce five columns: from 87 students in Natural Sciences to 154 in Health Sciences. The table shows at a glance what later analyses found, such as the higher stress (3.4) and dropout thoughts (19%) in Health Sciences and the lower stress in Humanities (3.0), and that the faculties are similar in age and the share of women. A table like this, in the first pages of the results, lets readers judge whether differences found later could reflect differences in who is in each group.

Exercise 10: Choosing methods

Use the chapter’s guide to choosing a method to decide which method suits each of these five research questions. Write your answers first, then open the model answer.

  1. How many hours a week do nurses in a hospital work, on average?
  2. Is the number of hours worked related to nurses’ burnout scores, allowing for age and ward?
  3. Burnout was measured every three months for two years: does it rise over time?
  4. Can the nurses who will leave the hospital within a year be identified from their records?
  5. Are there groups of nurses with similar patterns of shifts, overtime, and sick days?
  1. Describe: a mean with a confidence interval, and a histogram (Chapters 4, 6, and 7). 2. Explain, with a numeric outcome and one observation per nurse: multiple regression (Chapter 8); if nurses are nested in wards and the ward matters, a mixed-effects model with a random intercept for ward (Chapter 10). 3. Explain change over time with repeated measurements of the same nurses: a mixed-effects model with time as a predictor and random effects for nurses (Chapter 10). 4. Predict new cases with a yes-or-no outcome: a classification model, such as logistic regression, evaluated on held-out nurses with the AUC and a chosen threshold (Chapters 11 and 12). 5. Find groups: clustering, such as k-means or a mixture model, on scaled variables, reported as descriptive profiles (Chapters 9 and 14). The same data can answer all five, but each question needs its own method, which is why the question comes first.

Check your understanding

Answer each question in your own words first, then click to see a model answer.

1. Why should every number in a thesis be traced to code?

So that every result can be checked, repeated, and updated. A number calculated by code cannot be mistyped, stays correct when the data or the analysis changes, and shows exactly how it was obtained. An examiner who asks “where does this number come from?” can be shown the line of code.

2. What is the difference between a result and a conclusion?

A result is what the analysis produced: an estimate, an interval, a p-value. A conclusion is what the result means for the research question, which depends on the design that produced the data, how the variables were measured, and the limitations of both. The same result, such as a significant coefficient, supports a causal conclusion in a randomised experiment and only an association in an observational study.

3. Why is the workshop conclusion the only causal one in the chapter?

Because the workshop invitation was the only thing assigned at random. Randomisation makes the invited and not-invited groups alike on average in everything, measured or not, so a difference that follows can be attributed to the invitation. Sleep, stress, and support were not assigned; students with more of them may differ in other ways, so those results remain associations.

4. What is a limitations section for?

To tell readers how far the conclusions reach: what could still explain the results (such as unmeasured confounders), how the measures could be wrong (such as self-reported sleep), and to whom the results apply (one university, one cohort). It is not an apology but part of the argument: it shows that the researcher understands the evidence, and it tells the next researcher what to improve.

5. What are reporting standards such as JARS, CONSORT, and STROBE for?

They list the information readers need to judge a study: who the participants were and how they were recruited, how the variables were measured, how the analysis was done, and the results with effect sizes and confidence intervals. Checking a draft against the relevant standard catches missing information before an examiner or reviewer does, and many journals require it.

Do it yourself

These tasks have no starter code and no answers. The code space is empty and ready to run; the data (study, students, semesters, questionnaire) and the helpers are already loaded.

1. The capstone: choose a research question that the book did not answer, such as whether exercise is associated with wellbeing, allowing for sleep, and carry it through the whole chain. State the hypothesis and the design, choose and justify the method, check its assumptions, report the result in a sentence with every number from code, write the conclusion the design supports, and name two limitations. Start here, then turn it into a Quarto report in the download project.

2. Check your capstone report against the JARS items for participants, measures, and results: the number of participants and how they were recruited, how each variable was measured (with its reliability, for scales), and the results with effect sizes and confidence intervals. List anything missing, and add it. No code is needed.

3. Exchange reports with a fellow student, and review theirs using the chain of reasoning: does each conclusion follow from its result and its design, and is each limitation named? Write three constructive comments, each pointing to a specific place and suggesting an improvement. No code is needed.

Work on your own computer

NoteDownload the Chapter 19 project

The complete thesis project, with the folders data-raw, R, data, and output: R/01-clean-data.R turns the raw survey export into clean data, R/02-analysis.R runs the analyses and makes the figure, and results.qmd is the results chapter, with every number written as inline code. exercises.txt lists the project tasks, and exercises.R and solutions.R hold the R parts of this page. You need Quarto (it comes with RStudio) and the packages listed in README.txt.

  • Download chapter19.zip, unzip it, and double-click chapter19.Rproj.
  • Or type this one line in RStudio’s Console:
usethis::use_course("https://polla-fattah.github.io/data2thesis_r/playground/chapter19.zip")