Playground: Chapter 17
Reproducible Research
This page practises the ideas of Chapter 17 in four parts. Practise the chapter works through the book’s exercises, with starter code, hints, and solutions. Go further goes beyond the book, with new questions and simulations. Check your understanding asks short questions whose answers open with a click. Do it yourself gives open tasks with no answers: you make your own work reproducible, as you will in your thesis.
Quarto documents and Shiny apps run on your own computer, so the download project is the main practice for this chapter: a small Quarto project with a results chapter, a parameterised faculty report, and the wellbeing dashboard, with tasks and finished versions. This page practises the parts that run in the browser: numbers that write themselves into sentences, protecting participants before sharing data, the garden of forking paths, and recording versions. In the exercises, replace each ______ with your own code and press Run Code.
Practise the chapter
These are the exercises at the end of Chapter 17, with the same numbers.
Exercise 1: A sentence that writes itself
The chapter’s tiny Quarto document writes the average sleep of three friends into a sentence with inline code. Quarto runs on your own computer, but the idea can be tried here: build the sentence with paste0(), as inline code would, round the mean to one decimal place, then change one friend’s sleep and build it again.
The second argument of round() is the number of decimal places. In the Quarto document, the same expression goes between backticks after r: `r round(mean(sleep), 1)`.
paste0("My three friends slept ", round(mean(sleep), 1), " hours on average last night.")The first sentence says 6.3 hours, and after the change, 6.7: the sentence followed the data without anyone retyping a number. Without rounding, it would say 6.333333, which no reader needs. In the download project, the file results.qmd and its first task do the same in a real Quarto document; render it, change the data, and render again.
Exercise 2: A Quarto report of an earlier analysis
Turn one of your analyses from an earlier chapter into a Quarto document with a numbered table, a numbered figure, and a sentence with inline numbers. Render it to HTML and to Word.
This exercise needs Quarto, so it is done in the download project, whose results.qmd is a worked example to start from.
A table chunk gets a label starting with tbl- and a caption; a figure chunk a label starting with fig-; both are referred to in the text with @:
---
title: "Wellbeing by study mode"
format:
html: default
docx: default
---
```{r}
#| label: tbl-mode
#| tbl-cap: "First-semester wellbeing by study mode."
knitr::kable(mode_summary, digits = 1)
```
```{r}
#| label: fig-mode
#| fig-cap: "Wellbeing by study mode."
ggplot(first_sem, aes(study_mode, wellbeing)) + geom_boxplot()
```
As @tbl-mode and @fig-mode show, part-time students' average wellbeing was
`r round(gap, 1)` points lower.Rendering produces an HTML file and a Word file from the same source. The finished solutions/results.qmd in the project shows a complete version.
Exercise 3: References and a citation style
Add two references to a references.bib file, cite them in the document, and switch the citation style with a csl: file.
This exercise needs Quarto, so it is done in the download project, which includes a references.bib.
The YAML header names the bibliography and, optionally, a style; citations use the keys in square brackets:
bibliography: references.bib
csl: apa.cslEffect sizes were interpreted with Cohen's guidelines [@cohen1988], and mixed
models were fitted with lme4 [@bates2015].Quarto adds a reference list at the end. Citation styles are .csl files, freely available from the Zotero Style Repository for thousands of journals; downloading another one, such as vancouver.csl, and changing the csl: line changes every citation and the reference list at once, without touching the text.
Exercise 4: A slider in the dashboard
Add a slider to the Shiny dashboard that chooses the semesters shown.
This exercise needs Shiny, so it is done in the download project, whose app.R is the chapter’s dashboard.
The slider goes in the user interface, and the server filters the data by its value:
# in the user interface, next to the faculty menu
sliderInput("semesters", "Semesters", min = 1, max = 4, value = c(1, 4), step = 1)
# in the server, in the reactive data
chosen <- reactive({
panel |>
filter(faculty == input$faculty,
semester >= input$semesters[1],
semester <= input$semesters[2])
})A range slider returns two numbers, the lower and upper ends. Because the plot and the table both use chosen(), they update together whenever the slider moves. The finished solutions/app.R in the project has the complete app.
Exercise 5: The smallest cells
Look at the smallest cells in the table of faculty, study mode, gender, and children. Decide which variable you would remove or group before sharing the data, and explain why.
The loop has already calculated the smallest cell for each version of the table; print it.
print(c(removed = drop, smallest_cell = smallest))Five of the 40 combinations describe fewer than 5 students; the smallest, such as part-time male students with children in Education, describe only 3. All the small cells are part-time students with children, a rare combination. Removing any one variable leaves no cell below 5: without faculty the smallest cell has 21 students, without children 11, without study mode 10, and without gender 7. The choice should depend on the research question: if the shared data is meant for studying students with children, keep that variable and remove, or group, another. Gender is often the least needed for the wellbeing questions and gives the smallest protection (7), so grouping faculties, which keeps every variable, may be better (Exercise 9).
Exercise 6: A sixth analysis
Add a sixth analysis to forking_paths(), a t-test that excludes the five highest scores, and describe how the rate of false positives for any analysis changes. (The simulation repeats 2,000 studies, so it takes a few seconds.)
keep is TRUE for every student except the five with the highest scores; the scores and groups must be selected with the same condition.
no_top_five = t.test(score[keep] ~ group[keep])$p.valueThe single planned analysis still gives a false positive in about 5% of the studies (5.5%), but the chance that any of the six analyses does so rises from 10.9% with the chapter’s five analyses to 12.4%. Each extra defensible analysis adds another chance of a lucky result, although there is still nothing to find. The rise is smaller than it would be for an unrelated test, because excluding the five highest scores often agrees with excluding outliers; Exercise 7 explores this.
Go further
These exercises go beyond the book.
Exercise 7: How many forking paths?
Extend the simulation to ten defensible analyses, and calculate the rate of false positives when a researcher tries the first 1, 2, 5, and 10 of them. Compare with what ten independent tests would give.
A study counts as a false positive if any of the analyses tried has a p-value below the usual threshold.
mean(apply(p_values[1:k, , drop = FALSE], 2, function(p) any(p < 0.05)))Trying 1 analysis gives false positives in 5.5% of the studies, 2 in 7.7%, 5 in 10.9%, and 10 in 13.9%. Ten independent tests would give 40% (1 − 0.95¹⁰). The analyses of one dataset are strongly related, since they all compare the same two groups on nearly the same scores, so each extra path adds less and less. Even so, a researcher who tries ten reasonable analyses and reports the best has almost three times the advertised risk of a false positive, which a reader of the paper cannot see. Preregistration makes the planned path visible.
Exercise 8: A report for each programme
Change the chapter’s parameterised faculty report into a report for each programme (Master’s and PhD), and render it for both.
This exercise needs Quarto, so it is done in the download project, which includes faculty-report.qmd.
In the YAML header, the parameter changes from faculty to programme, and the code filters on it:
params:
programme: "Master's"chosen <- first_sem |> filter(programme == params$programme)It is then rendered once for each value, from the Terminal or with a short R loop:
for (p in c("Master's", "PhD")) {
quarto::quarto_render("programme-report.qmd",
execute_params = list(programme = p),
output_file = paste0("report-", p, ".html"))
}One source file produces as many reports as there are programmes, all guaranteed to use the same analysis.
Exercise 9: Grouping instead of removing
Instead of removing a variable, group the five faculties into two broader groups, Sciences (Natural Sciences and Health Sciences) and Humanities and social sciences (the other three), and check that no cell below 5 remains.
Count the cells with fewer than 5 students, the usual threshold for small cells.
sum(grouped$n < 5)No cell below 5 remains: the smallest cell, part-time male students with children in the sciences, now has 8 students. Grouping keeps all four variables, so the shared data can still answer questions about gender, study mode, and children; it gives up detail on the one variable that least needs it. Other protections work too, such as sharing age in bands instead of years, or adding small amounts of noise, but every one trades some usefulness for privacy, and the right balance depends on what the data will be used for and on what participants agreed to.
Exercise 10: Recording the software
Record the R version and the versions of the packages this page uses, and write them into the methods sentence a thesis would include.
The function that returns the installed version of a package takes its name in quotation marks and is named after what it does.
dplyr_version <- as.character(packageVersion("dplyr"))The sentence reports the versions actually used, built by code so that it is always right. In the browser, the R version is webR’s, which may differ from the one on your computer; that is exactly why versions must be recorded where the analysis runs. sessionInfo() lists everything loaded, for an appendix, and renv goes further, recording every package version in renv.lock so that the project can be restored exactly. citation("dplyr") gives the reference to cite.
Check your understanding
Answer each question in your own words first, then click to see a model answer.
1. What is the difference between a reproducible result and a replicable one?
A result is reproducible if the same data and the same code give the same result when someone else runs them. It is replicable if a new study, with new data collected in the same way, finds the same thing. Reproducibility checks the analysis; replication checks the finding. A result can be perfectly reproducible and still fail to replicate.
2. Why should numbers in the text of a thesis not be typed by hand?
Because typed numbers do not change when the data or the analysis changes. After a correction to the data, a hand-typed mean, p-value, or sample size silently goes out of date, and errors of copying creep in. Inline code calculates every number from the data when the document is rendered, so the text always matches the analysis.
3. What does renv record, and why does it matter?
The exact version of every package a project uses, in the file renv.lock. Packages change between versions, sometimes in their defaults or results; with the lock file, the project can be restored with the same versions on another computer or years later, so the analysis runs, and gives the same answers, as when it was written.
4. Why is removing names and student numbers not enough to anonymise data?
Because combinations of ordinary variables can identify people. A part-time male student with children in one faculty may be one of only three people, and anyone who knows the department could recognise him. Small cells must be protected, by grouping categories, removing variables that are not needed, or sharing only summaries, and what may be shared is limited by what participants consented to.
5. What does preregistration protect against?
Against choosing the analysis after seeing the results, the garden of forking paths, in which trying several reasonable analyses and reporting the one that works raises the rate of false positives far above 5%. By recording the hypotheses and planned analyses in advance, preregistration separates confirmatory results from exploratory ones, so readers know which findings were predicted.
Do it yourself
These tasks have no starter code and no answers. The code space is empty and ready to run; the data (students and semesters) and the dplyr package are already loaded.
1. Choose a result from an earlier chapter, such as the difference in wellbeing between part-time and full-time students, and write the reporting sentence as code: every number in it (the means, the difference, the test statistic, the p-value, and the sample sizes) calculated and inserted by paste0() or sprintf(), rounded as a journal would print them.
2. In the download project, turn one analysis from Chapters 6 to 10 into a Quarto report with a numbered table, a figure, inline numbers, and one citation, and render it to Word. Then add a second input to the Shiny dashboard, such as study mode or programme, and explain in comments how the inputs, the reactive data, and the outputs are linked.
3. Write a preregistration for a study of your own, using the headings of the AsPredicted template: the main question and hypotheses, the dependent variable, the conditions or groups, the planned analyses, the rules for excluding data, the sample size and how it was chosen, and anything else to preregister. No code is needed; write it in a document.
Work on your own computer
A small Quarto project: results.qmd, a results chapter with inline numbers, a numbered table and figure, and citations; faculty-report.qmd, a report with a faculty parameter; and app.R, the wellbeing dashboard. exercises.txt lists the Quarto and Shiny tasks, and the solutions folder has finished versions. exercises.R and solutions.R hold the R parts of this page. You need Quarto (it comes with RStudio) and the packages listed in README.txt.
- Download chapter17.zip, unzip it, and double-click
chapter17.Rproj. - Or type this one line in RStudio’s Console:
usethis::use_course("https://polla-fattah.github.io/data2thesis_r/playground/chapter17.zip")