Lecture slides

Putting It All Together

Putting It All Together

From the raw export to the last table, as one project

Chapter 19

Polla Fattah

By the end of today you can

  • organise a project that runs from raw data to finished results;
  • combine cleaning, description, modelling, and visualisation in one pipeline;
  • trace the chain of reasoning from question to conclusion and limitation;
  • report results as a thesis expects, following published standards;
  • choose a method for a research question;
  • recognise common challenges, and know where to learn more.

A thesis is an argument

The results chapter is the evidence.

Each claim rests on a chain: question, hypothesis, design, variables, analysis, result, conclusion, limitation.

It convinces when every link can be seen, and every number traced back to the raw data.

Rebuild it as one project

Analyses written over two years, as each method was learned, are scattered across scripts.

Before writing the results chapter: rebuild everything, from the raw export to the last table, as one project that runs from start to finish.

Then every number is right, and any examiner’s question can be answered by pointing to the code.

The shape of a research project

flowchart TB
  subgraph R1[" "]
    direction LR
    A["Question and design<br>Ch 5"] --> B["Import<br>Ch 1–2"] --> C["Clean<br>Ch 3"] --> D["Explore<br>Ch 4, 6"]
  end
  subgraph R2[" "]
    direction LR
    E["Test and model<br>Ch 7–10"] --> G["Report and share<br>Ch 17–18"]
    F["Predict and discover<br>Ch 11–16"] --> G
  end
  R1 --> R2

Not a straight line, but code makes going back easy.

Organising the project

wellbeing-thesis/
├── wellbeing-thesis.Rproj      the RStudio Project
├── README.txt                  what it is and how to run it
├── data-raw/wellbeing_raw.xlsx the export, never edited by hand
├── data/                       clean data, created by scripts
├── R/01-clean-data.R           raw export -> clean tables
├── R/02-analysis.R             models and figures
├── output/                     figures and tables, created by scripts
├── results.qmd                 the results chapter
├── references.bib
└── renv.lock                   package versions

Four principles

Principle Means
raw data is read-only data/ and output/ can be deleted and recreated
scripts are numbered in running order, one job each
paths are relative here::here() works on any computer
a README a few lines on what the project is and how to run it

The downloadable Chapter 19 project has exactly this structure.

From raw export to clean data

responses <- raw |>
  filter(!str_detect(str_to_upper(`Q1_Student ID`), "^TEST")) |>
  select(-`Response ID`) |>
  distinct() |>
  rename(student_id = `Q1_Student ID`, workshop = `Workshop group`,
         considering_dropout = `Y1_Considered leaving?`) |>
  rename_with(~ item_names, .cols = Q11_1:Q11_22) |>
  mutate(across(all_of(item_names), ~ as.integer(na_if(.x, "99"))))

The cleaning of Chapter 3 in one script: it runs in seconds.

The semester records

semesters_clean <- responses |>
  select(student_id, matches("_S[1-4]$")) |>
  pivot_longer(-student_id, names_to = c(".value", "semester"),
               names_sep = "_S") |>
  mutate(sleep_hours = parse_number(str_replace(sleep_hours, ",", ".")),
         sleep_hours = if_else(sleep_hours > 24, NA, sleep_hours),
         gpa         = if_else(gpa > 4, NA, gpa)) |>
  filter(!if_all(gpa:wellbeing, is.na))

The most important line: the check

same(questionnaire_clean |> arrange(student_id),
     questionnaire |> arrange(student_id))
#> [1] TRUE
same(semesters_clean |> arrange(student_id, semester),
     semesters |> arrange(student_id, semester))
#> [1] TRUE

In a real project there is nothing to compare with: check counts, ranges, missing values, and a few rows (Chapters 3 and 6).

One analysis table

study <- students |>
  left_join(scores, join_by(student_id)) |>
  left_join(semesters |> filter(semester == 1), join_by(student_id))

One row per student: background, questionnaire scores, and first-semester records.

Helper functions instead of copy and paste

mean_sd <- function(x) sprintf("%.1f (%.1f)", mean(x, na.rm = TRUE),
                                               sd(x, na.rm = TRUE))
percent <- function(x) sprintf("%.0f%%", 100 * mean(x, na.rm = TRUE))

describe <- function(d) {
  tibble(Characteristic = c("Students", "Age, mean (SD)", ...),
         Value = c(nrow(d), mean_sd(d$age), ...))
}

The same describe() makes every column. Write a small function whenever you would copy and paste.

Table 1: the sample

Characteristic Master’s PhD All
Students 426 174 600
Age, mean (SD) 27.9 (3.3) 34.2 (5.1) 29.7 (4.8)
Women 52% 53% 52%
Part-time 26% 39% 30%
Stress (1-5), mean (SD) 3.2 (0.7) 3.2 (0.7) 3.2 (0.7)
Support (1-5), mean (SD) 3.2 (0.8) 3.2 (0.7) 3.2 (0.8)
Wellbeing, semester 1 60.9 (12.1) 59.3 (11.9) 60.5 (12.0)
Considered dropping out 15% 14% 15%

Formatting p-values

format_p <- function(p) {
  if (p < 0.001) "p < .001"
  else paste("p =", sub("^0", "", sprintf("%.3f", p)))
}

Theses report p-values in a fixed style. A function applies it every time, without mistakes.

RQ3: the workshop and its persistence

workshop_model <- lmer(wellbeing ~ factor(semester) * workshop +
                         (time | student_id), data = panel)
gaps$wellbeing <- predict(workshop_model, newdata = gaps, re.form = NA)

Wellbeing was similar before the workshop (difference 0.4 points). Invited students’ wellbeing was 5.2 points higher in semester 2, narrowing to 4.1 in semester 3 and 2.6 in semester 4.

RQ5: explaining GPA

gpa_model <- lm(gpa ~ sleep_hours + study_hours + stress + support,
                data = study)
gpa_table <- tidy(gpa_model, conf.int = TRUE)

Each additional hour of sleep was associated with a GPA 0.11 points higher (95% CI 0.08 to 0.13, p < .001), holding study hours, stress, and support constant. The four predictors explained 25% of the variation (n = 587).

RQ9: considering dropout

dropout_model <- glm(
  I(considering_dropout == "Yes") ~ stress + support + financial_worry +
    employment + study_mode, data = study, family = binomial)
tidy(dropout_model, conf.int = TRUE, exponentiate = TRUE)

Each one-point increase in stress was associated with 3.6 times the odds of considering dropping out (95% CI 2.3 to 5.9); each point of support with 0.33 times the odds.

I() uses a condition as the outcome; exponentiate = TRUE gives odds ratios.

One figure for the thesis

patchwork: + side by side, / stacked, & applies a theme to all panels.

Saving the figure

(panel_a + panel_b) +
  plot_annotation(tag_levels = "A") &
  theme_minimal(base_size = 12)

ggsave(here::here("output", "figure-1.png"),
       width = 18, height = 8, units = "cm", dpi = 300)

At the size and resolution the thesis requires (Chapter 4).

A result is not yet a conclusion

Between them lie:

  • the design that produced the data;
  • how the variables were measured;
  • the limits of both.

The chain of reasoning makes each link visible.

The chain of reasoning: question to analysis

Step Workshop (RQ3) GPA (RQ5) Dropout (RQ9)
hypothesis invited higher in semester 2 more sleep, higher GPA stress raises the odds
design randomised, longitudinal observational observational
analysis mixed model multiple regression logistic regression

The chain of reasoning: result to limitation

Step Workshop (RQ3) GPA (RQ5) Dropout (RQ9)
result +5.2 in sem. 2, +2.6 by sem. 4 0.11 per hour of sleep OR 3.6 per point of stress
conclusion causal: the effect faded an association an association
limitation one university; invitation, not attendance self-reported sleep; unmeasured confounders considering, not leaving

Only the workshop was randomised, so only its conclusion is causal.

The results chapter

The tables, figures, and sentences go into results.qmd, with every number written as inline code (Chapter 17).

When the data changes, or an examiner asks for a different model, the code changes and the chapter renders again.

Reporting standards

Standard For
APA JARS quantitative research in psychology and related fields
CONSORT randomised trials in health research
STROBE observational studies

They list what readers need to judge a study, including effect sizes and confidence intervals. Check your draft against the relevant one.

Choosing a method: the goal

flowchart LR
  Q{What is the goal?} -->|Describe| D["Summaries and plots, Ch 4, 6"]
  Q -->|Explain or compare| O["Next slide"]
  Q -->|Predict new cases| P["Machine learning, Ch 11–15"]
  Q -->|Find structure| U["PCA, factors, clusters, Ch 9, 14"]
  Q -->|Forecast over time| T["Time series, Ch 16"]

Choosing a method: explaining or comparing

flowchart LR
  O["Outcome type?"] -->|Numeric| N["Repeated or nested data?"]
  O -->|Yes / No| L["Chi-square, logistic regression, Ch 7–8"]
  N -->|No| R["t-test, ANOVA, regression, Ch 7–8"]
  N -->|Yes| M["Mixed-effects models, Ch 10"]

The study’s questions and methods

Question Method Ch
RQ1 graduate life plots, descriptive statistics 4, 6
RQ2 sleep below 7 h one-sample t-test 7
RQ3 workshop t-test; mixed model 7, 10
RQ4 faculties differ ANOVA, post-hoc 7, 8
RQ5 explaining GPA multiple regression 8, 13
RQ6 questionnaire factor analysis, alpha 9

The study’s questions and methods (continued)

Question Method Ch
RQ7 student profiles clustering, mixture models 9, 14
RQ8 change over time mixed models 10
RQ9 considering dropout logistic regression; classification 8, 11, 12, 15
RQ10 predicting final GPA regularised regression, boosting 13
RQ11 challenges in words text coding with a language model 18
RQ12 counselling visits time series forecasting 16

Challenges every researcher meets: data

Challenge Response
messy data clean in code, check each step, never touch the raw file
missing data ask why before acting; leavers are rarely random
small samples report effect sizes with intervals; keep models simple; plan power

Challenges every researcher meets: analysis

Challenge Response
many tests decide main analyses in advance; correct; label exploration
assumptions check with plots; robust, non-parametric, or mixed models
correlation is not causation think through confounders; randomise where possible
prediction and explanation different aims, judged differently
communication clear tables and figures; export with write_csv() or write_sav()

Where to go next

Topic For Packages
structural equation modelling CFA, path models lavaan
Bayesian statistics prior knowledge, full uncertainty brms, rstanarm
survival analysis time until an event survival
meta-analysis combining studies metafor
text analysis words, sentiment, topics tidytext, quanteda
spatial data maps sf
publication tables formatted results gtsummary, modelsummary

You do not have to learn alone

  • free books: R for Data Science, Regression and Other Stories, ISLR, Tidy Modeling with R, Forecasting: Principles and Practice;
  • Posit Community and Stack Overflow for questions;
  • R-Ladies and local R user groups;
  • TidyTuesday: a new dataset every week to practise on.

A clear question with a small reproducible example is a skill, and the same one that makes an AI assistant useful.

In your field: nutrition science

tooth <- ToothGrowth |> mutate(dose = factor(dose))
tooth |> summarise(mean = mean(len), sd = sd(len), n = n(),
                   .by = c(supp, dose))
summary(aov(len ~ supp * dose, data = tooth))

Prepare, describe, model, report: a complete mini-project in a few lines.

Tooth length depends on dose, form, and their combination: orange juice helps more at low doses, not at the highest.

Practical lab: the Chapter 19 project

Download the complete thesis project from the playground.

Run the cleaning and analysis scripts, then render results.qmd.

Practical exercises 1–3: the pipeline

  1. Change one cleaning rule, render again, and note which numbers change.
  2. Add a row to Table 1 for students with children.
  3. Write a reporting sentence with inline numbers for the effect of support on GPA.

Practical exercises 4–6: figures, methods, reasoning

  1. Add a third panel to the thesis figure: wellbeing by faculty.
  2. Choose a method for a question from your own field.
  3. Complete the chain of reasoning for RQ2, sleep below 7 hours.

Try this yourself

For your own thesis:

  • set up the project structure today, with a README;
  • move the raw data into data-raw/ and make it read-only;
  • write the chain of reasoning for your main research question;
  • decide which reporting standard applies.

Troubleshooting guide (Part 1)

Symptom Likely cause
the project runs only on your computer absolute paths; use here::here()
a number in the thesis disagrees with the code typed by hand
cleaned data looks wrong but no error no checks at the end of the cleaning script
the same code pasted three times a helper function is needed

Troubleshooting guide (Part 2)

Symptom Likely cause
a causal claim from observational data the design was ignored in the conclusion
an examiner asks what was missing missing data and attrition not reported
the most advanced method chosen by default choose by goal, outcome, and structure
no limitations section every conclusion has limits; state them

Completion checklist

Misconceptions to leave behind (Part 1)

Misconception Better mental model
a significant result answers the question it is one link; design and measures decide its meaning
the discussion should defend the results it should also state their limits

Misconceptions to leave behind (Part 2)

Misconception Better mental model
formatting numbers by hand is quicker quicker once, wrong the next time
choose the most advanced method choose the simplest that fits the question

The chapter in one sentence

Build the thesis as one project that runs from the raw data to every number in the text, and keep each result attached to the design and limits that give it meaning.

The whole course in one picture

flowchart TB
  subgraph R1[" "]
    direction LR
    A["A question worth answering"] --> B["Clean, checked data"] --> C["The simplest fitting method"]
  end
  subgraph R2[" "]
    direction LR
    D["Results with size and uncertainty"] --> E["Honest conclusions and limits"]
  end
  R1 --> R2

From data to thesis: every step written in code, every claim traceable.

Questions

What is the first thing you will change in your own project?

Which link in your chain of reasoning is weakest?