Playground: Chapter 11
Introduction to Machine Learning in R
This page practises the ideas of Chapter 11 in four parts. Practise the chapter works through the book’s exercises, with starter code, hints, and solutions. Go further goes beyond the book, with new questions and some of R’s built-in datasets. Check your understanding asks short questions whose answers open with a click. Do it yourself gives open tasks with an empty code space and no answers: you design the test of a model yourself, as you will in your own thesis.
tidymodels is too large to run in a browser, so this page practises the chapter’s ideas with functions that come with R: splitting the data by hand, fitting with glm(), and measuring the AUC and cross-validated AUC with two short functions defined below. Two of the book’s exercises need tidymodels itself; they are marked as tasks for the download project, which uses tidymodels exactly as the chapter does. Because base R drops students with missing values (545 of 600 remain) where the chapter’s recipe fills them in, the numbers here differ a little from the chapter’s.
Practise the chapter
These are the exercises at the end of Chapter 11, with the same numbers. In the setup, ml_data holds one row per student with every predictor and dropout (1 for “Yes”), auc() calculates the AUC, and cv_auc() the 10-fold cross-validated AUC of a logistic regression.
Exercise 1: The model without the questionnaire
Remove the questionnaire scores (stress, burnout, support, satisfaction) from the data and fit the logistic regression again. Report how much the cross-validated AUC falls, and what that says about the questionnaire.
In a formula, a full stop after the tilde means “every other column of the data”.
cv_auc(dropout ~ ., data = no_questionnaire)With every predictor, the cross-validated AUC is 0.83; without the four questionnaire scores, it falls to 0.72. The background variables and semester records still predict something, but the questionnaire carries a large part of the signal: how students feel about their stress, burnout, support, and satisfaction tells more about who considers dropping out than their faculty, age, or grades. (With tidymodels, as in the chapter, the fall is from 0.80 to 0.69.)
Exercise 2: A different split
Split the data 80/20 instead of 75/25, with a different seed, and describe how much the test AUC of the logistic model changes, and why it might.
prop is the share of students used for training, written as a proportion. The last line repeats the 80/20 split with ten different seeds.
test_auc(prop = 0.8, seed = 7)The 75/25 split gives a test AUC of 0.78 and the 80/20 split with seed 7 gives 0.77, almost the same. But the ten seeds give AUCs from 0.77 to 0.88, which is the real lesson: an 80/20 test set holds only about 109 students, of whom about 15 are at risk, so the AUC depends heavily on which few at-risk students happen to land in it. A single test AUC from a small test set is a rough estimate. That is why cross-validation, which averages over many splits, is used to compare models, and why a thesis should report the uncertainty of its test result.
Exercise 3: Tuning k-nearest neighbours
Tune k-nearest neighbours over neighbors = c(5, 11, 21, 41) with cross-validation. Identify the best value, and compare its AUC with logistic regression.
This exercise needs tidymodels and the kknn package, so it is done in the download project, with the chapter’s workflow.
knn_wf <- workflow() |>
add_recipe(dropout_recipe) |>
add_model(nearest_neighbor(neighbors = tune()) |> set_mode("classification"))
set.seed(2026)
knn_tuning <- tune_grid(knn_wf, dropout_folds,
grid = tibble(neighbors = c(5, 11, 21, 41)),
metrics = metric_set(roc_auc))
collect_metrics(knn_tuning)The cross-validated AUC rises steadily with the number of neighbours: 0.62 with 5, 0.67 with 11, 0.71 with 21, and 0.73 with 41, so 41 is the best of the four. Even so, it stays below logistic regression (0.80). Averaging over more neighbours smooths out the noise of individual students, and the smoothest model wins: the relationship between the predictors and dropout is simple enough that a linear model captures it better than a method built for complicated patterns.
Exercise 4: Zero-variance predictors
Add step_zv(all_predictors()) to the recipe (it removes predictors with only one value). Look up what “zero variance” means, and explain when this step is useful. This exercise uses a recipe, so it too is in the download project; write your explanation first, then open the model answer.
A predictor with zero variance takes the same value for every case, so it carries no information and cannot help any model; some models even fail when they meet one, because they cannot scale it (its standard deviation is zero). In the wellbeing data, the step changes nothing: the prepared data has 25 columns with or without it. It becomes useful when some columns can be constant without anyone planning it. That happens when a rare category of a factor produces a dummy variable of all zeros in a small training set or a cross-validation fold, or after filtering the data (for example, keeping only full-time students makes study_mode constant). Adding step_zv() to a recipe is a cheap safeguard.
Exercise 5: Explaining accuracy to an administrator
In two or three sentences, explain to a university administrator why a model with 85% accuracy might still be useless for finding students at risk. Write your answer first, then open the model answer.
“Only about 15% of our students consider dropping out, so a ‘model’ that simply predicts that nobody will is already right 85% of the time, while finding not a single student at risk. Accuracy mostly rewards getting the many safe students right. What matters for us is how many of the students at risk the model finds, and how many false alarms it raises on the way, which is what measures such as the AUC and the sensitivity (Chapter 12) describe.”
Exercise 6: More training points
Repeat the polynomial simulation with 50 training points instead of 15. Describe how the error curve on new data changes, and explain why more data allows a more flexible model.
The blank is the number of training points. Run the code with 15 as well, to compare the two error curves.
train_points <- make_points(50)With 15 points, the error on new data is lowest at degree 3 (0.35) and explodes for the most flexible models, to over 500 at degree 12. With 50 points, the curve is flat and low from degree 3 to 12 (between 0.32 and 0.34), with its minimum at degree 5, and even degree 12 predicts new points well. With more data, a flexible curve has too many points to bend through every one of them, so it can no longer chase the noise; its extra flexibility is spent on the real pattern. The best flexibility is not a fixed property of a method: it depends on how much data there is.
Go further
These exercises go beyond the book.
Exercise 7: Data leakage
A researcher adds 200 columns of pure random noise to the data, picks the 5 that correlate most strongly with dropout, and tests a model built on them. Do it twice: once choosing the columns with all the data, and once with the training data only.
Choosing the columns is part of building the model, so it must use only the rows the model is trained on.
chosen_with_training <- order(-abs(cor(noise[train_rows, ], y[train_rows])))[1:5]Choosing with all the data gives a “test” AUC of 0.69 from columns of pure noise, which cannot predict anything; choosing with the training data only gives 0.50, exactly what noise deserves. Among 200 random columns, some correlate with dropout in the test students by chance, and choosing with all the data picked exactly those, so the test set was no longer new to the model. This is data leakage: any step that looks at the test data, whether selecting predictors, filling in missing values, or scaling, lets information about it into the model and makes the test too optimistic. Recipes and workflows prevent it by learning every step from the training data only.
Exercise 8: Cross-validation by hand
Build five-fold cross-validation yourself, with glm() only: assign every student to one of five folds, fit on four folds, calculate the AUC on the fifth, and repeat for each fold.
The model is fitted on every fold except fold k, and evaluated on fold k only. “Not equal” is !=, and “equal” is ==.
fold_auc <- sapply(1:5, function(k) {
fit <- glm(dropout ~ stress + burnout + support + satisfaction + financial_worry,
data = ml_data[folds != k, ], family = binomial)
auc(predict(fit, ml_data[folds == k, ], type = "response"),
ml_data$dropout[folds == k])
})The five folds give AUCs between 0.82 and 0.86, with an average of 0.85. Every student was used for evaluation exactly once, always by a model that had not seen them. The spread across folds, a standard deviation of about 0.02, shows how much a single test result could vary by luck. This simple model with five questionnaire-based predictors does at least as well as the model with every predictor in Exercise 1, a reminder that more predictors do not automatically mean better predictions.
Exercise 9: Noisier data, simpler model
Repeat the chapter’s polynomial simulation with 15 training points, but with twice as much noise (sd = 0.6). At which degree is the error on new data now lowest?
The noise is the standard deviation of rnorm(); the chapter used 0.3.
data.frame(x = x, y = true_curve(x) + rnorm(n, sd = 0.6))With twice the noise, the lowest error on new data is at degree 1, a straight line, where the chapter’s less noisy data favoured degree 3. When the noise is large compared with the pattern, 15 points cannot reveal the wave reliably, and any curve flexible enough to follow it mostly follows the noise instead. Noisy data and small samples both push the best model towards simplicity; clean data and large samples allow more flexibility.
Exercise 10: Who survived the Titanic?
R’s built-in Titanic data counts the 2,201 people on board by class, sex, age group, and survival. Expand it to one row per person, split it 75/25, fit a logistic regression of survival on class, sex, and age, and compare its test AUC and accuracy with a model that predicts that nobody survived.
Each row of the original table is a count, Freq; repeating each row Freq times gives one row per person. predict() returns probabilities with the same type as in Chapter 8.
prob <- predict(fit, titanic[-train_rows, ], type = "response")The model’s test AUC is 0.75, and its accuracy is 79%, against 68% for predicting that nobody survived. The odds ratios tell the story of “women and children first”: women had about 11 times the odds of surviving that men had, adults about 0.4 times the odds of children, and passengers in second and third class and the crew lower odds than those in first class. With only three categorical predictors, the model can separate just 14 groups of passengers, which limits how well it can rank individuals.
Check your understanding
Answer each question in your own words first, then click to see a model answer.
1. What is the difference between a model built for prediction and one built for explanation?
A model for explanation asks why: which variables are related to the outcome, how strongly, and whether the relationship could be causal, so its coefficients and their uncertainty matter. A model for prediction asks how well: how accurately it forecasts the outcome for new cases, whatever its coefficients mean. A good predictor can be useless for explanation (it may rely on variables that are only markers), and a good explanation can predict poorly.
2. Why is the test set used only once?
Because every time a decision is based on the test result, such as trying another model or another setting, the test set helps to build the model and stops being new data. After several rounds, the chosen model is the one that happened to suit those particular test cases, and its test result is too optimistic. Using the test set once, at the very end, keeps it an honest estimate of performance on new cases.
3. What does cross-validation estimate?
How well a way of building a model, with its data preparation, method, and settings, will perform on new data, estimated from the training data alone. Each case is predicted by a model that did not see it, and the results are averaged over the folds. It is used to compare models and tune settings without touching the test set.
4. What is overfitting, in one sentence?
Overfitting is when a model learns the noise of its training data along with the real pattern, so that it performs very well on the data it learned from and poorly on new data.
5. What is data leakage?
Letting information that would not be available at the time of prediction, or information from the test set, into the building of a model. Examples are choosing predictors or scaling with all the data before splitting, or using a variable measured after the outcome, such as the workshop in the chapter. Leakage makes a model look better than it can ever be in practice.
Do it yourself
These tasks have no starter code and no answers. Each code space below is empty and ready to run: write your own code, as you would for a thesis. The data (students, semesters, questionnaire, scores, and ml_data), the functions auc() and cv_auc(), and the dplyr package are already loaded, and R’s built-in datasets are always available.
1. Build a predictive model for another yes-or-no outcome, such as lives_away: prepare the data, split it, compare two sets of predictors with cross-validation on the training data, and report the chosen model’s AUC on the test data, alongside the accuracy of a model that always predicts the more common answer.
2. Using the numbers from your model in Task 1, explain in a short paragraph why accuracy would be misleading for your outcome, or why it would not.
3. Write the prediction section of a methods chapter for your model in Task 1: the data used, the outcome and predictors, how the data was split, how models were compared, and how the final model was tested. No code is needed; write it on paper or in a document.
Work on your own computer
The project contains the data and all four parts of this page as an R script. Its Practise part uses tidymodels exactly as the chapter does, including the two exercises that cannot run in a browser; the other parts use base R, as on this page. Answers to the questions are in solutions.R, and the open tasks have space to write your code.
- Download chapter11.zip, unzip it, and double-click
chapter11.Rproj. - Or type this one line in RStudio’s Console:
usethis::use_course("https://polla-fattah.github.io/data2thesis_r/playground/chapter11.zip")