Playground: Chapter 13
Predictive Regression
This page practises the ideas of Chapter 13 in four parts. Practise the chapter works through the book’s exercises, with starter code, hints, and solutions. Go further goes beyond the book, with new questions and some of R’s built-in datasets. Check your understanding asks short questions whose answers open with a click. Do it yourself gives open tasks with an empty code space and no answers: you build and judge the predictions yourself, as you will in your own thesis.
tidymodels, glmnet, and xgboost are too large to run in a browser, so this page works with functions that come with R. Ridge regression is calculated from its formula, boosting is built by hand from small rpart trees, and overfitting is shown with lm(). Three of the book’s exercises need glmnet or xgboost; they are marked as tasks for the download project, which uses tidymodels exactly as the chapter does.
In the setup, gpa_data has one row for each of the 563 students with a final GPA, with 38 predictors: the questionnaire items, age, financial worry, and the year-one records, with missing values filled in by medians. It is split at random into train (75%) and test, and rmse() calculates the root mean squared error. The numbers therefore differ a little from the chapter’s.
Practise the chapter
These are the exercises at the end of Chapter 13, with the same numbers.
Exercise 1: MAE and RMSE by hand
Calculate the MAE and RMSE of the chapter’s tiny example by hand. Then change one prediction so that it is off by 1.0 GPA point, and explain which measure changes more, and why.
The MAE ignores the sign of each error, so it needs the function for the absolute value.
c(mae = mean(abs(errors)), rmse = sqrt(mean(errors^2)))The MAE is 0.22 and the RMSE 0.25, as in the chapter (yardstick gives the same). With the last student’s prediction moved to 3.5, one miss of 1.0 point, the MAE rises by half, to 0.34, but the RMSE almost doubles, to 0.48. Squaring makes a 1-point error count 100 times as much as a 0.1-point error, where the MAE counts it 10 times as much. Report the RMSE when large misses are especially costly, and the MAE when a typical error is what matters; reporting both shows whether a few big misses drive the result.
Exercise 2: Sixty students or 150
Repeat the chapter’s 60-student experiment with 150 students, and report the gap between ordinary regression and the lasso. The ordinary regression runs here; the lasso needs glmnet, so its numbers come from the download project (see the solution).
The loop runs the experiment once for each sample size in the vector.
for (n in c(60, 150)) {With 60 students, ordinary regression fits its own sample closely (an RMSE of 0.12) but predicts new students with an RMSE of 0.32, no better than the baseline of 0.33. With 150 students, the test RMSE improves to 0.23: more students give the 38 coefficients more information and less room to chase noise. In the download project, on the chapter’s split, the lasso gives 0.20 with 60 students and 0.20 with 150, against 0.53 and 0.26 for ordinary regression. The gap shrinks from 0.32 to 0.07 as the sample grows: regularisation matters most when the sample is small compared with the number of predictors, which is exactly the situation of many theses.
Exercise 3: Ridge and lasso coefficients
Fit a ridge regression (mixture 0) with the best ridge penalty, compare its coefficients with the lasso’s, and count how many are exactly zero.
This exercise needs glmnet, so it is done in the download project. Go further, below, calculates ridge regression by hand.
best_ridge <- show_best(glmnet_tuning |> filter_parameters(parameters = tibble(mixture = 0)),
metric = "rmse", n = 1)
ridge_fit <- workflow(gpa_recipe,
linear_reg(penalty = best_ridge$penalty, mixture = 0) |>
set_engine("glmnet")) |>
fit(data = gpa_train)
tidy(ridge_fit) |> filter(term != "(Intercept)") |> summarise(zero = sum(estimate == 0))The best ridge model has a penalty of about 0.055 and a cross-validated RMSE of 0.213, a little worse than the lasso’s 0.204. None of its 51 coefficients is zero, while the lasso sets 41 of the 51 to zero and keeps 10. Both give the two year-one GPAs the largest coefficients (ridge: 0.12 and 0.10). Ridge spreads the rest of the credit thinly over every predictor; the lasso gives it to a few and drops the others. Ridge keeps everything and shrinks; the lasso shrinks and selects.
Exercise 4: Scale scores instead of items
Replace the 22 questionnaire items with the four scale scores, as in Chapter 11, and describe how the cross-validated RMSE of the lasso changes.
This exercise needs glmnet, so it is done in the download project.
With the four scale scores, the lasso’s best cross-validated RMSE is 0.203; with the 22 items, it was 0.204: no difference worth mentioning. The items add no predictive information that the scale scores lack, which is what a good scale should achieve: the averaging keeps the signal and removes item noise. For prediction, the smaller model with four scores is simpler to explain and to use, and it is as accurate. Year-one GPA carries most of the prediction either way.
Exercise 5: More trees for XGBoost
Tune XGBoost with trees = c(100, 500, 1000) and a learning rate of 0.03, and describe what happens to the RMSE as trees are added.
This exercise needs xgboost, so it is done in the download project. Go further builds boosting by hand in the browser.
xgb_wf <- workflow(gpa_recipe,
boost_tree(trees = tune(), learn_rate = 0.03) |>
set_engine("xgboost") |> set_mode("regression"))
set.seed(2026)
xgb_tuning <- tune_grid(xgb_wf, resamples = gpa_folds,
grid = tibble(trees = c(100, 500, 1000)),
metrics = metric_set(rmse))
collect_metrics(xgb_tuning)The cross-validated RMSE is 0.224 with 100 trees and 0.223 with 500 and with 1,000: nearly all the improvement is made in the first 100 trees, and after that each extra tree adds almost nothing. None of the three beats the lasso (0.204). With a small learning rate, boosting improves slowly and steadily, and the number of trees needed depends on the rate; here the relationship is so close to linear that the flexible trees have little to add to what a penalised linear model already captures.
Exercise 6: Explaining the RMSE to advisers
Write two sentences for the academic advisers explaining what an RMSE of 0.2 means for an individual student’s predicted GPA. Write your answer first, then open the model answer.
“The model’s predictions of final GPA are typically about 0.2 grade points away from the GPA a student actually achieves, so a student predicted to finish on 3.0 will usually end somewhere between about 2.8 and 3.2, and occasionally further away. The prediction is a useful starting point for a conversation, not a verdict: two students with the same prediction can finish half a grade point apart.” (If the errors are roughly normal, about two-thirds of students finish within one RMSE of their prediction and about 95% within two, from 2.6 to 3.4 in the example.)
Go further
These exercises go beyond the book.
Exercise 7: Ridge regression by hand
Ridge regression has a formula: the coefficients are \((X^\top X + \lambda I)^{-1} X^\top y\), where \(I\) is the identity matrix. With \(\lambda = 0\), it is ordinary regression. Calculate the ridge coefficients of the two year-one GPAs (standardised) for a range of penalties, and watch them shrink.
The identity matrix has ones on its diagonal and zeros elsewhere; the base R function that builds it takes the number of rows and is named after the diagonal.
b <- solve(t(X) %*% X + lambda * diag(ncol(X))) %*% t(X) %*% yWith no penalty, the coefficients are 0.120 for semester 1 GPA and 0.154 for semester 2; with a penalty of 100 they shrink to 0.111 and 0.130, with 1,000 to 0.055 and 0.059, and with 10,000 almost to zero. Two things happen at once: both coefficients shrink, and they become more equal. The two GPAs correlate at 0.69, so ridge shares the credit between them instead of letting one take it all. Adding \(\lambda\) to the diagonal is also what keeps the calculation stable when predictors are highly correlated, where ordinary regression’s \((X^\top X)^{-1}\) becomes unreliable.
Exercise 8: One sample of 60, or ten?
The chapter’s 60-student result came from one random sample. Repeat the experiment with 10 different samples of 60, fitting ordinary regression with all 38 predictors and with only the two year-one GPAs, and compare how much the test errors vary.
apply(results, 2, f) applies f to each column; the spread of the ten test errors is their standard deviation.
round(apply(results, 2, sd), 3)With all 38 predictors, the test RMSE ranges from 0.27 to 0.39 across the ten samples (standard deviation 0.036), and averages 0.34, worse than the baseline; with only the two GPAs, it stays between 0.197 and 0.205 (standard deviation 0.003). The flexible model is not only worse on average but also far more variable: its result depends heavily on which 60 students happened to be sampled. This is the variance part of the bias-variance trade-off from Chapter 11, measured directly, and it is what the lasso’s penalty reduces.
Exercise 9: Boosting by hand
Boosting starts with the average, fits a one-split tree to the errors, adds a fraction of its predictions (the learning rate), and repeats. Run 100 rounds, tracking the RMSE on the training and test students.
Each new tree is also applied to the test students, with the same function as for the training students.
pred_test <- pred_test + learn_rate * predict(tree, test)Both errors fall fast at first: after 10 trees, the test RMSE is 0.21, from 0.30 after one. The test error is lowest, 0.205, after 16 trees, then creeps up slightly and levels off at 0.207, while the training error keeps falling, to 0.187 after 100 trees. The widening gap is overfitting beginning. The number of trees is therefore a hyperparameter, to be chosen by cross-validation, and a smaller learning rate makes the steps smaller, so more trees are needed but the final model is usually better. XGBoost does the same, far faster and with extra refinements.
Exercise 10: The shape of a tree
R’s built-in trees data records the girth, height, and timber volume of 31 black cherry trees. Compare two models for predicting volume with leave-one-out cross-validation: a straight-line model, and a model on the logarithms of all three variables.
The log-log model predicts the logarithm of the volume, so its predictions must be turned back into cubic feet with the function that undoes log().
log_log = loo_rmse(log(Volume) ~ log(Girth) + log(Height), back = exp)The straight-line model predicts a left-out tree’s volume with an RMSE of 4.3 cubic feet; the log-log model with 2.6, about 40% better. Its coefficients explain why: volume grows with girth to the power of about 2.0 and with height to the power of about 1.1, just as the volume of a cylinder grows with the square of its diameter and with its height. The best predictive model here came from thinking about the shape of the relationship, not from a more complicated method. Leave-one-out cross-validation, which leaves out each tree in turn, suits a sample this small.
Check your understanding
Answer each question in your own words first, then click to see a model answer.
1. Why does ordinary regression overfit when there are many predictors and few cases?
Because each predictor adds a coefficient that can be adjusted to fit the sample, and with many coefficients and few cases, the model has enough freedom to fit the noise of that particular sample as well as the real pattern. Correlated predictors make it worse: their coefficients become large and unstable, cancelling out on the training data but not on new data.
2. What is the difference between ridge regression and the lasso?
Both penalise large coefficients and shrink them towards zero. Ridge penalises the squared coefficients: it shrinks every coefficient and shares the credit among correlated predictors, but keeps all of them in the model. The lasso penalises the absolute coefficients: it also sets some to exactly zero, so it selects a smaller set of predictors.
3. What does the penalty \(\lambda\) do, and how is its value chosen?
It sets how much large coefficients cost. With \(\lambda = 0\), the model is ordinary regression; as \(\lambda\) grows, the coefficients shrink more, until, at a very large value, the model predicts the mean for everyone. It is a hyperparameter, so its value is chosen by cross-validation: the value that gives the lowest error on the left-out folds.
4. How does boosting differ from a random forest?
A random forest grows many deep trees independently, each on a different bootstrap sample, and averages them, which reduces variance. Boosting grows small trees one after another, each fitted to the errors of the model so far and added with a small weight, which gradually reduces bias. A forest’s trees could be grown at the same time; a boosted model’s trees must be grown in order.
5. Why are a model’s predictions pulled towards the average, so that the highest achievers are predicted too low and the lowest too high?
Because part of every student’s outcome is unpredictable. A student with an exceptionally high final GPA usually had some good luck that no predictor could foresee, so the model, which predicts only the predictable part, gives them a value closer to the average. The same happens at the bottom. Regularisation strengthens the effect, because shrinking the coefficients pulls every prediction further towards the mean. This regression to the mean is normal, and it is why a prediction should not be read as a ceiling or a floor for a student.
Do it yourself
These tasks have no starter code and no answers. Each code space below is empty and ready to run: write your own code, as you would for a thesis. The data (gpa_data, train, and test), the function rmse(), and the dplyr, tidyr, and rpart packages are already loaded, and R’s built-in datasets are always available.
1. Predict semester 4 wellbeing, rather than GPA, from year-one information. Build the data as the setup does, with wellbeing in semester 4 as the outcome, compare a model with a few chosen predictors, a model with all of them, and boosting by hand, and report the test RMSE of each alongside the baseline. (In the download project, compare the lasso and XGBoost as well.)
2. Choose a small model from Task 1 and explain, in a short paragraph, why its predictors are not a list of the causes of wellbeing, using what you know about the data.
3. Write two sentences for the university’s student services explaining what your wellbeing model’s RMSE means for one student, in the units of the wellbeing scale. No code is needed; write it on paper or in a document.
Work on your own computer
The project contains the data and all four parts of this page as an R script. Its Practise part uses tidymodels, glmnet, and xgboost exactly as the chapter does, including the three exercises that cannot run in a browser; the other parts use base R, as on this page. Answers to the questions are in solutions.R, and the open tasks have space to write your code.
- Download chapter13.zip, unzip it, and double-click
chapter13.Rproj. - Or type this one line in RStudio’s Console:
usethis::use_course("https://polla-fattah.github.io/data2thesis_r/playground/chapter13.zip")