Playground: Chapter 12

Classification Models

This page practises the ideas of Chapter 12 in four parts. Practise the chapter works through the book’s exercises, with starter code, hints, and solutions. Go further goes beyond the book, with new questions and some of R’s built-in datasets. Check your understanding asks short questions whose answers open with a click. Do it yourself gives open tasks with an empty code space and no answers: you choose how a classifier is used, as you will in your own thesis.

tidymodels is too large to run in a browser, so this page works with functions that come with R: decision trees with rpart, logistic regression with glm(), and confusion matrices counted by hand. The setup builds cross-validated predictions for every student, like the chapter’s cv_predictions: each student’s probability comes from a model that did not see them. Three of the book’s exercises need tidymodels packages (ranger, kknn, and themis); they are marked as tasks for the download project, which uses tidymodels exactly as the chapter does. Because base R drops students with missing values (545 of 600 remain), the numbers here differ a little from the chapter’s.

Practise the chapter

These are the exercises at the end of Chapter 12, with the same numbers. In the setup, ml_data holds the students and their predictors, and cv_pred holds each student’s cross-validated probability of considering dropout (prob), the truth (truth, 1 for “Yes”), and their faculty.

Exercise 1: A tree of depth 2

Draw the decision tree for the wellbeing data with a depth of 2. Name the predictors it uses, and explain the tree in plain words, as you would to a counsellor.

NoteHint

maxdepth is the number of questions a student can be asked on the way from the top of the tree to a leaf.

TipSolution
small_tree <- rpart(dropout ~ ., data = tree_data, method = "class",
                    control = rpart.control(maxdepth = 2))

The tree uses two predictors, support and burnout. In plain words: “Students who feel reasonably supported by their supervisor (a support score of 2.55 or more) rarely consider dropping out, about 9 in 100. Among students who feel poorly supported, those who are also burnt out (a burnout score of about 3.4 or more) are the group to watch: more than half of them consider dropping out. Poorly supported students who are not burnt out are in between, at about 18 in 100.” The tidymodels version in the download project, which fills in missing values instead of dropping students, finds the same two questions with slightly different cut-offs.

Exercise 2: Tuning the forest’s mtry

Tune the random forest’s mtry over c(2, 5, 10) with cross-validation, and compare the best value with the default.

This exercise needs tidymodels and the ranger package, so it is done in the download project.

forest_wf <- workflow() |>
  add_recipe(tree_recipe) |>
  add_model(rand_forest(trees = 500, mtry = tune()) |>
              set_engine("ranger") |> set_mode("classification"))
set.seed(2026)
forest_tuning <- tune_grid(forest_wf, dropout_folds,
                           grid = tibble(mtry = c(2, 5, 10)),
                           metrics = metric_set(roc_auc))
collect_metrics(forest_tuning)

The cross-validated AUCs are 0.783 for mtry = 2, 0.780 for 5, and 0.775 for 10, and the default (the square root of the 20 predictors, rounded down to 4) gives 0.781. The differences are far smaller than their standard errors (about 0.035), so mtry hardly matters here: the default is fine. Small values make the trees more different from each other, which helps the ensemble when a few strong predictors would otherwise dominate every tree.

Exercise 3: Very many neighbours

Extend the k-NN grid to neighbors = c(81, 161, 301), describe what happens to the AUC, and explain what happens to a k-NN model as \(k\) approaches the number of training students.

This exercise needs tidymodels and the kknn package, so it is done in the download project.

set.seed(2026)
knn_tuning <- tune_grid(knn_wf, dropout_folds,
                        grid = tibble(neighbors = c(41, 81, 161, 301)),
                        metrics = metric_set(roc_auc))
collect_metrics(knn_tuning)

The AUC keeps rising: 0.73 with 41 neighbours, 0.75 with 81, 0.77 with 161, and 0.79 with 301, close to logistic regression. With more neighbours, each prediction averages over more students and becomes smoother, which suits this data, where the pattern is simple and the noise large. But the rise cannot go on for ever. Each cross-validation fold trains on about 404 students; if every one of them voted equally, every new student would get the same prediction, the overall share at risk, and the AUC would fall to 0.5. The kknn engine avoids this at 301 only because it weights closer neighbours more heavily. A k-NN model with \(k\) near the number of students has stopped being “nearest neighbours” at all.

Exercise 4: The F1 score by hand

Using the tiny example’s confusion matrix (a questionnaire flags 12 of 100 students; 8 of the 12 are among the 10 truly at risk), calculate the F1 score yourself from precision and recall: \(F_1 = 2 \times \text{precision} \times \text{recall} / (\text{precision} + \text{recall})\).

NoteHint

Precision divides the true positives by all the students flagged; recall divides them by all the students truly at risk. The denominator of F1 is the sum of the two.

TipSolution
precision_tiny <- 8 / 12
recall_tiny    <- 8 / 10
f1 <- 2 * precision_tiny * recall_tiny / (precision_tiny + recall_tiny)

Precision is 0.67 and recall 0.80, so F1 is 0.73, as the chapter says. F1 is the harmonic mean of the two, which always lies closer to the smaller one: a model cannot earn a high F1 by excelling at one and neglecting the other. With a precision of 0.1 and a recall of 1.0, for example, the ordinary average is 0.55, but F1 is only 0.18.

Exercise 5: Contacting only 10% of students

Suppose the counselling service can only contact 10% of new students. Build the threshold table from the cross-validated predictions, recommend a threshold, and report the recall and precision it would give.

NoteHint

Precision divides the correctly flagged students by all the flagged students, which is the number of TRUE values in flag.

TipSolution
precision = sum(flag & cv_pred$truth == 1) / sum(flag)

At 0.5, the model flags 7% of students; at 0.4, 11%, just over the limit. From the table, the recommendation is therefore 0.5: it finds 31% of the students at risk, and 60% of the contacted students are truly at risk. A finer search between the two shows that 0.45 still stays within the limit (it flags 9% of students) and finds a few more (37%, with a precision of 0.60), so a threshold of about 0.45 uses the available time best. The limit on resources costs a lot: at the threshold of 0.2 that the chapter chose, the model would find 64% of students at risk, but it would contact 26% of all students. The table makes that trade-off visible, so that the service can decide whether more counselling time would be worth it. (The chapter’s threshold_table, from tidymodels, gives nearly the same numbers.)

Exercise 6: Downsampling instead of upsampling

Replace step_upsample() with step_downsample(), and compare the sensitivity and AUC with upsampling.

This exercise needs tidymodels and the themis package, so it is done in the download project.

downsample_recipe <- dropout_recipe |>
  step_downsample(considering_dropout, under_ratio = 1)
downsample_wf <- workflow() |> add_recipe(downsample_recipe) |> add_model(logistic_reg())
set.seed(2026)
fit_resamples(downsample_wf, resamples = dropout_folds,
              metrics = metric_set(roc_auc, sensitivity, precision)) |>
  collect_metrics()

At the 0.5 threshold, downsampling gives a sensitivity of 0.75 and a precision of 0.29; upsampling gives 0.67 and 0.31; the ordinary model, 0.31 and 0.56. The AUC hardly changes: 0.79 with either kind of resampling, against 0.80 without. Both kinds of resampling shift the probabilities upwards, which works like a lower threshold, and neither makes the model better at telling students apart. Downsampling throws away most of the “No” students, which is why its AUC can be slightly lower; with a small dataset, that loss of information is its main drawback.

Exercise 7: A cheaper missed student

Repeat the cost analysis with a cost of 3 for a missed student instead of 10. Report the new cheapest threshold, and compare it with the decision-theory rule.

NoteHint

The blank is the new cost of a missed student. The last number is the decision-theory threshold: the cost of a false alarm divided by the sum of the two costs.

TipSolution
cost_miss <- 3

The decision-theory rule now gives 1 / (1 + 3) = 0.25, much higher than the 0.09 for a cost of 10: when missing a student matters less, fewer students are worth contacting. The cost curve’s cheapest threshold here is 0.44, but the plot shows why that number should not be taken literally: the curve is almost flat between about 0.20 and 0.46, where every threshold costs between 168 and 177, and 0.24 costs only 2% more than the minimum (171 against 168). With 78 students at risk, the exact minimum of a flat curve is a matter of chance. (The tidymodels version in the download project finds its minimum at 0.28, close to the rule.) Both methods agree on what matters: the threshold rises sharply as the cost of a miss falls, and a report should give the costs and a range of reasonable thresholds rather than one precise number.

Go further

These exercises go beyond the book.

Exercise 8: Confusion matrices by hand

Ten students have the predicted probabilities and true outcomes below. Count the four cells of the confusion matrix, and the recall and precision, at thresholds of 0.5, 0.35, and 0.12.

NoteHint

A student is flagged when their probability is at or above the threshold.

TipSolution
flag <- prob >= t

At 0.5, the model flags 4 students, 3 of them at risk: recall 3/5 = 0.60, precision 3/4 = 0.75. At 0.35, it flags 6: recall 0.80, precision 0.67. At 0.12, it flags 9 and finds all 5 students at risk (recall 1.0), but only 5 of the 9 are at risk (precision 0.56). Each step down finds another student at risk and adds false alarms. Counting by hand once makes every later confusion matrix easy to read.

Exercise 9: When a false alarm has a cost

Suppose a flagged student receives a formal letter, which some students find upsetting, so a false alarm costs 2 instead of 1, while a missed student still costs 10. Calculate the cost curve, find the cheapest threshold, and compare it with the decision-theory rule.

NoteHint

Each kind of error is multiplied by its own cost.

TipSolution
total_cost <- cost_miss * misses + cost_false_alarm * false_alarms

The cheapest threshold is 0.14, and the decision-theory rule gives 2 / (2 + 10) = 0.17: close. Doubling the cost of a false alarm roughly doubled the threshold, compared with the chapter’s costs, and the model would now flag about a third of students instead of about half. Changing the consequence of being flagged, from a friendly conversation to a formal letter, changes the right threshold, which is why the threshold must be chosen with the people who will act on it.

Exercise 10: Does it work for every faculty?

At the chapter’s threshold of 0.2, calculate the recall and precision of the model separately for each faculty.

NoteHint

Recall divides the students found by all the students truly at risk in that faculty.

TipSolution
recall = sum(flag & truth == 1) / sum(truth == 1)

Recall ranges from 0.50 in Social Sciences (8 of 16 students at risk found) to 0.83 in Humanities (10 of 12), and precision from 0.29 to 0.45. Taken at face value, students at risk in Social Sciences would be missed half the time. But each faculty has only 9 to 28 students at risk, so these shares rest on very few students and could differ this much by chance. The right conclusion is not that the model is unfair to one faculty, but that a check like this needs far more data before a model is used, and that it must be done: an average recall can hide a group that is served much worse.

Exercise 11: Classifying iris species

R’s built-in iris data measures the flowers of 150 irises of three species. Split it 100/50, classify the species with a decision tree and with k-nearest neighbours written by hand (\(k = 5\), on scaled measurements), and compare the confusion matrices.

NoteHint

order(distances) lists the training flowers from nearest to farthest; the first k of them vote.

TipSolution
nearest <- train_y[order(distances)[1:k]]

Both classifiers recognise every setosa. k-NN classifies 46 of the 50 test flowers correctly (92%), confusing 4 versicolor flowers with virginica; the tree gets 48 right (96%), with one mistake each way. The tree is also easy to explain: petals shorter than 2.45 cm mean setosa, and among the rest, petals narrower than 1.65 cm mean versicolor. With only 50 test flowers, a difference of two flowers is not evidence that one method is better; what the confusion matrices show clearly is which species are hard to tell apart, which accuracy alone would hide.

Check your understanding

Answer each question in your own words first, then click to see a model answer.

1. What is the difference between precision and recall?

Recall (sensitivity) is the share of the truly positive cases that the model flags: how many of the students at risk are found. Precision is the share of the flagged cases that are truly positive: how often a contact is justified. Lowering the threshold usually raises recall and lowers precision, so the two must be read together.

2. Why is choosing a classification threshold a value judgement and not only a statistical decision?

Because the threshold decides how many cases of each kind of error the decisions will produce, and what those errors cost is a judgement about consequences for real people: a missed student who leaves, an unnecessary conversation, a stigmatising letter. Statistics can show the trade-off, but weighing the costs belongs to the people who will act on the decisions, and the choice should be stated and justified.

3. What does a random forest average, and why does averaging help?

It averages the votes of hundreds of decision trees, each grown on a different bootstrap sample of the data and allowed a random subset of predictors at each split. Each tree overfits in its own way, but because the trees differ, their errors partly cancel out when they vote, and the average is far more stable and accurate than any single tree.

4. Why does k-nearest neighbours need the predictors to be normalised?

Because it decides which students are “nearest” by distance, and distance adds up differences in every predictor. Without normalisation, a predictor measured in large units, such as caffeine in milligrams, would dominate the distance, and one in small units, such as a 1-to-5 stress score, would hardly count. Normalising puts every predictor on the same scale, so each contributes fairly.

5. What does upsampling change, and what does it not change?

It changes the balance of the training data by repeating rare-class cases, which pushes the predicted probabilities upwards and raises the sensitivity at a given threshold. It does not usually make the model better at telling the classes apart, so the AUC hardly moves, and it makes the predicted probabilities no longer mean what they say. For a model that produces probabilities, choosing the threshold directly is usually simpler.

Do it yourself

These tasks have no starter code and no answers. Each code space below is empty and ready to run: write your own code, as you would for a thesis. The data (ml_data and cv_pred) and the dplyr and rpart packages are already loaded, and R’s built-in datasets are always available.

1. Invent a scenario for the dropout model: for example, a screening questionnaire followed by a costly one-hour interview, or a free online resource sent by email. Decide the costs of a miss and of a false alarm and justify them, choose a threshold from the cross-validated predictions, and report the recall, precision, and share of students flagged.

2. In the download project, compare three classifiers (for example, logistic regression, a random forest, and k-NN) on a new outcome, such as lives_away, with a workflow set, and argue for one of them, using both the AUC and how easily each can be explained.

3. Write the paragraph that reports a classifier for a thesis: the model, how it was evaluated, its AUC, the threshold chosen and the reasons for it, and the recall and precision at that threshold, with the numbers from your work in Task 1. No code is needed; write it on paper or in a document.

Work on your own computer

NoteDownload the Chapter 12 project

The project contains the data and all four parts of this page as an R script. Its Practise part uses tidymodels exactly as the chapter does, including the three exercises that cannot run in a browser; the other parts use base R, as on this page. Answers to the questions are in solutions.R, and the open tasks have space to write your code.

  • Download chapter12.zip, unzip it, and double-click chapter12.Rproj.
  • Or type this one line in RStudio’s Console:
usethis::use_course("https://polla-fattah.github.io/data2thesis_r/playground/chapter12.zip")