Playground: Chapter 15
Neural Networks
This page practises the ideas of Chapter 15 in four parts. Practise the chapter works through the book’s exercises, with starter code, hints, and solutions. Go further goes beyond the book, with new questions and some of R’s built-in datasets. Check your understanding asks short questions whose answers open with a click. Do it yourself gives open tasks with an empty code space and no answers: you decide whether a network is worth it, as you will in your own thesis.
tidymodels is too large to run in a browser, but the nnet package, which is the engine behind the chapter’s mlp() models, runs here, so almost everything on this page is real network training. The download project uses tidymodels exactly as the chapter does. The setup below splits the data 75/25 and normalises the predictors with the training means and standard deviations; because it keeps only students with every value (548 of 600), the numbers differ a little from the chapter’s.
Practise the chapter
These are the exercises at the end of Chapter 15, with the same numbers. In the setup, train and test hold the students, with dropout (1 for “Yes”) and nine normalised predictors; auc() and sigmoid() are defined as in the earlier chapters.
Exercise 1: A neuron with new weights
Change the neuron’s weights so that support matters twice as much as in the chapter, recalculate the outputs for the two students (stress 4 and support 2; stress 2 and support 4), and explain what a negative weight means.
The chapter’s weight for support was −0.8; twice as much, in the same direction, is twice that number.
neuron <- function(stress, support) {
sigmoid(-4 + 1.2 * stress - 1.6 * support)
}The high-stress, low-support student’s output falls from 0.31 to 0.08, and the low-stress, high-support student’s from 0.008 to 0.0003. A negative weight means that the input lowers the output: every extra point of support pushes the total, and so the predicted risk, down. Doubling the weight doubles how strongly support pulls the total down, but because the sigmoid is curved, the effect on the probability depends on where the total starts.
Exercise 3: When the seed matters
Fit a network with 5 hidden neurons three times with different seeds, with a weight decay of 1, then do the same with a weight decay of 0.01 and with none. Explain when the seed matters, and why.
The seed sets the random starting weights; inside the function, use the value that sapply() passes in.
set.seed(seed)With a weight decay of 1, the three seeds give test AUCs of 0.900, 0.899, and 0.900: the seed hardly matters. With a decay of 0.01, they give 0.839, 0.773, and 0.813; with none, 0.832, 0.800, and 0.848. Training starts from random weights, and with little or no penalty, a network with 5 hidden neurons has many different sets of weights that fit the training data about equally well, so each seed ends somewhere different, and each overfits in its own way. A strong penalty pulls every run towards the same small-weight solution. (The tuned network in the download project, with 10 hidden neurons and a decay of 3.2, gives exactly the same test AUC, 0.853, for all three seeds, against 0.75 to 0.79 with a decay of 0.01.) If results change with the seed, the model is not regularised enough, and the seed must be reported.
Exercise 4: Epochs
Train a 20-neuron network without weight decay for 10, 100, and 1,000 epochs (maxit in nnet), and describe how the training and test AUCs change.
Inside the loop, the number of epochs takes each value in turn; pass it to maxit.
net <- nnet(dropout ~ ., data = train, size = 20, decay = 0,
maxit = epochs, entropy = TRUE, trace = FALSE)After 10 epochs, the training AUC is 0.88 and the test AUC 0.84: the network has learned the broad pattern and nothing more. After 100 epochs, the training AUC is a perfect 1.00 and the test AUC has fallen to 0.74, and 1,000 epochs change nothing further. The network, with 221 weights for 411 students, first learns the pattern and then memorises the noise. Stopping early is itself a form of regularisation, but a fragile one, since the best moment depends on the seed and the data; weight decay is the more reliable defence. (The download project’s tidymodels version shows the same: 0.72 after 10 epochs, 0.64 after 100.)
Exercise 5: Explaining the choice to the research group
In two or three sentences, explain to the research group why the thesis does not use a neural network for the dropout model. Write your answer first, then open the model answer.
“A carefully tuned neural network predicted dropout no better than logistic regression: both reached a cross-validated AUC of about 0.80, and the small difference on the test students is within chance. The network’s hundreds of weights cannot be interpreted, while logistic regression gives odds ratios that show how each factor relates to the risk of dropping out, so we kept logistic regression and report the network as a model we tried.” The answer names the comparison and its result, and the reason for the choice when the accuracy is equal.
Exercise 6: The learning rate
Repeat the chapter’s gradient-descent loop with a learning rate of 0.1 and of 10, and then of 100. Describe how the loss curve changes, and explain what goes wrong when the learning rate is too large.
The blank is the middle learning rate from the exercise. The chapter used 1; the answer from glm() is a stress weight of 2.43.
for (learning_rate in c(0.1, 10, 100)) {With 0.1, the steps are small: the loss falls smoothly but slowly, still 0.44 after 100 steps, and 2,000 steps are needed to reach the answer. With 10, the first step overshoots (the stress weight jumps to 2.92, past the answer of 2.43) and the loss settles within a few steps; this tiny, well-behaved problem tolerates a large rate. With 100, each step overshoots so far that the weight swings from 29 to −0.8 to 42, the loss bounces up and down, and after a few steps the probabilities reach exactly 0 and 1, so the log loss becomes infinite and R reports NaN. A learning rate that is too large makes gradient descent jump across the minimum instead of walking down to it; one that is too small wastes time. In real networks, with thousands of weights, the safe range is much narrower than here.
Go further
These exercises go beyond the book.
Exercise 7: Gradient descent with two predictors
Train a neuron with two inputs, stress and support, by gradient descent on the training students, and compare its weights with glm(). With several inputs, the gradient of every weight has the same form: the average prediction error multiplied by that weight’s input.
The prediction error is the predicted probability minus the actual outcome; the outcomes are in y. t(X) %*% error multiplies each input by the errors and adds them up, for all three weights at once.
b <- b - 1 * t(X) %*% (p - y) / nrow(X)After 2,000 steps, the weights are −2.246 for the bias, 0.915 for stress, and −0.649 for support, identical to glm() to four decimal places. Writing the step with matrices handles any number of inputs in the same line of code, which is how neural network software does it. With only 100 steps, the weights are already close (0.909 and −0.647), because this problem has one smooth minimum; networks with hidden layers have many, which is why they need more care.
Exercise 9: Weight decay with a small sample
Networks need more data than regression. Train networks with 5 hidden neurons on only 60 training students, with weight decays of 0, 0.1, 1, and 10 and three seeds each, and compare them with logistic regression on the same 60 students.
Inside the loop, the weight decay takes each value in turn; pass it on to nnet().
net <- nnet(dropout ~ ., data = small, size = 5, decay = decay,
maxit = 1000, entropy = TRUE, trace = FALSE)Without weight decay, the test AUC is 0.76 to 0.81, depending on the seed; with 0.1, 0.81 to 0.85; with 1, 0.889 for every seed; with 10, 0.85, as the penalty begins to flatten real patterns too. Logistic regression on the same 60 students (9 of whom considered dropping out) gives 0.85. With so little data, a network is worse than logistic regression unless it is heavily penalised, and even the best penalty gains little. The right amount of decay is found by cross-validation, and in small samples it is large.
Exercise 10: A network for timber volume
R’s built-in trees data records the girth, height, and timber volume of 31 black cherry trees. Fit a regression network (linout = TRUE gives a numeric output instead of a probability) and a linear regression, and compare them with leave-one-out cross-validation.
For a numeric outcome, the output neuron must be linear rather than a sigmoid, which would squeeze every prediction between 0 and 1.
net <- nnet(Volume ~ Girth + Height, data = fit_data, size = 3, decay = 0.1,
linout = TRUE, maxit = 2000, trace = FALSE)The straight line predicts a left-out tree’s volume with an RMSE of 4.3 cubic feet; the network with 3 hidden neurons does worse, 5.5. With 31 trees, the network’s 13 weights are too many to estimate reliably, even with some weight decay, while the linear model needs only 3. Chapter 13’s playground found a better model still, by taking logarithms, with an RMSE of 2.6. Flexible methods need data to be flexible with; with small samples, a simple model, or a thoughtfully transformed one, wins.
Check your understanding
Answer each question in your own words first, then click to see a model answer.
1. Why is a single neuron with a sigmoid activation the same as a logistic regression?
Because it computes the same thing: a weighted sum of the inputs plus a constant, passed through the sigmoid, which turns any number into a probability. The bias is the intercept, the weights are the coefficients, and training it by minimising the log loss finds the same values as glm(). Networks become more than logistic regression only when neurons are combined in layers.
2. What does the learning rate control?
How far the weights move at each step of gradient descent, in the direction that reduces the error. Too small, and training is slow; too large, and each step overshoots the minimum, so the error bounces around or grows without limit. It is a setting chosen before training, not something the network learns.
3. What does weight decay prevent?
Overfitting. It adds a penalty on large weights to the error that training minimises, like the ridge penalty in Chapter 13. Without it, a network with many weights can keep adjusting them until it memorises the noise of its training data; with it, the weights stay small, the network’s function stays smooth, and different random starts end at similar solutions.
4. Why must the seed be set, and reported, when training a neural network?
Because training starts from random weights, and a network can end at different solutions depending on where it starts, especially with little regularisation. Setting the seed makes the result reproducible; reporting it lets others repeat it. If results change noticeably with the seed, that is itself a warning that the model is not stable enough to trust.
5. When are neural networks worth using?
When there is a lot of data and the relationship is complex, with many interacting inputs whose combinations matter, as in images, sound, and text, where deep networks learn features that no one could specify by hand. For tabular data of a few hundred cases with a smooth relationship, like most thesis data, they rarely beat regression, and they cost interpretability, stability, and tuning effort.
Do it yourself
These tasks have no starter code and no answers. Each code space below is empty and ready to run: write your own code, as you would for a thesis. The data (train, test, and dropout_data), the functions auc() and sigmoid(), and the dplyr and nnet packages are already loaded, and R’s built-in datasets are always available.
1. Build a network for a different outcome, such as whether a student lives away from their family, or their semester 1 GPA (with linout = TRUE). Choose the number of hidden neurons and the weight decay by cross-validation, compare the network fairly with a regression model on the same folds, and decide which you would use. (In the download project, do this with tidymodels.)
2. Explain to a friend who has never studied statistics what “training” a neural network means, using your own gradient-descent loop as the example: what the loop starts from, what each step does, and how it knows when it is done. Write the explanation as comments above the code.
3. For each of these three research projects, decide whether a neural network is worth trying, and why: (a) predicting the diagnosis of a skin lesion from 20,000 labelled photographs; (b) predicting which of 300 survey respondents will change jobs, from 12 questionnaire scores; (c) predicting next week’s hospital admissions from ten years of daily records. No code is needed; write your answers on paper or in a document.
Work on your own computer
The project contains the data and all four parts of this page as an R script. Its Practise part uses tidymodels exactly as the chapter does; the other parts use base R and nnet, as on this page. Answers to the questions are in solutions.R, and the open tasks have space to write your code.
- Download chapter15.zip, unzip it, and double-click
chapter15.Rproj. - Or type this one line in RStudio’s Console:
usethis::use_course("https://polla-fattah.github.io/data2thesis_r/playground/chapter15.zip")