15 Neural Networks

Neural networks are the method behind most of what is called artificial intelligence today: image recognition, speech recognition, translation, and the large language models of Chapter 18. Their reputation leads many researchers to assume that a neural network must predict better than a traditional model on any data. The assumption deserves to be tested rather than accepted, and testing it requires understanding what a neural network is. The answer is less mysterious than the reputation suggests: a neural network is built from units that each perform a logistic regression, and it learns by repeatedly adjusting its weights to reduce its errors.

This chapter builds a neural network from a single “neuron”, shows by hand what training does, fits a network to the dropout data with the same tidymodels workflow as Chapters 11 and 12, and compares it fairly with the simpler models. It ends with deep learning: what it is, when it helps, and where to learn more. In the story, a member of Elaf’s research group asks why she does not simply use AI, since a neural network would surely do better than logistic regression, and her supervisor suggests that she find out, carefully.

TipBy the end of this chapter you will be able to
  • Explain how a neuron combines its inputs, and why a single neuron with a sigmoid activation is a logistic regression.
  • Describe a neural network’s layers, weights, and activation functions, and how it is trained.
  • Fit and tune a neural network with tidymodels (mlp()), using weight decay to prevent overfitting.
  • Compare a neural network fairly with simpler models, and judge when one is worth using.
  • Explain what deep learning is, and when it is (and is not) the right tool.

15.1 A single neuron

A neuron (or unit) in a neural network does three things. It multiplies each input by a weight, adds up the results together with a constant called the bias, and passes the total through an activation function, which turns it into the neuron’s output.

Take a neuron with two inputs, a student’s stress and support scores, with weights 1.2 and −0.8 and a bias of −4. For a student with stress 4 and support 2, the total is

\[ -4 + 1.2 \times 4 - 0.8 \times 2 = -0.8. \]

The activation function here is the sigmoid (also called the logistic function), which squeezes any number into the range 0 to 1:

\[ \text{sigmoid}(z) = \frac{1}{1 + e^{-z}}. \]

In R:

sigmoid <- function(z) 1 / (1 + exp(-z))

neuron <- function(stress, support) {
  sigmoid(-4 + 1.2 * stress - 0.8 * support)
}

neuron(stress = 4, support = 2)
[1] 0.3100255
neuron(stress = 2, support = 4)
[1] 0.008162571

The high-stress, low-support student gets an output of 0.31; the low-stress, high-support student, 0.008. If the output is read as the probability of considering dropout, this neuron is exactly a logistic regression (Chapter 8): the bias is the intercept, the weights are the coefficients, and the sigmoid turns the log odds into a probability. Everything a single neuron can do, Chapter 8 has already done.

15.2 From neurons to networks

The power of neural networks comes from connecting many neurons in layers. In the most common design, the multilayer perceptron, the inputs feed into a hidden layer of neurons, each with its own weights; the outputs of the hidden neurons then feed into an output layer, which produces the prediction. Figure 15.1 shows a network with three inputs and three hidden neurons.

flowchart LR
  I1((Stress)) --> H1((H1))
  I1 --> H2((H2))
  I1 --> H3((H3))
  I2((Support)) --> H1
  I2 --> H2
  I2 --> H3
  I3((Sleep)) --> H1
  I3 --> H2
  I3 --> H3
  H1 --> O((P of<br/>dropout))
  H2 --> O
  H3 --> O
Figure 15.1: A neural network with one hidden layer. Every arrow carries a weight; each hidden neuron and the output neuron add up their weighted inputs and apply an activation function.

Each hidden neuron learns its own combination of the inputs: one might respond to high stress, another to the combination of low support and short sleep. The output neuron then combines these. Because each neuron applies a curved activation function, the network as a whole can represent curved relationships and interactions that a single logistic regression cannot. With enough hidden neurons, a network can approximate almost any relationship between inputs and output. That flexibility is its strength, and, as you will see, its danger.

The activation function matters. Figure 15.2 shows the two most common: the sigmoid, used in this chapter, and the ReLU (rectified linear unit), which simply replaces negative totals with zero and is the standard choice in deep networks.

Two panels. Left: the sigmoid, an S-shaped curve rising from 0 on the left to 1 on the right, crossing 0.5 at an input of 0. Right: the ReLU, flat at 0 for negative inputs, then a straight line rising for positive inputs.
Figure 15.2: Two activation functions: the sigmoid squeezes any input into the range 0 to 1; the ReLU sets negative inputs to 0 and passes positive inputs unchanged.

15.2.1 How a network learns

A network starts with small random weights, so its first predictions are useless. Training adjusts the weights step by step to reduce the prediction error on the training data. At each step, an algorithm works out, for every weight, whether a small increase or decrease would reduce the error (the gradient), and moves all the weights a little in the helpful direction. This is gradient descent. One pass through the training data is an epoch, and training usually runs for hundreds of epochs.

The idea is clearest when one step is worked by hand. Take a single neuron, which is a logistic regression, and six students, with their stress measured from the group’s average:

six <- data.frame(stress  = c(2.0, 2.5, 3.0, 3.5, 4.0, 4.5) - 3.25,
                  dropout = c(0,   0,   1,   0,   1,   1))

The error to be reduced is the log loss, the measure of fit that logistic regression uses: it is small when the neuron gives high probabilities to the students who did consider dropping out and low probabilities to those who did not. Training starts with both weights at zero, so every student receives a probability of 0.5:

log_loss <- function(b0, b1) {
  p <- sigmoid(b0 + b1 * six$stress)
  -mean(six$dropout * log(p) + (1 - six$dropout) * log(1 - p))
}

b0 <- 0
b1 <- 0
log_loss(b0, b1)
[1] 0.6931472

For this neuron, the gradient has a simple form: for each weight, the average of the prediction errors (predicted probability minus the actual outcome), each multiplied by the input that the weight belongs to (1 for the bias):

p <- sigmoid(b0 + b1 * six$stress)
gradient_b0 <- mean(p - six$dropout)
gradient_b1 <- mean((p - six$dropout) * six$stress)
c(gradient_b0, gradient_b1)
[1]  0.0000000 -0.2916667

The gradient for the stress weight is negative, which means that increasing the weight would reduce the error: the students who considered dropping out have higher stress, and the neuron does not yet know it. One step of gradient descent moves each weight against its gradient, by an amount set by the learning rate:

learning_rate <- 1
b0 <- b0 - learning_rate * gradient_b0
b1 <- b1 - learning_rate * gradient_b1
c(b0 = b0, b1 = b1, loss = log_loss(b0, b1))
       b0        b1      loss 
0.0000000 0.2916667 0.6157970 

After one step, the stress weight has become positive and the loss has fallen from 0.693 to 0.616. Training is nothing more than this step, repeated. The loop below repeats it 2,000 times and records the loss:

b0 <- 0
b1 <- 0
loss_history <- numeric(2000)
for (step in 1:2000) {
  p  <- sigmoid(b0 + b1 * six$stress)
  b0 <- b0 - learning_rate * mean(p - six$dropout)
  b1 <- b1 - learning_rate * mean((p - six$dropout) * six$stress)
  loss_history[step] <- log_loss(b0, b1)
}
rbind(gradient_descent = c(b0, b1),
      glm = coef(glm(dropout ~ stress, data = six, family = binomial)))
                   (Intercept)   stress
gradient_descent  6.111907e-17 2.428055
glm              -1.790641e-16 2.428055
ggplot(data.frame(step = 1:2000, loss = loss_history), aes(x = step, y = loss)) +
  geom_line(colour = "#2f6793", linewidth = 1) +
  scale_x_log10() +
  labs(x = "Training step", y = "Log loss") +
  theme_minimal(base_size = 12)
A curve of log loss against training step on a logarithmic scale. It starts at about 0.69, falls steeply over the first hundred steps, and flattens out by about step 1,000.
Figure 15.3: Log loss of a single neuron during 2,000 steps of gradient descent (logarithmic horizontal axis). The loss falls quickly at first and then levels off as the weights approach their best values.

After 2,000 small steps, the weights found by gradient descent are the coefficients that R’s glm() calculates for the same logistic regression. A neural network with thousands of weights is trained in exactly this way. The gradients are harder to write down, and a method called backpropagation calculates them efficiently, but each step still moves every weight a little in the direction that reduces the error.

Two consequences follow. First, because training starts from random weights, two runs can give slightly different networks, so set.seed() matters. Second, a network with many weights can keep reducing its training error long after it has learned the real pattern, by memorising the noise. The usual defence is weight decay: a penalty on large weights, exactly like the ridge penalty of Chapter 13. In tidymodels it is called penalty.

15.3 A neural network for the dropout question

The data, split, recipe, and folds are those of Chapters 11 and 12. Neural networks need normalised inputs, like the penalised regressions of Chapter 13, and the recipe already provides them.

library(tidymodels)
tidymodels_prefer()

scores <- questionnaire |>
  mutate(
    stress_4     = 6 - stress_4,
    stress       = rowMeans(pick(stress_1:stress_6), na.rm = TRUE),
    burnout      = rowMeans(pick(burnout_1:burnout_6), na.rm = TRUE),
    support      = rowMeans(pick(support_1:support_6), na.rm = TRUE),
    satisfaction = rowMeans(pick(satisfaction_1:satisfaction_4), na.rm = TRUE)
  ) |>
  select(student_id, stress, burnout, support, satisfaction)

dropout_data <- students |>
  left_join(scores, join_by(student_id)) |>
  left_join(semesters |> filter(semester == 1) |> select(-semester), join_by(student_id)) |>
  mutate(considering_dropout = factor(considering_dropout, levels = c("Yes", "No"))) |>
  select(-student_id, -supervisor_id, -workshop, -workshop_sessions)

set.seed(2026)
dropout_split <- initial_split(dropout_data, prop = 0.75, strata = considering_dropout)
dropout_train <- training(dropout_split)
dropout_test  <- testing(dropout_split)

dropout_recipe <- recipe(considering_dropout ~ ., data = dropout_train) |>
  step_impute_median(all_numeric_predictors()) |>
  step_dummy(all_nominal_predictors()) |>
  step_normalize(all_numeric_predictors())

set.seed(2026)
dropout_folds <- vfold_cv(dropout_train, v = 10, strata = considering_dropout)

In tidymodels, a multilayer perceptron is mlp(). Its main arguments are hidden_units (the number of neurons in the hidden layer), penalty (the weight decay), and epochs. The nnet engine, which comes with R, fits networks with one hidden layer; MaxNWts raises its limit on the number of weights.

15.3.1 A network that memorises

A network with 20 hidden neurons and no weight decay comes first, as a warning:

big_net <- mlp(hidden_units = 20, penalty = 0, epochs = 1000) |>
  set_engine("nnet", MaxNWts = 5000) |>
  set_mode("classification")

set.seed(2026)
big_fit <- fit(workflow(dropout_recipe, big_net), data = dropout_train)

augment(big_fit, new_data = dropout_train) |> roc_auc(considering_dropout, .pred_Yes)
# A tibble: 1 × 3
  .metric .estimator .estimate
  <chr>   <chr>          <dbl>
1 roc_auc binary             1
augment(big_fit, new_data = dropout_test) |> roc_auc(considering_dropout, .pred_Yes)
# A tibble: 1 × 3
  .metric .estimator .estimate
  <chr>   <chr>          <dbl>
1 roc_auc binary         0.641

A perfect AUC on the training data, and 0.64 on new students: the same overfitting as the one-nearest-neighbour model of Chapter 11. The network has 521 weights and only 449 training students, so it has more than enough freedom to memorise every one of them.

15.3.2 Tuning size and weight decay

The number of hidden neurons and the weight decay are hyperparameters, tuned with cross-validation as in Chapters 11 to 13. Weight decay is tried on a logarithmic scale from 0.01 to about 30:

net_spec <- mlp(hidden_units = tune(), penalty = tune(), epochs = 500) |>
  set_engine("nnet", MaxNWts = 5000) |>
  set_mode("classification")

net_wf <- workflow(dropout_recipe, net_spec)

net_grid <- expand_grid(hidden_units = c(1, 3, 5, 10, 20),
                        penalty = 10^seq(-2, 1.5, by = 0.5))

set.seed(2026)
net_tuning <- tune_grid(net_wf, resamples = dropout_folds, grid = net_grid,
                        metrics = metric_set(roc_auc))
autoplot(net_tuning) + theme_minimal(base_size = 12)
Line chart of AUC against weight decay on a logarithmic scale, with one line per number of hidden units. With little weight decay, the AUC is lower and varies, lowest for the larger networks. With more weight decay, the lines rise and level off at a similar value for networks of every size.
Figure 15.4: Cross-validated ROC AUC of neural networks with different numbers of hidden neurons and amounts of weight decay.

Figure 15.4 tells a clear story. With little weight decay, the networks overfit, and the bigger networks overfit most. With enough weight decay, networks of every size do about equally well. The penalty, not the size, is what matters most here.

show_best(net_tuning, metric = "roc_auc", n = 3)
# A tibble: 3 × 8
  hidden_units penalty .metric .estimator  mean     n std_err .config         
         <dbl>   <dbl> <chr>   <chr>      <dbl> <int>   <dbl> <chr>           
1           10    3.16 roc_auc binary     0.804    10  0.0300 pre0_mod30_post0
2           10    1    roc_auc binary     0.804    10  0.0298 pre0_mod29_post0
3           20    3.16 roc_auc binary     0.803    10  0.0299 pre0_mod38_post0
best_net <- select_best(net_tuning, metric = "roc_auc")

15.3.3 A fair comparison

The best network has a cross-validated AUC of 0.804. For comparison, logistic regression on the same folds:

set.seed(2026)
logistic_cv <- fit_resamples(workflow(dropout_recipe, logistic_reg()),
                             resamples = dropout_folds, metrics = metric_set(roc_auc))
collect_metrics(logistic_cv)
# A tibble: 1 × 6
  .metric .estimator  mean     n std_err .config        
  <chr>   <chr>      <dbl> <int>   <dbl> <chr>          
1 roc_auc binary     0.804    10  0.0303 pre0_mod0_post0

The two are practically identical: 0.804 for the network and 0.804 for logistic regression. The final test on the held-out students confirms it:

set.seed(2026)
final_net <- net_wf |>
  finalize_workflow(best_net) |>
  last_fit(dropout_split, metrics = metric_set(roc_auc, accuracy))
collect_metrics(final_net)
# A tibble: 2 × 4
  .metric  .estimator .estimate .config        
  <chr>    <chr>          <dbl> <chr>          
1 accuracy binary         0.841 pre0_mod0_post0
2 roc_auc  binary         0.853 pre0_mod0_post0

The network reaches a test AUC of 0.85, against 0.84 for logistic regression on the same test students (Chapter 11). With only 151 test students, of whom 23 considered dropping out, a difference of this size is well within chance. The answer to the research group is therefore that a neural network does not predict dropout better. Chapters 12 and 13 found the same for random forests, support vector machines, and boosting. The wellbeing data has a smooth pattern, which logistic regression captures fully, leaving nothing extra for a more flexible model to find.

And the neural network has real costs. Even the tuned network, with 10 hidden neurons, has 261 weights, which cannot be interpreted: there is no equivalent of an odds ratio saying how much each predictor matters. Without strong weight decay, its results depend on the random starting weights. And it needed careful tuning to avoid overfitting. Logistic regression is kept, and the thesis reports that a neural network was tried and did not improve on it, which is itself a useful finding.

15.4 When neural networks are worth using

Neural networks are not magic, but they are not overrated either; they are powerful in particular situations. They excel with unstructured data, such as images, sound, and text, where the raw inputs (pixels, sound samples, words) mean little on their own and the useful features must be learned from them; here they outperform every other method by a wide margin. They need very large datasets, with many thousands or millions of cases, to fit models with thousands of weights reliably. And they can capture complex, non-linear relationships that simpler models miss, as the “In your field” box below shows.

For typical research data, a table with a few hundred cases and a few dozen variables, well-tuned regression, random forests, and boosting usually do as well or better, and are easier to explain. That is the situation of most theses.

15.5 Deep learning

Deep learning means neural networks with many hidden layers, often dozens or hundreds, and millions or billions of weights. Each layer learns features built on the features of the layer before: in an image network, the first layers detect edges, later layers shapes, and the last layers whole objects. Special architectures suit special data: convolutional networks for images, and transformers for text. The large language models behind ChatGPT and Claude (Chapter 18) are transformer networks with billions of weights, trained on enormous amounts of text.

Deep learning needs more data, more computing power (often a graphics card), and more technical setup than anything else in this book. In R, there are two main tools. The keras3 package is an interface to the Keras library, which runs on Python’s TensorFlow, JAX, or PyTorch; it is the most widely documented route, but it requires Python to be installed alongside R. The torch package runs the same engine as Python’s PyTorch without needing Python, and the brulee package lets tidymodels fit deep networks through torch, with the same mlp() interface used in this chapter.

For most researchers, the practical way to use deep learning is not to train a network from scratch, but to use a pretrained model, already trained by someone else on huge datasets, for a task such as transcribing interviews, recognising objects in images, or classifying text. Chapter 18 does exactly this with a language model, to code the study’s open-ended survey answers.

NoteIn your field: engineering

The concrete data of Chapter 13 has the kind of curved, interacting relationships where a neural network can shine: strength rises steeply over the first weeks and then levels off, and the effect of each ingredient depends on the others. For regression, mlp() is used with set_mode("regression"):

data(concrete, package = "modeldata")
set.seed(1)
concrete_split <- initial_split(concrete)

concrete_recipe <- recipe(compressive_strength ~ ., data = training(concrete_split)) |>
  step_log(age) |>
  step_normalize(all_numeric_predictors())

concrete_net <- mlp(hidden_units = 10, penalty = 0.1, epochs = 1000) |>
  set_engine("nnet", MaxNWts = 5000) |>
  set_mode("regression")

set.seed(1)
concrete_lm_fit <- last_fit(workflow(concrete_recipe, linear_reg()), concrete_split)
set.seed(1)
concrete_net_fit <- last_fit(workflow(concrete_recipe, concrete_net), concrete_split)
collect_metrics(concrete_lm_fit)
# A tibble: 2 × 4
  .metric .estimator .estimate .config        
  <chr>   <chr>          <dbl> <chr>          
1 rmse    standard       7.21  pre0_mod0_post0
2 rsq     standard       0.815 pre0_mod0_post0
collect_metrics(concrete_net_fit)
# A tibble: 2 × 4
  .metric .estimator .estimate .config        
  <chr>   <chr>          <dbl> <chr>          
1 rmse    standard       5.04  pre0_mod0_post0
2 rsq     standard       0.909 pre0_mod0_post0

The step step_log(age) takes the logarithm of the age, which already straightens its curve for the linear model. Even so, the network’s prediction error is clearly smaller: an RMSE of 5.0 megapascals, against 7.2 for the linear model. Here, there is structure for its flexibility to capture.

15.6 Common misconceptions

The reputation of neural networks produces several misconceptions that a thesis should avoid.

  • “A neural network is always more accurate.” On typical tables of research data it often predicts no better than regression, as this chapter found.
  • “Neural networks learn in a mysterious way.” Training is gradient descent: many small steps that each reduce the error, as the six-student example showed.
  • “A network that fits the training data perfectly has learned the pattern.” A perfect training score is usually a sign of memorisation; only data the network has not seen can show what it learned.
  • “Neural networks and AI are the same thing.” Neural networks are one family of methods; large language models are a particular, very large kind of network.

15.7 Chapter review

15.7.1 Summary

  • A neuron multiplies its inputs by weights, adds a bias, and applies an activation function. A single neuron with a sigmoid activation is a logistic regression.
  • A multilayer perceptron connects neurons in layers: inputs, one or more hidden layers, and an output layer. Hidden neurons with curved activations let the network represent curves and interactions.
  • Networks are trained by gradient descent: starting from random weights, each step moves every weight a little in the direction that reduces the error. For a single neuron, gradient descent arrives at the logistic regression coefficients. With many weights, networks overfit easily; weight decay (a penalty on large weights) prevents it.
  • In tidymodels, mlp() with the nnet engine fits a network with one hidden layer; tune hidden_units and penalty with cross-validation.
  • On the dropout data, a tuned neural network predicts no better than logistic regression, and is harder to interpret. Neural networks shine on images, sound, and text, on very large datasets, and on strongly non-linear problems.
  • Deep learning uses many layers; in R, it is available through keras3 (with Python) and torch. Pretrained models are often the practical route for researchers.

15.7.2 Key terms

Neural network, neuron (unit), weight, bias, activation function, sigmoid, ReLU, multilayer perceptron, input layer, hidden layer, output layer, training, log loss, gradient, gradient descent, learning rate, epoch, weight decay, deep learning, convolutional network, transformer, pretrained model.

15.8 Exercises

The playground has these and more, with hints and solutions.

  1. Change the neuron’s weights so that support matters twice as much as in the chapter, recalculate the outputs for the two students, and explain what a negative weight means.
  2. Fit a network with hidden_units = 1 and a weight decay of 1, compare its cross-validated AUC with logistic regression, and explain why they might be so similar.
  3. Fit the chapter’s best network three times with different seeds, then do the same with a weight decay of 0.01. Explain when the seed matters, and why.
  4. Train the 20-neuron network without weight decay for 10, 100, and 1000 epochs, and describe how the training and test AUCs change.
  5. In two or three sentences, explain to the research group why the thesis does not use a neural network for the dropout model.
  6. Repeat the gradient-descent loop with a learning rate of 0.1 and of 10. Describe how the loss curve changes, and explain what goes wrong when the learning rate is too large.

15.9 Further reading

  • An Introduction to Statistical Learning (James et al. 2021) has a clear chapter on deep learning, including convolutional networks, with the ideas explained through examples.
  • Deep Learning with R (Chollet et al. 2022) is the practical guide to keras in R, by the creator of Keras and the authors of its R interface.
  • Deep Learning (Goodfellow et al. 2016) is the standard, more mathematical textbook, free to read online.
  • Tidy Modeling with R (Kuhn and Silge 2022) shows how neural networks fit into the tidymodels workflow alongside other models.

References

Chollet, François, Tomasz Kalinowski, and J. J. Allaire. 2022. Deep Learning with r. 2nd ed. Manning.
Goodfellow, Ian, Yoshua Bengio, and Aaron Courville. 2016. Deep Learning. MIT Press. https://www.deeplearningbook.org.
James, Gareth, Daniela Witten, Trevor Hastie, and Robert Tibshirani. 2021. An Introduction to Statistical Learning: With Applications in r. 2nd ed. Springer. https://doi.org/10.1007/978-1-0716-1418-1.
Kuhn, Max, and Julia Silge. 2022. Tidy Modeling with r: A Framework for Modeling in the Tidyverse. O’Reilly Media. https://www.tmwr.org.