Appendix B: Glossary

This appendix collects the key terms from every chapter’s review, with a short definition of each and the chapters where it is introduced or used, followed by a table of the R functions used most often in the book. Definitions are written in plain language; the chapters give the full explanations and examples.

Key terms

A

accuracy
The share of cases a classification model classifies correctly. Misleading when one outcome is rare. (Chapter 11)
activation function
The function a neuron applies to its weighted total to produce its output, such as the sigmoid or the ReLU. (Chapter 15)
adjusted Rand index
A measure of agreement between two clusterings or classifications, corrected for chance: 1 means identical groupings, 0 means chance agreement. (Chapter 14)
adjusted R-squared
R-squared corrected for the number of predictors, so that adding useless predictors does not make a model look better. (Chapter 8)
AI coding assistant
An AI tool that writes, explains, or completes code, in a chat or inside the editor. Its suggestions must be checked. (Chapter 18)
AIC
Akaike information criterion: a measure for comparing models that balances fit against complexity; lower is better. (Chapter 10)
alternative hypothesis
The claim that there is an effect or a difference, tested against the null hypothesis. (Chapters 5, 7)
annotation
Text, arrows, or shapes added to a plot to point out something specific. (Chapter 4)
anonymisation
Removing or altering information so that the people in a dataset cannot be identified. (Chapter 17)
ANOVA
Analysis of variance: a test of whether the means of three or more groups differ, by comparing variation between groups with variation within them. (Chapter 8)
API
Application programming interface: a way for one program to send requests to another, such as R sending text to a language model. (Chapter 18)
API key
A secret code that identifies you to an online service’s API. Keep it in .Renviron, never in a script. (Chapter 18)
argument
A value given to a function inside its brackets, such as na.rm = TRUE, that controls what the function does. (Chapter 1)
ARIMA
A family of time series models that forecast from the autocorrelation of the series and past random shocks. (Chapter 16)
artificial intelligence
Computer systems that perform tasks usually needing human intelligence, such as understanding language; in this book, mainly large language models. (Chapter 18)
aspect ratio
The ratio of a graph’s width to its height. It changes how steep lines and slopes appear. (Chapter 4)
assignment
Storing a value in an object with <-, as in x <- 5. (Chapter 1)
attrition
The loss of participants from a study over time. (Chapters 5, 6)
autocorrelation
The correlation between a time series and itself a number of steps (lags) earlier. (Chapter 16)

B

bar chart
A plot of counts or values as bars, one per category. (Chapter 4)
baseline model
The simplest possible model, such as predicting the mean for everyone, used as a benchmark for real models. (Chapter 13)
Bayesian information criterion (BIC)
A criterion for comparing models that rewards fit and penalises complexity. In mclust, higher BIC is better. Also: BIC. (Chapter 14)
benchmark forecast
A simple forecast, such as the mean or the value a season earlier, that any serious method should beat. (Chapter 16)
between-group variation
How much group means differ from the overall mean; the numerator of the ANOVA F statistic. (Chapter 8)
bias
In a neural network, the constant added to a neuron’s weighted total (like an intercept). More generally, a systematic error. (Chapters 5, 15)
bias-variance trade-off
The balance between a model too simple to capture the pattern (high bias) and one so flexible that it follows the noise of each sample (high variance). Error on new data is lowest in between. (Chapter 11)
BibTeX
A text format for bibliographic references, used by Quarto to format citations and reference lists. (Chapter 17)
bimodal
Having two peaks. (Chapter 6)
bin
One of the intervals into which a histogram divides the values. (Chapter 4)
boosting
Building a model by adding many small models (usually trees) one at a time, each correcting the errors of the model so far. (Chapter 13)
bootstrap
Estimating uncertainty by repeatedly resampling the data with replacement and recalculating a statistic. (Chapter 7)
bootstrap sample
A sample of the same size as the data, drawn from it at random with replacement. (Chapter 12)
border point
In DBSCAN, a case near a core point that belongs to its cluster but is not itself a core point. (Chapter 14)
box plot
A plot that shows the median, the quartiles, and unusual values of a variable. (Chapter 4)
broom
A package that turns model results into tidy data frames, with tidy(), glance(), and augment(). (Chapter 19)

C

case
One unit that is measured in a study, such as a student, a patient, or a school; one row of a data table. (Chapter 2)
causal question
A research question about what leads to what. It needs a design that rules out other explanations, ideally an experiment. (Chapter 5)
central limit theorem
The result that the sampling distribution of a mean is approximately normal for large enough samples, whatever the shape of the data. (Chapter 7)
centring
Subtracting the mean from a variable, so that zero means “average”; often used before fitting interactions. (Chapter 8)
chain of reasoning
The sequence of steps behind a research conclusion: question, hypothesis, design, variables, analysis, result, conclusion, and limitation. (Chapter 19)
chi-square test
A test for categorical data: whether counts fit expected proportions (goodness of fit), or whether two categorical variables are related (independence). (Chapter 7)
chunk option
A setting for a code chunk in a Quarto document, written as #| option: value, such as echo: false. (Chapter 17)
citation
A reference to a source in the text, written in Quarto as [@key]. (Chapter 17)
citation style (CSL)
A file in the Citation Style Language that sets how citations and references are formatted, such as APA. (Chapter 17)
class weights
Making errors on a rare class count more when a model is fitted, one remedy for imbalanced outcomes. (Chapter 12)
classification
Predicting a category, such as whether a student will consider dropping out. (Chapter 11)
cluster analysis
Methods that group cases so that cases in the same group are similar to each other. (Chapter 9)
cluster sampling
Sampling whole groups, such as all the students of randomly chosen supervisors. (Chapter 5)
code chunk
A block of code in a Quarto document that is run when the document is rendered. (Chapter 17)
codebook
A description of every variable in a dataset or, in qualitative coding, of every theme and how to assign it. (Chapters 2, 17, 18)
coefficient path
A plot of how each coefficient of a regularised model changes as the penalty changes. (Chapter 13)
Cohen’s d
An effect size for the difference between two means, in standard deviations. (Chapter 7)
Cohen’s kappa
A measure of agreement between two coders or classifiers, corrected for the agreement expected by chance. (Chapter 18)
comment
Text in a script after #, which R ignores; used to explain the code. (Chapter 1)
commit
In git, a saved snapshot of a project, with a message describing the change. (Chapter 17)
confidence interval
A range of plausible values for a population quantity, calculated from a sample, with a stated level of confidence such as 95%. (Chapter 7)
confirmatory analysis
An analysis planned before seeing the data, to test a specific hypothesis. (Chapters 5, 17)
confirmatory factor analysis
Factor analysis that tests a factor structure specified in advance. (Chapter 9)
confounder
A variable related to both the predictor and the outcome, which can create or hide an apparent effect. (Chapters 5, 6, 8)
confusion matrix
A table of predicted against actual classes, showing true and false positives and negatives. (Chapter 12)
Console
The RStudio pane where R commands are run and results appear. (Chapter 1)
construct
An idea that cannot be observed directly, such as stress or wellbeing, and must be measured indirectly. (Chapters 5, 9)
construct validity
Whether a score behaves as the construct should: related to similar measures, less related to different ones. (Chapter 5)
content validity
Whether the items of a measure cover the whole construct. (Chapter 5)
convenience sampling
Sampling whoever is easy to reach. Common, and the weakest basis for generalising. (Chapter 5)
convolutional network
A type of deep neural network designed for images. (Chapter 15)
core point
In DBSCAN, a case with at least minPts cases within distance eps. (Chapter 14)
correlation coefficient
A number from −1 to 1 measuring the strength and direction of a straight-line relationship between two variables. (Chapter 6)
correlation matrix
A table of the correlations between every pair of variables. (Chapter 9)
count outcome
An outcome that counts events (0, 1, 2, …), usually modelled with a Poisson model. (Chapter 10)
covariance matrix
A table of the variances and covariances of several variables; in a mixture model it sets the shape of a cluster. (Chapter 14)
Cronbach’s alpha
A measure of the internal consistency (reliability) of a scale made of several items. (Chapter 9)
cross-loading
An item that loads substantially on more than one factor. (Chapter 9)
cross-reference
A reference in a Quarto document, such as @fig-wellbeing, that becomes a numbered link. (Chapter 17)
cross-sectional design
A design that measures every variable once, at one time. (Chapter 5)
cross-validation
Estimating how well a model predicts new data by repeatedly fitting it on part of the training data and testing it on the rest. (Chapter 11)
CSV file
A plain-text file of comma-separated values, the most common format for sharing data tables. (Chapter 2)

D

data cleaning
Preparing raw data for analysis: removing test and duplicate cases, making categories consistent, marking missing values, and correcting or removing impossible values, with every decision recorded. (Chapter 3)
data frame
R’s table of data: columns are variables, rows are cases. (Chapters 1, 2)
data leakage
Information from the test data reaching the model during training, which makes its performance look better than it is. (Chapter 11)
data matrix
The arrangement of data as a table with one row per case and one column per variable. (Chapter 2)
data structure
The way data is organised in R: vector, factor, data frame, matrix, or list. (Chapter 2)
data type
The kind of value: numeric, integer, character, or logical. (Chapter 1)
DBSCAN
A clustering method that finds dense regions of cases and labels cases in sparse regions as noise. (Chapter 14)
decision tree
A model that predicts by a series of yes-or-no questions about the predictors. (Chapters 11, 12)
decomposition
Splitting a time series into trend, seasonal, and remainder components. (Chapter 16)
deep learning
Neural networks with many hidden layers, used for images, sound, and text. (Chapter 15)
dendrogram
The tree diagram produced by hierarchical clustering. (Chapter 9)
density
In clustering, how closely packed cases are in a region. (Chapter 14)
density plot
A smooth version of a histogram, showing the shape of a distribution. (Chapter 4)
descriptive question
A research question about what is: how much, how many, how often. (Chapter 5)
descriptive statistics
Numbers that summarise data, such as means, medians, and standard deviations. (Chapter 6)
design effect
How many times more grouped observations are needed to give the same information as independent ones: 1 + (m - 1) × ICC, where m is the group size. (Chapter 10)
deviation
The distance of a value from the mean, \(x_i - \bar{x}\). The standard deviation summarises the deviations of all values. (Chapter 6)
diagnostic plot
A plot used to check a model’s assumptions, such as residuals against fitted values. (Chapter 8)
dictionary method
Classifying texts by counting keywords from a list for each category. (Chapter 18)
diminishing returns
A relationship that levels off, so each extra unit of the predictor adds less than the one before. (Chapter 8)
directional hypothesis
A hypothesis that predicts the direction of an effect, such as “higher” or “lower”. (Chapter 5)
disclosure
Stating in a publication how AI tools or other aids were used. (Chapter 18)
distance
How different two cases are, calculated from their values on several variables. (Chapters 9, 12)
diverging palette
Colours running in two directions away from a meaningful midpoint, such as zero, used for values that can be positive or negative. (Chapter 4)
DOI
Digital object identifier: a permanent identifier for a publication or dataset, such as 10.1126/science.1213847. (Chapter 17)
downsampling
Balancing classes by randomly removing cases of the common class from the training data. (Chapter 12)
dpi
Dots per inch: the resolution of a saved image; 300 dpi is usual for print. (Chapter 4)
dummy variable
A 0/1 variable representing one category of a categorical predictor. (Chapter 11)

E

effect size
A measure of how large an effect is, independent of sample size, such as Cohen’s d or eta squared. (Chapter 7)
eigenvalue
In PCA, the amount of variance captured by a component. (Chapter 9)
elastic net
Regularised regression that mixes the ridge and lasso penalties. (Chapter 13)
elbow method
Choosing the number of clusters where adding more stops reducing within-cluster distance much. (Chapter 9)
ensemble
A model that combines the predictions of many models, such as a random forest. (Chapter 12)
epoch
One pass through the training data when training a neural network. (Chapter 15)
eps
In DBSCAN, the radius of the neighbourhood around each case. (Chapter 14)
error message
R’s message when something goes wrong, which usually says what and where. (Chapter 2)
eta squared
An effect size for ANOVA: the share of the variation in the outcome explained by the groups. (Chapter 8)
expected count
In a chi-square test, the count expected in a cell if the null hypothesis were true. (Chapter 7)
explanation
Using a model to understand why something happens, rather than to predict new cases. (Chapter 11)
explanatory graph
A carefully designed graph made for readers, to show a finding clearly and accurately. (Chapter 4)
exploratory analysis
Exploring data to find patterns and generate questions, rather than to test planned hypotheses. (Chapters 5, 6, 17)
exploratory graph
A quick graph made by the researcher to understand the data: to check distributions, find unusual values, and notice patterns. (Chapter 4)
exponential smoothing (ETS)
Forecasting with weighted averages of past observations, with more weight on recent ones; ETS models the error, trend, and season. Also: ETS. (Chapter 16)
external validation
Checking clusters or predictions against known categories or outcomes. (Chapter 14)
external validity
Whether a study’s results apply beyond it, to other people, places, and times. (Chapter 5)

F

F statistic
In ANOVA, the ratio of between-group to within-group variation. (Chapter 8)
F1 score
A single measure combining precision and recall (their harmonic mean). (Chapter 12)
facet
A small panel of a plot showing one subgroup; facet_wrap() makes one panel per group. (Chapter 4)
factor
In R, a categorical variable with a fixed set of levels. In factor analysis, a hidden (latent) variable behind several items. (Chapters 2, 9)
factor analysis
A method that explains the correlations among items by a smaller number of hidden factors. (Chapter 9)
false negative
A case that belongs to the positive class but is predicted negative. (Chapter 12)
false positive
A case predicted positive that belongs to the negative class. (Chapter 12)
falsifiability
The property of a claim that some possible result would contradict it. A hypothesis that fits every possible result tells us nothing. (Chapter 5)
fixed effect
In a mixed model, an effect assumed the same for everyone, such as the average change over time. (Chapter 10)
fold
One of the parts into which the data is split for cross-validation. (Chapter 11)
forecast horizon
How far ahead a forecast is made. (Chapter 16)
Fourier terms
Pairs of sine and cosine waves used as predictors to describe a seasonal pattern. (Chapter 16)
function
A named piece of code that takes arguments and returns a result, such as mean(). (Chapter 1)

G

garden of forking paths
The many reasonable choices in an analysis. Choosing among them after seeing the results inflates the chance of a false positive. (Chapter 17)
Gaussian mixture model
A model that describes data as a mix of normal distributions, giving each case a probability of belonging to each cluster. (Chapter 14)
generalisation
The ability of a model to perform well on new data, not only on the data it learned from. (Chapter 11)
generalised linear mixed model
A mixed-effects model for outcomes that are not normal, such as counts or yes/no outcomes. (Chapter 10)
geom
In ggplot2, the geometric shape that represents data, such as points, lines, or bars. (Chapter 4)
Gini impurity
A measure of how mixed the classes are in a group; decision trees choose splits that reduce it. (Chapter 12)
git
The standard version control system, which records the history of a project. (Chapter 17)
GitHub
A website for storing git projects online, sharing them, and working on them together. (Chapter 17)
gradient
The direction and rate at which the error changes as each weight changes. (Chapter 15)
gradient descent
Training a model by repeatedly moving the weights a small step in the direction that reduces the error. (Chapter 15)
grammar of graphics
The idea behind ggplot2: a plot is built from data, aesthetic mappings, and geometric layers. (Chapter 4)
graphical perception
How people read values from graphs. Positions and lengths are judged most accurately, then angles, areas, and colours. (Chapter 4)
grouped summary
Summary statistics calculated separately for each group, for example with summarise(.by = ...). (Chapter 3)

H

hallucination
Plausible but invented content produced by a language model, such as a function or reference that does not exist. (Chapter 18)
heat map
A grid of coloured cells showing values, such as a correlation matrix. (Chapter 4)
helper function
A small function written to avoid repeating the same code, such as one that formats means. (Chapter 19)
hidden layer
A layer of neurons between the inputs and the output of a neural network. (Chapter 15)
hierarchical clustering
Clustering that repeatedly merges the most similar groups, producing a dendrogram. (Chapter 9)
histogram
A plot of the distribution of a numeric variable, as bars counting the values in each bin. (Chapter 4)
hyperparameter
A setting of a model that is chosen before fitting rather than learned from the data, such as the number of neighbours in k-NN. (Chapter 11)
hypothesis
A prediction, stated before the data is analysed, of what the data will show; precise enough to be wrong. (Chapter 5)

I

identifier
A variable that uniquely identifies each case, such as student_id. (Chapter 3)
imbalanced outcome
An outcome in which one class is much rarer than the other. (Chapter 11)
imputation
Filling in missing values with estimated ones, such as the median. (Chapter 11)
independence
The assumption that observations do not influence each other; in a chi-square test, the hypothesis that two variables are unrelated. (Chapter 7)
index
In a tsibble, the column that holds time. (Chapter 16)
inferential analysis
Using a sample to draw conclusions about a population. (Chapter 6)
inline code
R code inside a sentence of a Quarto document, replaced by its result when rendered. (Chapter 17)
input
In Shiny, a control such as a menu or slider whose value the user chooses. (Chapter 17)
input layer
The inputs (predictors) of a neural network. (Chapter 15)
interaction
When the effect of one predictor depends on the value of another. (Chapter 8)
interaction plot
A plot of group means that shows whether the effect of one factor depends on another. (Chapter 8)
intercept
The predicted value of the outcome when all predictors are zero. (Chapter 8)
internal consistency
The extent to which the items of a scale agree with each other, often measured with Cronbach’s alpha. (Chapters 5, 9)
internal validation
Judging clusters by how compact and separated they are, using the data alone, as with the silhouette. (Chapter 14)
internal validity
Whether a study can rule out other explanations for its results. Highest in randomised experiments. (Chapter 5)
interquartile range
The range of the middle 50% of the values: the third quartile minus the first. (Chapter 6)
inter-rater reliability
The extent to which two people coding or rating the same material agree, often measured with Cohen’s kappa. (Chapters 5, 18)
interval
A level of measurement with equal distances between values but no true zero, such as a wellbeing index. (Chapter 5)
intraclass correlation
The share of the total variation that lies between groups (such as students or supervisors). (Chapter 10)

J

join
Combining two tables by matching rows on a key, such as student_id. (Chapter 3)

K

Kaiser rule
Keeping components with an eigenvalue above 1; a rough guide that often keeps too many. (Chapter 9)
kernel
In a support vector machine, a function that allows curved boundaries between classes. (Chapter 12)
k-means
A clustering method that assigns each case to the nearest of k cluster centres and moves the centres to the means. (Chapter 9)
k-nearest neighbours
Predicting a case from the k most similar cases in the training data. (Chapters 11, 12)
k-nearest-neighbour distance plot
A sorted plot of each case’s distance to its k-th nearest neighbour, used to choose eps for DBSCAN. (Chapter 14)
knitr
The R package that runs the code in Quarto and R Markdown documents. (Chapter 17)
Kruskal-Wallis test
A non-parametric alternative to one-way ANOVA. (Chapter 8)

L

lag
The number of time steps between an observation and an earlier one it is compared with. (Chapter 16)
large language model
A neural network trained on huge amounts of text to predict the next token, used in AI assistants. (Chapter 18)
lasso
Regularised regression with a penalty on the absolute size of the coefficients, which sets some of them to exactly zero. (Chapter 13)
latent variable
A variable that cannot be observed directly, such as stress, and is assumed to cause part of the answers to the items that measure it. (Chapter 9)
LaTeX
A typesetting system whose notation is used to write equations, as in $\bar{x}$. (Chapter 17)
layer
In ggplot2, one part of a plot added with +, such as a set of points or a trend line. In a neural network, a group of neurons. (Chapter 4)
leaf
A final node of a decision tree, which gives the prediction. (Chapter 12)
learning rate
In boosting and neural networks, how large a step each update takes. (Chapters 13, 15)
least squares
Choosing a regression line that minimises the sum of the squared residuals. (Chapter 8)
level
One of the categories of a factor. (Chapter 2)
level of measurement
What the values of a variable mean (nominal, ordinal, interval, or ratio), which decides the summaries and tests that make sense. (Chapter 5)
Levene’s test
A test of whether groups have equal variances. (Chapter 8)
licence
A statement of what others may do with shared data or code, such as CC BY or MIT. (Chapter 17)
likelihood ratio test
A test comparing two nested models by how much better the larger one fits. (Chapter 10)
line chart
A plot of values connected by lines, often over time. (Chapter 4)
linear regression
A model of a numeric outcome as a straight-line function of one or more predictors. (Chapter 8)
list
An R object that can hold elements of different types and sizes, such as the results of a test. (Chapter 2)
loading
The correlation between an item and a factor or component. (Chapter 9)
local model
A language model that runs on your own computer, so data does not leave it. (Chapter 18)
lockfile
A file, such as renv’s renv.lock, that records the exact package versions of a project. (Chapter 17)
log loss
The measure of error used to fit logistic regression and classification networks: small when high probabilities are given to the cases that did occur and low probabilities to those that did not. (Chapter 15)
logical indexing
Selecting elements with a TRUE/FALSE condition, as in x[x > 5]. (Chapter 2)
logistic regression
A regression model for a yes/no outcome, which predicts the probability of “yes”. (Chapter 8)
log-odds
The logarithm of the odds; the scale on which logistic regression coefficients are estimated. (Chapter 8)
long format
Data with one row per observation, such as one row per student per semester. (Chapter 3)
longitudinal design
A design that measures the same people repeatedly over time. (Chapter 5)

M

machine learning
Methods that learn patterns from data to make predictions or find structure, judged by performance on new data. (Chapter 11)
MAE
Mean absolute error: the average size of the prediction errors. (Chapter 13)
main effect
The effect of one factor, averaged over the levels of another. (Chapter 8)
Mann-Whitney U test
A non-parametric test comparing two independent groups. (Chapter 7)
mapping
In ggplot2, linking a variable to a visual property (an aesthetic), such as position or colour. Also: aesthetic mapping. (Chapter 4)
margin
In a support vector machine, the empty band between the classes and the boundary. (Chapter 12)
Markdown
A simple way of formatting text with symbols, such as **bold** and # Heading. (Chapter 17)
matrix
A two-dimensional table of values of one type. (Chapter 2)
mean
The average: the sum of the values divided by their number. (Chapter 6)
mean method
A benchmark forecast that predicts the average of past observations. (Chapter 16)
median
The middle value when the values are sorted. (Chapters 4, 6)
mediator
A variable on the path between a predictor and an outcome, through which the predictor has its effect. (Chapter 5)
membership probability
The probability that a case belongs to each cluster, as given by a mixture model. Also: soft assignment. (Chapter 14)
method choice
Choosing an analysis method from the goal, the type of outcome, and how observations are related. (Chapter 19)
minPts
In DBSCAN, the number of cases a neighbourhood must contain for a case to be a core point. (Chapter 14)
misclassification cost
The cost assigned to each kind of error of a classifier, such as missing a case or raising a false alarm. The costs determine the best threshold. (Chapter 12)
missing at random
Missingness that depends only on observed variables, not on the missing values themselves. (Chapter 6)
missing code
A value such as 99 or -9 used in raw data to mark a missing answer. (Chapter 3)
missing completely at random
Missingness unrelated to any variable, observed or not. (Chapter 6)
missing not at random
Missingness that depends on the missing values themselves. (Chapter 6)
missing value (NA)
R’s marker for a value that is not available. Also: NA. (Chapter 1)
mixed-effects model
A regression model with both fixed effects and random effects, for repeated or nested data. (Chapter 10)
mixing probability
In a mixture model, the share of cases belonging to a component. (Chapter 14)
mixture
In regularised regression, the mix of lasso and ridge penalties, from 0 (ridge) to 1 (lasso). (Chapter 13)
mixture component
One of the normal distributions in a Gaussian mixture model. (Chapter 14)
mode
The most common value. (Chapter 6)
model specification
In tidymodels, the description of a model (type, engine, mode) before it is fitted. (Chapter 11)
moderator
A variable that changes the strength or direction of the relationship between a predictor and an outcome. (Chapter 5)
mtry
In a random forest, the number of predictors each split may choose from. (Chapter 12)
multicollinearity
Strong correlation among predictors, which makes regression coefficients unstable. (Chapter 13)
multilayer perceptron
A neural network with one or more hidden layers of neurons. (Chapter 15)
multilevel model
Another name for a mixed-effects model, emphasising nested levels. (Chapter 10)
multiple regression
Linear regression with more than one predictor. (Chapter 8)
multiple testing
Running many tests in one study. Each test carries its own risk of a false alarm, so the chance of at least one grows with the number of tests. (Chapter 7)
multivariate analysis
Methods that analyse many variables at once, such as PCA and factor analysis. (Chapter 9)

N

naive method
A benchmark forecast that repeats the last observed value. (Chapter 16)
nested data
Data in which cases are grouped within units, such as students within supervisors. (Chapter 10)
neural network
A model made of connected neurons in layers, whose weights are learned from data. (Chapter 15)
neuron (unit)
The basic element of a neural network: it weights its inputs, adds a bias, and applies an activation function. Also: neuron, unit. (Chapter 15)
node
A point in a decision tree where a question is asked or a prediction is made. (Chapter 12)
noise
In DBSCAN, cases in sparse regions that belong to no cluster. (Chapter 14)
nominal
A level of measurement for categories with no order, such as faculty. (Chapter 5)
non-directional hypothesis
A hypothesis that predicts a difference or relationship without saying in which direction. (Chapter 5)
non-parametric test
A test that does not assume a particular distribution, often based on ranks. (Chapter 7)
normal distribution
The symmetric, bell-shaped distribution described by a mean and a standard deviation. (Chapter 6)
normalisation
Rescaling variables, usually to z-scores, so that they are on the same scale. (Chapter 11)
null distribution
The distribution of a statistic that would be expected if the null hypothesis were true, used to judge how surprising the observed result is. (Chapter 7)
null hypothesis
The claim of no effect or no difference, which a test tries to reject. (Chapters 5, 7)

O

object
A named value stored in R, such as a number, a vector, or a data frame. (Chapter 1)
oblique rotation
A factor rotation that allows the factors to correlate. (Chapter 9)
odds
The probability of an event divided by the probability of it not happening. (Chapter 8)
odds ratio
The ratio of the odds in two groups; in logistic regression, the multiplicative change in odds per unit of a predictor. (Chapter 8)
one-sample t-test
A test of whether a mean differs from a given value. (Chapter 7)
one-tailed test
A test that looks for a difference in one direction only. (Chapter 7)
open science
Practices that make research transparent and reusable: sharing data and code, preregistration, and open access. (Chapter 17)
operationalisation
The decision about how a construct is measured: which questions, records, or instruments turn it into a variable. (Chapter 5)
ordered factor
A factor whose levels have an order, created with factor(..., ordered = TRUE). (Chapter 5)
ordinal
A level of measurement for ordered categories whose distances are unknown, such as none, part-time, and full-time employment. (Chapter 5)
outcome
The variable a model explains or predicts. (Chapters 5, 8)
outlier
A value far from the others. (Chapter 6)
output
In Shiny, a plot, table, or text that the server produces for the user interface. (Chapter 17)
output format
The kind of document Quarto produces, such as HTML, Word, or PDF. (Chapter 17)
output layer
The final layer of a neural network, which produces the prediction. (Chapter 15)
overfitting
A model learning the noise in its training data, so it performs well there and poorly on new data. (Chapters 11, 13)
overplotting
Points drawn on top of each other in a plot, hiding how many there are. (Chapter 4)

P

package
A collection of R functions, data, and documentation that adds features to R. (Chapter 1)
paired t-test
A test comparing two measurements on the same cases. (Chapter 7)
Pandoc
The document converter that Quarto uses to produce HTML, Word, PDF, and other formats. (Chapter 17)
parallel analysis
Choosing the number of factors by comparing eigenvalues with those from random data. (Chapter 9)
parameter
A quantity describing a population (Chapter 7); a value learned by a model, such as a coefficient (Chapter 11); or an input to a Quarto report (Chapter 17). (Chapters 7, 11, 17)
patchwork
A package that combines ggplot2 plots into one figure. (Chapter 19)
penalty (lambda)
In regularised regression or a neural network, the strength of the penalty on large coefficients or weights. Also: penalty, lambda. (Chapter 13)
permutation importance
A predictor’s importance measured by how much shuffling its values reduces a model’s accuracy. (Chapter 12)
permutation test
A test that builds the null distribution by shuffling group labels many times, and compares the observed result with the shuffled ones. (Chapter 7)
pipe
The operator |>, which passes the result on its left to the function on its right. (Chapter 3)
pipeline
The series of steps, in code, from raw data to results. (Chapter 19)
Poisson model
A regression model for counts. (Chapter 10)
population
The whole group a study wants to draw conclusions about. (Chapters 5, 7)
post-hoc test
A test run after ANOVA to find which groups differ, such as Tukey’s test. (Chapter 8)
power
The probability that a study detects an effect of a given size, if the effect is real. A common target is 80%. (Chapters 5, 7)
precision
The share of cases predicted positive that are truly positive. (Chapter 12)
predicted probability
A model’s estimated probability of an outcome for a case. (Chapter 8)
prediction
Using a model to estimate the outcome for new cases. (Chapter 11)
prediction interval
A range within which a future observation is expected to fall with a stated probability. (Chapter 16)
predictive analysis
Analysis aimed at predicting new cases rather than explaining or testing. (Chapter 6)
predictor
A variable used to explain or predict the outcome. (Chapters 5, 8)
preregistration
Publicly recording hypotheses and planned analyses before seeing the data. (Chapters 5, 17)
pretrained model
A model already trained by others on large datasets, used as it is or adapted. (Chapter 15)
principal component
A combination of variables that captures as much of their variation as possible. (Chapter 9)
principal component analysis
A method that summarises many correlated variables with a few principal components. (Chapter 9)
project structure
The organisation of a project’s folders and files, such as data-raw/, R/, and output/. (Chapter 19)
prompt
The text given to a language model. (Chapter 18)
push
In git, copying commits to an online repository such as GitHub. (Chapter 17)
p-value
The probability of results at least as extreme as those observed, if the null hypothesis were true. (Chapter 7)

Q

Q-Q plot
A plot comparing a variable’s values with those expected from a normal distribution. (Chapter 6)
quadratic term
A squared predictor in a regression model, which allows a curved relationship. (Chapter 8)
qualitative coding
Assigning themes or categories to texts such as open-ended answers. (Chapter 18)
qualitative palette
A set of clearly different colours of similar strength, used to distinguish categories. (Chapter 4)
quartile
The values that divide sorted data into four equal parts. (Chapter 6)
Quarto
A publishing system that turns documents with text and code into reports, books, websites, and slides. (Chapter 17)
quasi-experiment
A study that compares groups receiving different treatments that were not assigned by chance. (Chapter 5)

R

R
The programming language and environment for statistics used in this book. (Chapter 1)
R Markdown
The predecessor of Quarto: documents (.Rmd) that combine text and R code. (Chapter 17)
radial basis function
A common kernel for support vector machines, which allows curved, rounded boundaries. (Chapter 12)
random effect
In a mixed model, variation between groups (such as students) described by a distribution rather than a separate estimate for each group. (Chapter 10)
random error
The difference between an estimate and the truth caused by which cases happened to be selected. It shrinks as the sample grows. (Chapter 5)
random forest
An ensemble of decision trees, each grown on a bootstrap sample with random subsets of predictors. (Chapter 12)
random intercept
A random effect that lets each group have its own average level. (Chapter 10)
random slope
A random effect that lets each group have its own effect of a predictor, such as its own rate of change. (Chapter 10)
randomised experiment
A study in which the researcher assigns the treatment by chance, so that the groups differ only by chance at the start. (Chapter 5)
range
The difference between the largest and smallest values. (Chapter 6)
rate ratio
In a Poisson model, the multiplicative change in the expected count per unit of a predictor. (Chapter 10)
ratio
A level of measurement with equal distances and a true zero, so that “twice as much” is meaningful, such as hours of sleep. (Chapter 5)
raw data
Data as it was collected or exported, before any cleaning; never edited by hand. (Chapter 19)
reactive expression
In Shiny, a calculation, made with reactive(), that is updated automatically when its inputs change. (Chapter 17)
reactivity
Shiny’s system for updating outputs automatically when inputs change. (Chapter 17)
recall
The share of truly positive cases that a model predicts as positive. Also: sensitivity. (Chapter 12)
recipe
In tidymodels, a list of data preparation steps, such as imputation and normalisation. (Chapter 11)
reference category
The category of a categorical predictor that the others are compared with. (Chapter 8)
registered report
A publication format in which a journal accepts a study’s plan before the data is collected. (Chapter 17)
regression
Modelling or predicting a numeric outcome. In machine learning, regression means predicting a number rather than a category. Also: Regression (prediction). (Chapters 11, 13)
regular expression
A pattern for matching text, such as "\\b(money|fee)". (Chapter 18)
regularisation
Penalising large coefficients to prevent overfitting, as in ridge and lasso regression. (Chapter 13)
relational question
A research question about which variables go together. It establishes association, not cause. (Chapter 5)
reliability
How consistently a scale measures what it measures. (Chapters 5, 9)
ReLU
Rectified linear unit: an activation function that sets negative values to zero. (Chapter 15)
remainder
In a time series decomposition, what is left after removing trend and season. (Chapter 16)
render
Running the code in a Quarto document and producing the finished output. (Chapter 17)
renv
A package that records and restores the package versions used by a project. (Chapter 17)
repeated measures
Several measurements of the same cases, such as each student in four semesters. (Chapter 10)
replicability
Whether a new study, with new data, reaches the same conclusions. (Chapter 17)
reporting sentence
A sentence in a results section that states a finding with its numbers, such as an estimate, interval, and p-value. (Chapter 19)
reporting standard
A published list of the information a research report must contain, such as JARS for psychology, CONSORT for randomised trials, and STROBE for observational studies. (Chapter 19)
repository
A project folder whose history git keeps; also its online copy on GitHub. (Chapter 17)
reproducibility
Whether the same data and code give exactly the same results. (Chapter 17)
reproducible research
Research whose results can be recomputed from the shared data and code. (Chapter 1)
research question
A focused question that says precisely what a study will find out, narrow enough that data can answer it. (Chapter 5)
research workflow
The path of a research project from question and data to reported results. (Chapter 19)
residual
The difference between an observed value and the value a model predicts. Also: error. (Chapters 8, 13)
reversed item
A questionnaire item worded in the opposite direction to the others, which must be reversed before scoring. (Chapters 3, 9)
ridge regression
Regularised regression with a penalty on the squared coefficients, which shrinks them all towards zero. (Chapter 13)
RMSE
Root mean squared error: the square root of the average squared prediction error. (Chapter 13)
ROC AUC
The area under the ROC curve: how often a model ranks a positive case above a negative one; 0.5 is chance, 1 is perfect. (Chapter 11)
ROC curve
A plot of sensitivity against the false positive rate across all classification thresholds. (Chapter 12)
rotation
In factor analysis, turning the factors to make the loadings easier to interpret. (Chapter 9)
R-squared
The share of the variation in the outcome that a model explains, from 0 to 1. Also: R². (Chapters 8, 13)
RStudio
The most widely used editor for R, used throughout this book. (Chapter 1)
RStudio Project
A folder that RStudio treats as the home of one piece of work, setting the working directory. (Chapter 1)

S

sample
The part of a population that is observed. (Chapters 5, 7)
sample description (“Table 1”)
A table describing the participants, usually the first table of a results chapter. Also: sample description. (Chapter 19)
sampling distribution
The distribution of a statistic over many possible samples. (Chapter 7)
sampling frame
A list of the members of a population, from which a sample can be drawn. (Chapter 5)
scale
In ggplot2, the control of how data values are turned into positions, colours, or sizes. (Chapter 4)
scale score
A score combining several questionnaire items, usually their mean. (Chapter 3)
scaling
Converting variables to a common scale, usually z-scores, before calculating distances. (Chapter 9)
scatter plot
A plot of two numeric variables as points. (Chapter 4)
scree plot
A plot of eigenvalues, used to choose the number of components or factors. (Chapter 9)
script
A file of R code that can be saved and rerun. (Chapter 1)
seasonal naive method
A benchmark forecast that repeats the value from the same season a cycle earlier. (Chapter 16)
seasonal period
The length of a repeating pattern, such as 52 weeks or 12 months. (Chapter 16)
seasonal plot
A plot that overlays the cycles of a time series to show its seasonal pattern. (Chapter 16)
seasonality
A pattern in a time series that repeats at a fixed period. (Chapter 16)
seasonally adjusted series
A time series with the seasonal component removed. (Chapter 16)
self-report
A measure based on what participants say about themselves, such as their estimated hours of sleep. (Chapter 5)
self-selection
When people choose their own group, such as volunteering for a workshop, so that the groups may differ from the start. (Chapter 5)
sequential palette
Colours running from light to dark, used for quantities that go in one direction. (Chapter 4)
server
In Shiny, the function that computes the outputs from the inputs. (Chapter 17)
setting
In ggplot2, giving a visual property a fixed value, such as colour = "blue", rather than mapping it to a variable. (Chapter 4)
Shiny
An R package for building interactive web applications and dashboards. (Chapter 17)
shrinkage
Pulling coefficients towards zero, as regularisation does. (Chapter 13)
sigmoid
The S-shaped function that turns any number into a value between 0 and 1. (Chapter 15)
significance level
The threshold (usually 0.05) below which a p-value counts as statistically significant. (Chapter 7)
silhouette
A measure of how much closer each case is to its own cluster than to the next nearest, from −1 to 1. (Chapters 9, 14)
simple random sampling
Sampling in which every member of the population has the same chance of being chosen. (Chapter 5)
simulation
Using random numbers to generate many possible outcomes, for example many possible futures of a time series. (Chapter 16)
singular fit
A warning that a mixed model estimates some variation as zero, often because the model is too complex for the data. (Chapter 10)
skewness
Asymmetry of a distribution; a long tail to the right is positive skew. (Chapter 6)
slope
The change in the predicted outcome for a one-unit change in a predictor. (Chapter 8)
small cells
Combinations of variables that describe very few people, who could be identified. (Chapter 17)
SMOTE
A method that balances classes by creating artificial cases of the rare class between existing ones. (Chapter 12)
specificity
The share of truly negative cases that a model predicts as negative. (Chapter 12)
stability
Whether clusters or results stay the same under small changes to the data or the method. (Chapter 14)
standard deviation
A measure of spread: roughly the typical distance of values from the mean. (Chapter 6)
standard error
The standard deviation of a sampling distribution; the uncertainty of an estimate. (Chapter 7)
statistic
A number calculated from a sample, such as a mean or a t value. (Chapter 7)
statistically significant
A result with a p-value below the significance level. (Chapter 7)
STL
Seasonal and trend decomposition using loess: a flexible method for decomposing a time series. (Chapter 16)
stratified sampling
Sampling in which the population is divided into groups (strata) and a random sample is drawn from each. (Chapter 5)
stratified split
A split of the data that keeps the same share of each outcome class in each part. (Chapter 11)
structured output
A language model’s answer in a fixed format, such as one category from a list, that code can use directly. (Chapter 18)
study design
The plan for who is measured, on what, when, and under which conditions. (Chapter 5)
supervised learning
Machine learning with a known outcome to predict. (Chapter 11)
support vector
A case close to the boundary of a support vector machine, which determines where the boundary lies. (Chapter 12)
support vector machine
A classifier that separates classes with the widest possible margin. (Chapter 12)
symmetrical
Having the same shape on both sides of the centre. (Chapter 6)
synthetic data
Artificial data with the structure of real data, used when the real data cannot be shared. (Chapter 17)
system prompt
Instructions given to a language model before the conversation, such as a task and a codebook. (Chapter 18)

T

temperature
A setting of a language model that controls how random its answers are; 0 gives the most likely answer. (Chapter 18)
test set
Data set aside to evaluate a final model once, at the end. (Chapter 11)
test-retest reliability
The extent to which a measure gives similar results for the same people on two occasions. (Chapter 5)
theme
In ggplot2, the non-data appearance of a plot, such as fonts and background. In qualitative coding, a category of meaning. (Chapter 4)
threshold
The probability above which a classifier predicts the positive class. (Chapter 12)
tibble
The tidyverse’s version of a data frame, which prints more neatly. (Chapters 2, 3)
tidy data
Data with one variable per column, one observation per row, and one value per cell. (Chapter 3)
tidyverse
A collection of R packages, such as dplyr and ggplot2, that share one way of working. (Chapter 3)
time plot
A plot of a time series against time. (Chapter 16)
time series
Observations of one quantity at regular intervals over time. (Chapter 16)
token
A word or part of a word, the unit of text a language model reads and writes. (Chapter 18)
training
Adjusting a model’s weights or parameters to fit the training data. (Chapter 15)
training cut-off
The date after which a language model’s training data contains no information. (Chapter 18)
training set
The data used to build and tune a model. (Chapter 11)
transformer
The neural network architecture behind large language models. (Chapter 15)
tree depth
The number of questions in a row a decision tree may ask. (Chapter 13)
trend
The long-term direction of a time series. (Chapter 16)
trend line
A line added to a scatter plot to show the overall relationship. (Chapter 4)
true negative
A negative case correctly predicted negative. (Chapter 12)
true positive
A positive case correctly predicted positive. (Chapter 12)
tsibble
A data frame for time series, which knows which column is time. (Chapter 16)
Tukey’s test
A post-hoc test comparing every pair of groups, with a correction for multiple comparisons. (Chapter 8)
tuning
Choosing hyperparameters by comparing their performance, usually with cross-validation. (Chapter 11)
two-sample t-test
A test comparing the means of two independent groups. (Chapter 7)
two-tailed test
A test that looks for a difference in either direction. (Chapter 7)
two-way ANOVA
ANOVA with two factors, including their interaction. (Chapter 8)
type conversion
Changing a value from one data type to another, such as text to numbers. (Chapter 2)
Type I error
Rejecting a true null hypothesis: a false alarm. (Chapter 7)
Type II error
Failing to reject a false null hypothesis: a missed effect. (Chapter 7)

U

uncertainty
In a mixture model, 1 minus a case’s largest membership probability. (Chapter 14)
underfitting
A model too simple to capture the real pattern in the data. (Chapter 11)
unimodal
Having one peak. (Chapter 6)
unit of analysis
What one row of the data represents, and what the question is about: a student, a semester, a supervisor. (Chapters 2, 5)
unsupervised learning
Machine learning without an outcome, looking for structure such as clusters. (Chapter 11)
upsampling
Balancing classes by repeating cases of the rare class in the training data. (Chapter 12)
user interface
In Shiny, what the user sees: inputs and places for outputs. (Chapter 17)

V

validity
Whether a measure measures what it claims to, or whether a study’s conclusions are justified. (Chapter 5)
value label
A label attached to a coded value, such as 1 = “Strongly disagree”, as in SPSS files. (Chapter 2)
variable
One characteristic measured on every case, such as age or faculty; one column of a data table. (Chapter 2)
variable label
A longer description attached to a variable, as in SPSS files. (Chapter 2)
variable selection
Choosing which predictors to keep in a model; the lasso does it automatically. (Chapter 13)
variance
The average squared distance from the mean; the square of the standard deviation. (Chapter 6)
vector
R’s basic structure: a sequence of values of the same type. (Chapters 1, 2)
vector format
An image format, such as PDF or SVG, that stays sharp at any size. (Chapter 4)
verb
A dplyr function that does one thing to a data frame, such as filter() or mutate(). (Chapter 3)
version control
Keeping a history of every change to a project’s files. (Chapter 17)

W

Ward’s method
A way of merging clusters in hierarchical clustering that keeps them as compact as possible. (Chapter 9)
weak learner
A simple model, such as a small tree, that is only slightly better than chance on its own; boosting combines many. (Chapter 13)
weight
In a neural network, the number by which an input is multiplied; learned during training. (Chapter 15)
weight decay
A penalty on large weights in a neural network, which prevents overfitting. (Chapter 15)
Welch test
The version of the two-sample t-test that does not assume equal variances; R’s default. (Chapter 7)
wide format
Data with repeated measurements side by side in columns, one row per case. (Chapter 3)
Wilcoxon signed-rank test
A non-parametric test comparing two measurements on the same cases. (Chapter 7)
within-group variation
How much values vary around their own group mean; the denominator of the ANOVA F statistic. (Chapter 8)
workflow
In tidymodels, a recipe and a model specification combined. (Chapter 11)
workflow set
In tidymodels, several workflows fitted and compared together. (Chapter 12)
working directory
The folder R reads files from and saves files to by default. (Chapter 1)

X

XGBoost
Extreme gradient boosting: a fast, widely used boosting method. (Chapter 13)

Y

YAML header
The settings at the top of a Quarto document, between two --- lines. (Chapter 17)

Z

z-score
A value expressed as the number of standard deviations from the mean. (Chapter 6)

R functions

Table 1.1 lists the functions used most often in the book, grouped by task, with the package they come from and the chapters that use them. base R functions are available without loading any package. For any function, ?name in the Console opens its help page.

Table 1.1: Frequently used R functions.
Task Function Package What it does Chapters
Getting started install.packages() base R Install a package from CRAN (once) 1
Getting started library() base R Load an installed package (every session) 1, 2, 3, and later
Getting started c() base R Combine values into a vector 1, 2, 3, and later
Getting started round() base R Round numbers to a number of decimal places 1, 2, 3, and later
Getting started head() base R Show the first rows of a data frame or first values of a vector 1, 2, 3, and later
Getting started str() base R Show the structure of an object 2
Getting started summary() base R Summarise a data frame or a model 2, 5, 7, and later
Getting started here() here Build a file path from the project’s folder 1, 3, 19
Importing and exporting read.csv() base R Read a CSV file 1, 2
Importing and exporting read_excel() readxl Read an Excel file 2, 3, 19
Importing and exporting read_sav() haven Read an SPSS file, with its labels 2
Importing and exporting write_csv() readr Save a data frame as a CSV file 3
Importing and exporting factor() base R Create a categorical variable with levels 2, 5, 7, and later
Cleaning and reshaping filter() dplyr Keep the rows that meet a condition 3, 4, 5, and later
Cleaning and reshaping select() dplyr Keep or drop columns 3, 4, 5, and later
Cleaning and reshaping arrange() dplyr Sort rows 3, 8, 9, and later
Cleaning and reshaping mutate() dplyr Create or change columns 3, 4, 5, and later
Cleaning and reshaping summarise() dplyr Calculate summaries, optionally by group (.by) 3, 4, 5, and later
Cleaning and reshaping count() dplyr Count the rows in each group 3, 6, 11, and later
Cleaning and reshaping case_when() dplyr Recode values with a series of conditions 3
Cleaning and reshaping left_join() dplyr Add columns from another table by matching a key 3, 4, 5, and later
Cleaning and reshaping distinct() dplyr Remove duplicate rows 3, 19
Cleaning and reshaping pivot_longer() tidyr Reshape from wide to long format 3, 4, 6, and later
Cleaning and reshaping pivot_wider() tidyr Reshape from long to wide format 3, 7, 10, and later
Cleaning and reshaping str_detect() stringr Test whether text matches a pattern 3, 19
Cleaning and reshaping parse_number() readr Extract a number from text 3, 19
Visualising ggplot() ggplot2 Start a plot from data and aesthetic mappings 4, 5, 6, and later
Visualising geom_histogram() ggplot2 Draw a histogram 4, 5, 6, and later
Visualising geom_point() ggplot2 Draw points (a scatter plot) 4, 5, 6, and later
Visualising geom_boxplot() ggplot2 Draw box plots 4, 7, 8
Visualising geom_smooth() ggplot2 Add a trend line 4, 8, 10, 19
Visualising facet_wrap() ggplot2 Split a plot into panels by group 4, 5, 6, and later
Visualising labs() ggplot2 Set titles and axis labels 4, 5, 6, and later
Visualising ggsave() ggplot2 Save a plot to a file 4, 19
Describing mean() base R Mean 1, 2, 3, and later
Describing median() base R Median 5, 6, 7
Describing sd() base R Standard deviation 4, 5, 6, and later
Describing quantile() base R Quantiles, such as quartiles 5, 6, 7, 16
Describing cor() base R Correlation coefficients 2, 4, 5, and later
Describing scale() base R Convert to z-scores 9, 14, 17
Testing set.seed() base R Make random results repeatable 5, 7, 8, and later
Testing t.test() base R One-sample, two-sample, and paired t-tests 2, 7, 17
Testing chisq.test() base R Chi-square tests 7
Testing wilcox.test() base R Mann-Whitney and Wilcoxon signed-rank tests 7, 17
Modelling aov() base R Analysis of variance 8, 19
Modelling TukeyHSD() base R Tukey’s post-hoc comparisons 8
Modelling lm() base R Linear regression 8, 10, 11, and later
Modelling glm() base R Logistic and other generalised linear models 8, 15, 19
Modelling confint() base R Confidence intervals for model coefficients 10
Modelling predict() base R Predictions from a model 8, 10, 11, and later
Modelling tidy() broom Model results as a data frame 8, 13, 19
Modelling lmer() lme4 Linear mixed-effects model 10, 19
Modelling glmer() lme4 Generalised linear mixed-effects model 10
Many variables prcomp() base R Principal component analysis 9
Many variables fa() psych Exploratory factor analysis 9
Many variables kmeans() base R k-means clustering 9, 14
Many variables hclust() base R Hierarchical clustering 9
Many variables Mclust() mclust Gaussian mixture model 14
Many variables dbscan() dbscan DBSCAN clustering 14
Machine learning initial_split() rsample (tidymodels) Split data into training and test sets 11, 12, 13, 15
Machine learning vfold_cv() rsample (tidymodels) Create cross-validation folds 11, 12, 13, 15
Machine learning recipe() recipes (tidymodels) Start a data preparation recipe 11, 12, 13, 15
Machine learning workflow() workflows (tidymodels) Combine a recipe and a model 11, 12, 13, 15
Machine learning fit() parsnip (tidymodels) Fit a model or workflow 11, 12, 13, 15
Machine learning tune_grid() tune (tidymodels) Tune hyperparameters with cross-validation 11, 12, 13, 15
Machine learning last_fit() tune (tidymodels) Fit on the training set and evaluate once on the test set 11, 12, 13, 15
Machine learning roc_auc() yardstick (tidymodels) Area under the ROC curve 11, 15
Machine learning conf_mat() yardstick (tidymodels) Confusion matrix 12
Time series as_tsibble() tsibble Create a time series data frame 16
Time series model() fabletools Fit one or more forecasting models 16
Time series forecast() fabletools Forecast from fitted models 16
Reporting and sharing kable() knitr Format a table for a report 17, 19
Reporting and sharing citation() base R How to cite R or a package 17
Reporting and sharing shinyApp() shiny Create a Shiny app from a user interface and server 17
Reporting and sharing chat_anthropic() ellmer Start a chat with a language model (Claude) 18