| Task | Function | Package | What it does | Chapters |
|---|---|---|---|---|
| Getting started | install.packages() |
base R | Install a package from CRAN (once) | 1 |
| Getting started | library() |
base R | Load an installed package (every session) | 1, 2, 3, and later |
| Getting started | c() |
base R | Combine values into a vector | 1, 2, 3, and later |
| Getting started | round() |
base R | Round numbers to a number of decimal places | 1, 2, 3, and later |
| Getting started | head() |
base R | Show the first rows of a data frame or first values of a vector | 1, 2, 3, and later |
| Getting started | str() |
base R | Show the structure of an object | 2 |
| Getting started | summary() |
base R | Summarise a data frame or a model | 2, 5, 7, and later |
| Getting started | here() |
here | Build a file path from the project’s folder | 1, 3, 19 |
| Importing and exporting | read.csv() |
base R | Read a CSV file | 1, 2 |
| Importing and exporting | read_excel() |
readxl | Read an Excel file | 2, 3, 19 |
| Importing and exporting | read_sav() |
haven | Read an SPSS file, with its labels | 2 |
| Importing and exporting | write_csv() |
readr | Save a data frame as a CSV file | 3 |
| Importing and exporting | factor() |
base R | Create a categorical variable with levels | 2, 5, 7, and later |
| Cleaning and reshaping | filter() |
dplyr | Keep the rows that meet a condition | 3, 4, 5, and later |
| Cleaning and reshaping | select() |
dplyr | Keep or drop columns | 3, 4, 5, and later |
| Cleaning and reshaping | arrange() |
dplyr | Sort rows | 3, 8, 9, and later |
| Cleaning and reshaping | mutate() |
dplyr | Create or change columns | 3, 4, 5, and later |
| Cleaning and reshaping | summarise() |
dplyr | Calculate summaries, optionally by group (.by) |
3, 4, 5, and later |
| Cleaning and reshaping | count() |
dplyr | Count the rows in each group | 3, 6, 11, and later |
| Cleaning and reshaping | case_when() |
dplyr | Recode values with a series of conditions | 3 |
| Cleaning and reshaping | left_join() |
dplyr | Add columns from another table by matching a key | 3, 4, 5, and later |
| Cleaning and reshaping | distinct() |
dplyr | Remove duplicate rows | 3, 19 |
| Cleaning and reshaping | pivot_longer() |
tidyr | Reshape from wide to long format | 3, 4, 6, and later |
| Cleaning and reshaping | pivot_wider() |
tidyr | Reshape from long to wide format | 3, 7, 10, and later |
| Cleaning and reshaping | str_detect() |
stringr | Test whether text matches a pattern | 3, 19 |
| Cleaning and reshaping | parse_number() |
readr | Extract a number from text | 3, 19 |
| Visualising | ggplot() |
ggplot2 | Start a plot from data and aesthetic mappings | 4, 5, 6, and later |
| Visualising | geom_histogram() |
ggplot2 | Draw a histogram | 4, 5, 6, and later |
| Visualising | geom_point() |
ggplot2 | Draw points (a scatter plot) | 4, 5, 6, and later |
| Visualising | geom_boxplot() |
ggplot2 | Draw box plots | 4, 7, 8 |
| Visualising | geom_smooth() |
ggplot2 | Add a trend line | 4, 8, 10, 19 |
| Visualising | facet_wrap() |
ggplot2 | Split a plot into panels by group | 4, 5, 6, and later |
| Visualising | labs() |
ggplot2 | Set titles and axis labels | 4, 5, 6, and later |
| Visualising | ggsave() |
ggplot2 | Save a plot to a file | 4, 19 |
| Describing | mean() |
base R | Mean | 1, 2, 3, and later |
| Describing | median() |
base R | Median | 5, 6, 7 |
| Describing | sd() |
base R | Standard deviation | 4, 5, 6, and later |
| Describing | quantile() |
base R | Quantiles, such as quartiles | 5, 6, 7, 16 |
| Describing | cor() |
base R | Correlation coefficients | 2, 4, 5, and later |
| Describing | scale() |
base R | Convert to z-scores | 9, 14, 17 |
| Testing | set.seed() |
base R | Make random results repeatable | 5, 7, 8, and later |
| Testing | t.test() |
base R | One-sample, two-sample, and paired t-tests | 2, 7, 17 |
| Testing | chisq.test() |
base R | Chi-square tests | 7 |
| Testing | wilcox.test() |
base R | Mann-Whitney and Wilcoxon signed-rank tests | 7, 17 |
| Modelling | aov() |
base R | Analysis of variance | 8, 19 |
| Modelling | TukeyHSD() |
base R | Tukey’s post-hoc comparisons | 8 |
| Modelling | lm() |
base R | Linear regression | 8, 10, 11, and later |
| Modelling | glm() |
base R | Logistic and other generalised linear models | 8, 15, 19 |
| Modelling | confint() |
base R | Confidence intervals for model coefficients | 10 |
| Modelling | predict() |
base R | Predictions from a model | 8, 10, 11, and later |
| Modelling | tidy() |
broom | Model results as a data frame | 8, 13, 19 |
| Modelling | lmer() |
lme4 | Linear mixed-effects model | 10, 19 |
| Modelling | glmer() |
lme4 | Generalised linear mixed-effects model | 10 |
| Many variables | prcomp() |
base R | Principal component analysis | 9 |
| Many variables | fa() |
psych | Exploratory factor analysis | 9 |
| Many variables | kmeans() |
base R | k-means clustering | 9, 14 |
| Many variables | hclust() |
base R | Hierarchical clustering | 9 |
| Many variables | Mclust() |
mclust | Gaussian mixture model | 14 |
| Many variables | dbscan() |
dbscan | DBSCAN clustering | 14 |
| Machine learning | initial_split() |
rsample (tidymodels) | Split data into training and test sets | 11, 12, 13, 15 |
| Machine learning | vfold_cv() |
rsample (tidymodels) | Create cross-validation folds | 11, 12, 13, 15 |
| Machine learning | recipe() |
recipes (tidymodels) | Start a data preparation recipe | 11, 12, 13, 15 |
| Machine learning | workflow() |
workflows (tidymodels) | Combine a recipe and a model | 11, 12, 13, 15 |
| Machine learning | fit() |
parsnip (tidymodels) | Fit a model or workflow | 11, 12, 13, 15 |
| Machine learning | tune_grid() |
tune (tidymodels) | Tune hyperparameters with cross-validation | 11, 12, 13, 15 |
| Machine learning | last_fit() |
tune (tidymodels) | Fit on the training set and evaluate once on the test set | 11, 12, 13, 15 |
| Machine learning | roc_auc() |
yardstick (tidymodels) | Area under the ROC curve | 11, 15 |
| Machine learning | conf_mat() |
yardstick (tidymodels) | Confusion matrix | 12 |
| Time series | as_tsibble() |
tsibble | Create a time series data frame | 16 |
| Time series | model() |
fabletools | Fit one or more forecasting models | 16 |
| Time series | forecast() |
fabletools | Forecast from fitted models | 16 |
| Reporting and sharing | kable() |
knitr | Format a table for a report | 17, 19 |
| Reporting and sharing | citation() |
base R | How to cite R or a package | 17 |
| Reporting and sharing | shinyApp() |
shiny | Create a Shiny app from a user interface and server | 17 |
| Reporting and sharing | chat_anthropic() |
ellmer | Start a chat with a language model (Claude) | 18 |
Appendix B: Glossary
This appendix collects the key terms from every chapter’s review, with a short definition of each and the chapters where it is introduced or used, followed by a table of the R functions used most often in the book. Definitions are written in plain language; the chapters give the full explanations and examples.
Key terms
A
- accuracy
- The share of cases a classification model classifies correctly. Misleading when one outcome is rare. (Chapter 11)
- activation function
- The function a neuron applies to its weighted total to produce its output, such as the sigmoid or the ReLU. (Chapter 15)
- adjusted Rand index
- A measure of agreement between two clusterings or classifications, corrected for chance: 1 means identical groupings, 0 means chance agreement. (Chapter 14)
- adjusted R-squared
- R-squared corrected for the number of predictors, so that adding useless predictors does not make a model look better. (Chapter 8)
- AI coding assistant
- An AI tool that writes, explains, or completes code, in a chat or inside the editor. Its suggestions must be checked. (Chapter 18)
- AIC
- Akaike information criterion: a measure for comparing models that balances fit against complexity; lower is better. (Chapter 10)
- alternative hypothesis
- The claim that there is an effect or a difference, tested against the null hypothesis. (Chapters 5, 7)
- annotation
- Text, arrows, or shapes added to a plot to point out something specific. (Chapter 4)
- anonymisation
- Removing or altering information so that the people in a dataset cannot be identified. (Chapter 17)
- ANOVA
- Analysis of variance: a test of whether the means of three or more groups differ, by comparing variation between groups with variation within them. (Chapter 8)
- API
- Application programming interface: a way for one program to send requests to another, such as R sending text to a language model. (Chapter 18)
- API key
-
A secret code that identifies you to an online service’s API. Keep it in
.Renviron, never in a script. (Chapter 18) - argument
-
A value given to a function inside its brackets, such as
na.rm = TRUE, that controls what the function does. (Chapter 1) - ARIMA
- A family of time series models that forecast from the autocorrelation of the series and past random shocks. (Chapter 16)
- artificial intelligence
- Computer systems that perform tasks usually needing human intelligence, such as understanding language; in this book, mainly large language models. (Chapter 18)
- aspect ratio
- The ratio of a graph’s width to its height. It changes how steep lines and slopes appear. (Chapter 4)
- assignment
-
Storing a value in an object with
<-, as inx <- 5. (Chapter 1) - attrition
- The loss of participants from a study over time. (Chapters 5, 6)
- autocorrelation
- The correlation between a time series and itself a number of steps (lags) earlier. (Chapter 16)
B
- bar chart
- A plot of counts or values as bars, one per category. (Chapter 4)
- baseline model
- The simplest possible model, such as predicting the mean for everyone, used as a benchmark for real models. (Chapter 13)
- Bayesian information criterion (BIC)
- A criterion for comparing models that rewards fit and penalises complexity. In mclust, higher BIC is better. Also: BIC. (Chapter 14)
- benchmark forecast
- A simple forecast, such as the mean or the value a season earlier, that any serious method should beat. (Chapter 16)
- between-group variation
- How much group means differ from the overall mean; the numerator of the ANOVA F statistic. (Chapter 8)
- bias
- In a neural network, the constant added to a neuron’s weighted total (like an intercept). More generally, a systematic error. (Chapters 5, 15)
- bias-variance trade-off
- The balance between a model too simple to capture the pattern (high bias) and one so flexible that it follows the noise of each sample (high variance). Error on new data is lowest in between. (Chapter 11)
- BibTeX
- A text format for bibliographic references, used by Quarto to format citations and reference lists. (Chapter 17)
- bimodal
- Having two peaks. (Chapter 6)
- bin
- One of the intervals into which a histogram divides the values. (Chapter 4)
- boosting
- Building a model by adding many small models (usually trees) one at a time, each correcting the errors of the model so far. (Chapter 13)
- bootstrap
- Estimating uncertainty by repeatedly resampling the data with replacement and recalculating a statistic. (Chapter 7)
- bootstrap sample
- A sample of the same size as the data, drawn from it at random with replacement. (Chapter 12)
- border point
- In DBSCAN, a case near a core point that belongs to its cluster but is not itself a core point. (Chapter 14)
- box plot
- A plot that shows the median, the quartiles, and unusual values of a variable. (Chapter 4)
- broom
-
A package that turns model results into tidy data frames, with
tidy(),glance(), andaugment(). (Chapter 19)
C
- case
- One unit that is measured in a study, such as a student, a patient, or a school; one row of a data table. (Chapter 2)
- causal question
- A research question about what leads to what. It needs a design that rules out other explanations, ideally an experiment. (Chapter 5)
- central limit theorem
- The result that the sampling distribution of a mean is approximately normal for large enough samples, whatever the shape of the data. (Chapter 7)
- centring
- Subtracting the mean from a variable, so that zero means “average”; often used before fitting interactions. (Chapter 8)
- chain of reasoning
- The sequence of steps behind a research conclusion: question, hypothesis, design, variables, analysis, result, conclusion, and limitation. (Chapter 19)
- chi-square test
- A test for categorical data: whether counts fit expected proportions (goodness of fit), or whether two categorical variables are related (independence). (Chapter 7)
- chunk option
-
A setting for a code chunk in a Quarto document, written as
#| option: value, such asecho: false. (Chapter 17) - citation
-
A reference to a source in the text, written in Quarto as
[@key]. (Chapter 17) - citation style (CSL)
- A file in the Citation Style Language that sets how citations and references are formatted, such as APA. (Chapter 17)
- class weights
- Making errors on a rare class count more when a model is fitted, one remedy for imbalanced outcomes. (Chapter 12)
- classification
- Predicting a category, such as whether a student will consider dropping out. (Chapter 11)
- cluster analysis
- Methods that group cases so that cases in the same group are similar to each other. (Chapter 9)
- cluster sampling
- Sampling whole groups, such as all the students of randomly chosen supervisors. (Chapter 5)
- code chunk
- A block of code in a Quarto document that is run when the document is rendered. (Chapter 17)
- codebook
- A description of every variable in a dataset or, in qualitative coding, of every theme and how to assign it. (Chapters 2, 17, 18)
- coefficient path
- A plot of how each coefficient of a regularised model changes as the penalty changes. (Chapter 13)
- Cohen’s d
- An effect size for the difference between two means, in standard deviations. (Chapter 7)
- Cohen’s kappa
- A measure of agreement between two coders or classifiers, corrected for the agreement expected by chance. (Chapter 18)
- comment
-
Text in a script after
#, which R ignores; used to explain the code. (Chapter 1) - commit
- In git, a saved snapshot of a project, with a message describing the change. (Chapter 17)
- confidence interval
- A range of plausible values for a population quantity, calculated from a sample, with a stated level of confidence such as 95%. (Chapter 7)
- confirmatory analysis
- An analysis planned before seeing the data, to test a specific hypothesis. (Chapters 5, 17)
- confirmatory factor analysis
- Factor analysis that tests a factor structure specified in advance. (Chapter 9)
- confounder
- A variable related to both the predictor and the outcome, which can create or hide an apparent effect. (Chapters 5, 6, 8)
- confusion matrix
- A table of predicted against actual classes, showing true and false positives and negatives. (Chapter 12)
- Console
- The RStudio pane where R commands are run and results appear. (Chapter 1)
- construct
- An idea that cannot be observed directly, such as stress or wellbeing, and must be measured indirectly. (Chapters 5, 9)
- construct validity
- Whether a score behaves as the construct should: related to similar measures, less related to different ones. (Chapter 5)
- content validity
- Whether the items of a measure cover the whole construct. (Chapter 5)
- convenience sampling
- Sampling whoever is easy to reach. Common, and the weakest basis for generalising. (Chapter 5)
- convolutional network
- A type of deep neural network designed for images. (Chapter 15)
- core point
-
In DBSCAN, a case with at least
minPtscases within distanceeps. (Chapter 14) - correlation coefficient
- A number from −1 to 1 measuring the strength and direction of a straight-line relationship between two variables. (Chapter 6)
- correlation matrix
- A table of the correlations between every pair of variables. (Chapter 9)
- count outcome
- An outcome that counts events (0, 1, 2, …), usually modelled with a Poisson model. (Chapter 10)
- covariance matrix
- A table of the variances and covariances of several variables; in a mixture model it sets the shape of a cluster. (Chapter 14)
- Cronbach’s alpha
- A measure of the internal consistency (reliability) of a scale made of several items. (Chapter 9)
- cross-loading
- An item that loads substantially on more than one factor. (Chapter 9)
- cross-reference
-
A reference in a Quarto document, such as
@fig-wellbeing, that becomes a numbered link. (Chapter 17) - cross-sectional design
- A design that measures every variable once, at one time. (Chapter 5)
- cross-validation
- Estimating how well a model predicts new data by repeatedly fitting it on part of the training data and testing it on the rest. (Chapter 11)
- CSV file
- A plain-text file of comma-separated values, the most common format for sharing data tables. (Chapter 2)
D
- data cleaning
- Preparing raw data for analysis: removing test and duplicate cases, making categories consistent, marking missing values, and correcting or removing impossible values, with every decision recorded. (Chapter 3)
- data frame
- R’s table of data: columns are variables, rows are cases. (Chapters 1, 2)
- data leakage
- Information from the test data reaching the model during training, which makes its performance look better than it is. (Chapter 11)
- data matrix
- The arrangement of data as a table with one row per case and one column per variable. (Chapter 2)
- data structure
- The way data is organised in R: vector, factor, data frame, matrix, or list. (Chapter 2)
- data type
- The kind of value: numeric, integer, character, or logical. (Chapter 1)
- DBSCAN
- A clustering method that finds dense regions of cases and labels cases in sparse regions as noise. (Chapter 14)
- decision tree
- A model that predicts by a series of yes-or-no questions about the predictors. (Chapters 11, 12)
- decomposition
- Splitting a time series into trend, seasonal, and remainder components. (Chapter 16)
- deep learning
- Neural networks with many hidden layers, used for images, sound, and text. (Chapter 15)
- dendrogram
- The tree diagram produced by hierarchical clustering. (Chapter 9)
- density
- In clustering, how closely packed cases are in a region. (Chapter 14)
- density plot
- A smooth version of a histogram, showing the shape of a distribution. (Chapter 4)
- descriptive question
- A research question about what is: how much, how many, how often. (Chapter 5)
- descriptive statistics
- Numbers that summarise data, such as means, medians, and standard deviations. (Chapter 6)
- design effect
- How many times more grouped observations are needed to give the same information as independent ones: 1 + (m - 1) × ICC, where m is the group size. (Chapter 10)
- deviation
- The distance of a value from the mean, \(x_i - \bar{x}\). The standard deviation summarises the deviations of all values. (Chapter 6)
- diagnostic plot
- A plot used to check a model’s assumptions, such as residuals against fitted values. (Chapter 8)
- dictionary method
- Classifying texts by counting keywords from a list for each category. (Chapter 18)
- diminishing returns
- A relationship that levels off, so each extra unit of the predictor adds less than the one before. (Chapter 8)
- directional hypothesis
- A hypothesis that predicts the direction of an effect, such as “higher” or “lower”. (Chapter 5)
- disclosure
- Stating in a publication how AI tools or other aids were used. (Chapter 18)
- distance
- How different two cases are, calculated from their values on several variables. (Chapters 9, 12)
- diverging palette
- Colours running in two directions away from a meaningful midpoint, such as zero, used for values that can be positive or negative. (Chapter 4)
- DOI
-
Digital object identifier: a permanent identifier for a publication or dataset, such as
10.1126/science.1213847. (Chapter 17) - downsampling
- Balancing classes by randomly removing cases of the common class from the training data. (Chapter 12)
- dpi
- Dots per inch: the resolution of a saved image; 300 dpi is usual for print. (Chapter 4)
- dummy variable
- A 0/1 variable representing one category of a categorical predictor. (Chapter 11)
E
- effect size
- A measure of how large an effect is, independent of sample size, such as Cohen’s d or eta squared. (Chapter 7)
- eigenvalue
- In PCA, the amount of variance captured by a component. (Chapter 9)
- elastic net
- Regularised regression that mixes the ridge and lasso penalties. (Chapter 13)
- elbow method
- Choosing the number of clusters where adding more stops reducing within-cluster distance much. (Chapter 9)
- ensemble
- A model that combines the predictions of many models, such as a random forest. (Chapter 12)
- epoch
- One pass through the training data when training a neural network. (Chapter 15)
- eps
- In DBSCAN, the radius of the neighbourhood around each case. (Chapter 14)
- error message
- R’s message when something goes wrong, which usually says what and where. (Chapter 2)
- eta squared
- An effect size for ANOVA: the share of the variation in the outcome explained by the groups. (Chapter 8)
- expected count
- In a chi-square test, the count expected in a cell if the null hypothesis were true. (Chapter 7)
- explanation
- Using a model to understand why something happens, rather than to predict new cases. (Chapter 11)
- explanatory graph
- A carefully designed graph made for readers, to show a finding clearly and accurately. (Chapter 4)
- exploratory analysis
- Exploring data to find patterns and generate questions, rather than to test planned hypotheses. (Chapters 5, 6, 17)
- exploratory graph
- A quick graph made by the researcher to understand the data: to check distributions, find unusual values, and notice patterns. (Chapter 4)
- exponential smoothing (ETS)
- Forecasting with weighted averages of past observations, with more weight on recent ones; ETS models the error, trend, and season. Also: ETS. (Chapter 16)
- external validation
- Checking clusters or predictions against known categories or outcomes. (Chapter 14)
- external validity
- Whether a study’s results apply beyond it, to other people, places, and times. (Chapter 5)
F
- F statistic
- In ANOVA, the ratio of between-group to within-group variation. (Chapter 8)
- F1 score
- A single measure combining precision and recall (their harmonic mean). (Chapter 12)
- facet
-
A small panel of a plot showing one subgroup;
facet_wrap()makes one panel per group. (Chapter 4) - factor
- In R, a categorical variable with a fixed set of levels. In factor analysis, a hidden (latent) variable behind several items. (Chapters 2, 9)
- factor analysis
- A method that explains the correlations among items by a smaller number of hidden factors. (Chapter 9)
- false negative
- A case that belongs to the positive class but is predicted negative. (Chapter 12)
- false positive
- A case predicted positive that belongs to the negative class. (Chapter 12)
- falsifiability
- The property of a claim that some possible result would contradict it. A hypothesis that fits every possible result tells us nothing. (Chapter 5)
- fixed effect
- In a mixed model, an effect assumed the same for everyone, such as the average change over time. (Chapter 10)
- fold
- One of the parts into which the data is split for cross-validation. (Chapter 11)
- forecast horizon
- How far ahead a forecast is made. (Chapter 16)
- Fourier terms
- Pairs of sine and cosine waves used as predictors to describe a seasonal pattern. (Chapter 16)
- function
-
A named piece of code that takes arguments and returns a result, such as
mean(). (Chapter 1)
G
- garden of forking paths
- The many reasonable choices in an analysis. Choosing among them after seeing the results inflates the chance of a false positive. (Chapter 17)
- Gaussian mixture model
- A model that describes data as a mix of normal distributions, giving each case a probability of belonging to each cluster. (Chapter 14)
- generalisation
- The ability of a model to perform well on new data, not only on the data it learned from. (Chapter 11)
- generalised linear mixed model
- A mixed-effects model for outcomes that are not normal, such as counts or yes/no outcomes. (Chapter 10)
- geom
- In ggplot2, the geometric shape that represents data, such as points, lines, or bars. (Chapter 4)
- Gini impurity
- A measure of how mixed the classes are in a group; decision trees choose splits that reduce it. (Chapter 12)
- git
- The standard version control system, which records the history of a project. (Chapter 17)
- GitHub
- A website for storing git projects online, sharing them, and working on them together. (Chapter 17)
- gradient
- The direction and rate at which the error changes as each weight changes. (Chapter 15)
- gradient descent
- Training a model by repeatedly moving the weights a small step in the direction that reduces the error. (Chapter 15)
- grammar of graphics
- The idea behind ggplot2: a plot is built from data, aesthetic mappings, and geometric layers. (Chapter 4)
- graphical perception
- How people read values from graphs. Positions and lengths are judged most accurately, then angles, areas, and colours. (Chapter 4)
- grouped summary
-
Summary statistics calculated separately for each group, for example with
summarise(.by = ...). (Chapter 3)
H
- hallucination
- Plausible but invented content produced by a language model, such as a function or reference that does not exist. (Chapter 18)
- heat map
- A grid of coloured cells showing values, such as a correlation matrix. (Chapter 4)
- helper function
- A small function written to avoid repeating the same code, such as one that formats means. (Chapter 19)
- hidden layer
- A layer of neurons between the inputs and the output of a neural network. (Chapter 15)
- hierarchical clustering
- Clustering that repeatedly merges the most similar groups, producing a dendrogram. (Chapter 9)
- histogram
- A plot of the distribution of a numeric variable, as bars counting the values in each bin. (Chapter 4)
- hyperparameter
- A setting of a model that is chosen before fitting rather than learned from the data, such as the number of neighbours in k-NN. (Chapter 11)
- hypothesis
- A prediction, stated before the data is analysed, of what the data will show; precise enough to be wrong. (Chapter 5)
I
- identifier
-
A variable that uniquely identifies each case, such as
student_id. (Chapter 3) - imbalanced outcome
- An outcome in which one class is much rarer than the other. (Chapter 11)
- imputation
- Filling in missing values with estimated ones, such as the median. (Chapter 11)
- independence
- The assumption that observations do not influence each other; in a chi-square test, the hypothesis that two variables are unrelated. (Chapter 7)
- index
- In a tsibble, the column that holds time. (Chapter 16)
- inferential analysis
- Using a sample to draw conclusions about a population. (Chapter 6)
- inline code
- R code inside a sentence of a Quarto document, replaced by its result when rendered. (Chapter 17)
- input
- In Shiny, a control such as a menu or slider whose value the user chooses. (Chapter 17)
- input layer
- The inputs (predictors) of a neural network. (Chapter 15)
- interaction
- When the effect of one predictor depends on the value of another. (Chapter 8)
- interaction plot
- A plot of group means that shows whether the effect of one factor depends on another. (Chapter 8)
- intercept
- The predicted value of the outcome when all predictors are zero. (Chapter 8)
- internal consistency
- The extent to which the items of a scale agree with each other, often measured with Cronbach’s alpha. (Chapters 5, 9)
- internal validation
- Judging clusters by how compact and separated they are, using the data alone, as with the silhouette. (Chapter 14)
- internal validity
- Whether a study can rule out other explanations for its results. Highest in randomised experiments. (Chapter 5)
- interquartile range
- The range of the middle 50% of the values: the third quartile minus the first. (Chapter 6)
- inter-rater reliability
- The extent to which two people coding or rating the same material agree, often measured with Cohen’s kappa. (Chapters 5, 18)
- interval
- A level of measurement with equal distances between values but no true zero, such as a wellbeing index. (Chapter 5)
- intraclass correlation
- The share of the total variation that lies between groups (such as students or supervisors). (Chapter 10)
J
- join
-
Combining two tables by matching rows on a key, such as
student_id. (Chapter 3)
K
- Kaiser rule
- Keeping components with an eigenvalue above 1; a rough guide that often keeps too many. (Chapter 9)
- kernel
- In a support vector machine, a function that allows curved boundaries between classes. (Chapter 12)
- k-means
- A clustering method that assigns each case to the nearest of k cluster centres and moves the centres to the means. (Chapter 9)
- k-nearest neighbours
- Predicting a case from the k most similar cases in the training data. (Chapters 11, 12)
- k-nearest-neighbour distance plot
-
A sorted plot of each case’s distance to its k-th nearest neighbour, used to choose
epsfor DBSCAN. (Chapter 14) - knitr
- The R package that runs the code in Quarto and R Markdown documents. (Chapter 17)
- Kruskal-Wallis test
- A non-parametric alternative to one-way ANOVA. (Chapter 8)
L
- lag
- The number of time steps between an observation and an earlier one it is compared with. (Chapter 16)
- large language model
- A neural network trained on huge amounts of text to predict the next token, used in AI assistants. (Chapter 18)
- lasso
- Regularised regression with a penalty on the absolute size of the coefficients, which sets some of them to exactly zero. (Chapter 13)
- latent variable
- A variable that cannot be observed directly, such as stress, and is assumed to cause part of the answers to the items that measure it. (Chapter 9)
- LaTeX
-
A typesetting system whose notation is used to write equations, as in
$\bar{x}$. (Chapter 17) - layer
-
In ggplot2, one part of a plot added with
+, such as a set of points or a trend line. In a neural network, a group of neurons. (Chapter 4) - leaf
- A final node of a decision tree, which gives the prediction. (Chapter 12)
- learning rate
- In boosting and neural networks, how large a step each update takes. (Chapters 13, 15)
- least squares
- Choosing a regression line that minimises the sum of the squared residuals. (Chapter 8)
- level
- One of the categories of a factor. (Chapter 2)
- level of measurement
- What the values of a variable mean (nominal, ordinal, interval, or ratio), which decides the summaries and tests that make sense. (Chapter 5)
- Levene’s test
- A test of whether groups have equal variances. (Chapter 8)
- licence
- A statement of what others may do with shared data or code, such as CC BY or MIT. (Chapter 17)
- likelihood ratio test
- A test comparing two nested models by how much better the larger one fits. (Chapter 10)
- line chart
- A plot of values connected by lines, often over time. (Chapter 4)
- linear regression
- A model of a numeric outcome as a straight-line function of one or more predictors. (Chapter 8)
- list
- An R object that can hold elements of different types and sizes, such as the results of a test. (Chapter 2)
- loading
- The correlation between an item and a factor or component. (Chapter 9)
- local model
- A language model that runs on your own computer, so data does not leave it. (Chapter 18)
- lockfile
-
A file, such as renv’s
renv.lock, that records the exact package versions of a project. (Chapter 17) - log loss
- The measure of error used to fit logistic regression and classification networks: small when high probabilities are given to the cases that did occur and low probabilities to those that did not. (Chapter 15)
- logical indexing
-
Selecting elements with a TRUE/FALSE condition, as in
x[x > 5]. (Chapter 2) - logistic regression
- A regression model for a yes/no outcome, which predicts the probability of “yes”. (Chapter 8)
- log-odds
- The logarithm of the odds; the scale on which logistic regression coefficients are estimated. (Chapter 8)
- long format
- Data with one row per observation, such as one row per student per semester. (Chapter 3)
- longitudinal design
- A design that measures the same people repeatedly over time. (Chapter 5)
M
- machine learning
- Methods that learn patterns from data to make predictions or find structure, judged by performance on new data. (Chapter 11)
- MAE
- Mean absolute error: the average size of the prediction errors. (Chapter 13)
- main effect
- The effect of one factor, averaged over the levels of another. (Chapter 8)
- Mann-Whitney U test
- A non-parametric test comparing two independent groups. (Chapter 7)
- mapping
- In ggplot2, linking a variable to a visual property (an aesthetic), such as position or colour. Also: aesthetic mapping. (Chapter 4)
- margin
- In a support vector machine, the empty band between the classes and the boundary. (Chapter 12)
- Markdown
-
A simple way of formatting text with symbols, such as
**bold**and# Heading. (Chapter 17) - matrix
- A two-dimensional table of values of one type. (Chapter 2)
- mean
- The average: the sum of the values divided by their number. (Chapter 6)
- mean method
- A benchmark forecast that predicts the average of past observations. (Chapter 16)
- median
- The middle value when the values are sorted. (Chapters 4, 6)
- mediator
- A variable on the path between a predictor and an outcome, through which the predictor has its effect. (Chapter 5)
- membership probability
- The probability that a case belongs to each cluster, as given by a mixture model. Also: soft assignment. (Chapter 14)
- method choice
- Choosing an analysis method from the goal, the type of outcome, and how observations are related. (Chapter 19)
- minPts
- In DBSCAN, the number of cases a neighbourhood must contain for a case to be a core point. (Chapter 14)
- misclassification cost
- The cost assigned to each kind of error of a classifier, such as missing a case or raising a false alarm. The costs determine the best threshold. (Chapter 12)
- missing at random
- Missingness that depends only on observed variables, not on the missing values themselves. (Chapter 6)
- missing code
-
A value such as
99or-9used in raw data to mark a missing answer. (Chapter 3) - missing completely at random
- Missingness unrelated to any variable, observed or not. (Chapter 6)
- missing not at random
- Missingness that depends on the missing values themselves. (Chapter 6)
- missing value (
NA) - R’s marker for a value that is not available. Also: NA. (Chapter 1)
- mixed-effects model
- A regression model with both fixed effects and random effects, for repeated or nested data. (Chapter 10)
- mixing probability
- In a mixture model, the share of cases belonging to a component. (Chapter 14)
- mixture
- In regularised regression, the mix of lasso and ridge penalties, from 0 (ridge) to 1 (lasso). (Chapter 13)
- mixture component
- One of the normal distributions in a Gaussian mixture model. (Chapter 14)
- mode
- The most common value. (Chapter 6)
- model specification
- In tidymodels, the description of a model (type, engine, mode) before it is fitted. (Chapter 11)
- moderator
- A variable that changes the strength or direction of the relationship between a predictor and an outcome. (Chapter 5)
mtry- In a random forest, the number of predictors each split may choose from. (Chapter 12)
- multicollinearity
- Strong correlation among predictors, which makes regression coefficients unstable. (Chapter 13)
- multilayer perceptron
- A neural network with one or more hidden layers of neurons. (Chapter 15)
- multilevel model
- Another name for a mixed-effects model, emphasising nested levels. (Chapter 10)
- multiple regression
- Linear regression with more than one predictor. (Chapter 8)
- multiple testing
- Running many tests in one study. Each test carries its own risk of a false alarm, so the chance of at least one grows with the number of tests. (Chapter 7)
- multivariate analysis
- Methods that analyse many variables at once, such as PCA and factor analysis. (Chapter 9)
N
- naive method
- A benchmark forecast that repeats the last observed value. (Chapter 16)
- nested data
- Data in which cases are grouped within units, such as students within supervisors. (Chapter 10)
- neural network
- A model made of connected neurons in layers, whose weights are learned from data. (Chapter 15)
- neuron (unit)
- The basic element of a neural network: it weights its inputs, adds a bias, and applies an activation function. Also: neuron, unit. (Chapter 15)
- node
- A point in a decision tree where a question is asked or a prediction is made. (Chapter 12)
- noise
- In DBSCAN, cases in sparse regions that belong to no cluster. (Chapter 14)
- nominal
- A level of measurement for categories with no order, such as faculty. (Chapter 5)
- non-directional hypothesis
- A hypothesis that predicts a difference or relationship without saying in which direction. (Chapter 5)
- non-parametric test
- A test that does not assume a particular distribution, often based on ranks. (Chapter 7)
- normal distribution
- The symmetric, bell-shaped distribution described by a mean and a standard deviation. (Chapter 6)
- normalisation
- Rescaling variables, usually to z-scores, so that they are on the same scale. (Chapter 11)
- null distribution
- The distribution of a statistic that would be expected if the null hypothesis were true, used to judge how surprising the observed result is. (Chapter 7)
- null hypothesis
- The claim of no effect or no difference, which a test tries to reject. (Chapters 5, 7)
O
- object
- A named value stored in R, such as a number, a vector, or a data frame. (Chapter 1)
- oblique rotation
- A factor rotation that allows the factors to correlate. (Chapter 9)
- odds
- The probability of an event divided by the probability of it not happening. (Chapter 8)
- odds ratio
- The ratio of the odds in two groups; in logistic regression, the multiplicative change in odds per unit of a predictor. (Chapter 8)
- one-sample t-test
- A test of whether a mean differs from a given value. (Chapter 7)
- one-tailed test
- A test that looks for a difference in one direction only. (Chapter 7)
- open science
- Practices that make research transparent and reusable: sharing data and code, preregistration, and open access. (Chapter 17)
- operationalisation
- The decision about how a construct is measured: which questions, records, or instruments turn it into a variable. (Chapter 5)
- ordered factor
-
A factor whose levels have an order, created with
factor(..., ordered = TRUE). (Chapter 5) - ordinal
- A level of measurement for ordered categories whose distances are unknown, such as none, part-time, and full-time employment. (Chapter 5)
- outcome
- The variable a model explains or predicts. (Chapters 5, 8)
- outlier
- A value far from the others. (Chapter 6)
- output
- In Shiny, a plot, table, or text that the server produces for the user interface. (Chapter 17)
- output format
- The kind of document Quarto produces, such as HTML, Word, or PDF. (Chapter 17)
- output layer
- The final layer of a neural network, which produces the prediction. (Chapter 15)
- overfitting
- A model learning the noise in its training data, so it performs well there and poorly on new data. (Chapters 11, 13)
- overplotting
- Points drawn on top of each other in a plot, hiding how many there are. (Chapter 4)
P
- package
- A collection of R functions, data, and documentation that adds features to R. (Chapter 1)
- paired t-test
- A test comparing two measurements on the same cases. (Chapter 7)
- Pandoc
- The document converter that Quarto uses to produce HTML, Word, PDF, and other formats. (Chapter 17)
- parallel analysis
- Choosing the number of factors by comparing eigenvalues with those from random data. (Chapter 9)
- parameter
- A quantity describing a population (Chapter 7); a value learned by a model, such as a coefficient (Chapter 11); or an input to a Quarto report (Chapter 17). (Chapters 7, 11, 17)
- patchwork
- A package that combines ggplot2 plots into one figure. (Chapter 19)
- penalty (lambda)
- In regularised regression or a neural network, the strength of the penalty on large coefficients or weights. Also: penalty, lambda. (Chapter 13)
- permutation importance
- A predictor’s importance measured by how much shuffling its values reduces a model’s accuracy. (Chapter 12)
- permutation test
- A test that builds the null distribution by shuffling group labels many times, and compares the observed result with the shuffled ones. (Chapter 7)
- pipe
-
The operator
|>, which passes the result on its left to the function on its right. (Chapter 3) - pipeline
- The series of steps, in code, from raw data to results. (Chapter 19)
- Poisson model
- A regression model for counts. (Chapter 10)
- population
- The whole group a study wants to draw conclusions about. (Chapters 5, 7)
- post-hoc test
- A test run after ANOVA to find which groups differ, such as Tukey’s test. (Chapter 8)
- power
- The probability that a study detects an effect of a given size, if the effect is real. A common target is 80%. (Chapters 5, 7)
- precision
- The share of cases predicted positive that are truly positive. (Chapter 12)
- predicted probability
- A model’s estimated probability of an outcome for a case. (Chapter 8)
- prediction
- Using a model to estimate the outcome for new cases. (Chapter 11)
- prediction interval
- A range within which a future observation is expected to fall with a stated probability. (Chapter 16)
- predictive analysis
- Analysis aimed at predicting new cases rather than explaining or testing. (Chapter 6)
- predictor
- A variable used to explain or predict the outcome. (Chapters 5, 8)
- preregistration
- Publicly recording hypotheses and planned analyses before seeing the data. (Chapters 5, 17)
- pretrained model
- A model already trained by others on large datasets, used as it is or adapted. (Chapter 15)
- principal component
- A combination of variables that captures as much of their variation as possible. (Chapter 9)
- principal component analysis
- A method that summarises many correlated variables with a few principal components. (Chapter 9)
- project structure
-
The organisation of a project’s folders and files, such as
data-raw/,R/, andoutput/. (Chapter 19) - prompt
- The text given to a language model. (Chapter 18)
- push
- In git, copying commits to an online repository such as GitHub. (Chapter 17)
- p-value
- The probability of results at least as extreme as those observed, if the null hypothesis were true. (Chapter 7)
Q
- Q-Q plot
- A plot comparing a variable’s values with those expected from a normal distribution. (Chapter 6)
- quadratic term
- A squared predictor in a regression model, which allows a curved relationship. (Chapter 8)
- qualitative coding
- Assigning themes or categories to texts such as open-ended answers. (Chapter 18)
- qualitative palette
- A set of clearly different colours of similar strength, used to distinguish categories. (Chapter 4)
- quartile
- The values that divide sorted data into four equal parts. (Chapter 6)
- Quarto
- A publishing system that turns documents with text and code into reports, books, websites, and slides. (Chapter 17)
- quasi-experiment
- A study that compares groups receiving different treatments that were not assigned by chance. (Chapter 5)
R
- R
- The programming language and environment for statistics used in this book. (Chapter 1)
- R Markdown
-
The predecessor of Quarto: documents (
.Rmd) that combine text and R code. (Chapter 17) - radial basis function
- A common kernel for support vector machines, which allows curved, rounded boundaries. (Chapter 12)
- random effect
- In a mixed model, variation between groups (such as students) described by a distribution rather than a separate estimate for each group. (Chapter 10)
- random error
- The difference between an estimate and the truth caused by which cases happened to be selected. It shrinks as the sample grows. (Chapter 5)
- random forest
- An ensemble of decision trees, each grown on a bootstrap sample with random subsets of predictors. (Chapter 12)
- random intercept
- A random effect that lets each group have its own average level. (Chapter 10)
- random slope
- A random effect that lets each group have its own effect of a predictor, such as its own rate of change. (Chapter 10)
- randomised experiment
- A study in which the researcher assigns the treatment by chance, so that the groups differ only by chance at the start. (Chapter 5)
- range
- The difference between the largest and smallest values. (Chapter 6)
- rate ratio
- In a Poisson model, the multiplicative change in the expected count per unit of a predictor. (Chapter 10)
- ratio
- A level of measurement with equal distances and a true zero, so that “twice as much” is meaningful, such as hours of sleep. (Chapter 5)
- raw data
- Data as it was collected or exported, before any cleaning; never edited by hand. (Chapter 19)
- reactive expression
-
In Shiny, a calculation, made with
reactive(), that is updated automatically when its inputs change. (Chapter 17) - reactivity
- Shiny’s system for updating outputs automatically when inputs change. (Chapter 17)
- recall
- The share of truly positive cases that a model predicts as positive. Also: sensitivity. (Chapter 12)
- recipe
- In tidymodels, a list of data preparation steps, such as imputation and normalisation. (Chapter 11)
- reference category
- The category of a categorical predictor that the others are compared with. (Chapter 8)
- registered report
- A publication format in which a journal accepts a study’s plan before the data is collected. (Chapter 17)
- regression
- Modelling or predicting a numeric outcome. In machine learning, regression means predicting a number rather than a category. Also: Regression (prediction). (Chapters 11, 13)
- regular expression
-
A pattern for matching text, such as
"\\b(money|fee)". (Chapter 18) - regularisation
- Penalising large coefficients to prevent overfitting, as in ridge and lasso regression. (Chapter 13)
- relational question
- A research question about which variables go together. It establishes association, not cause. (Chapter 5)
- reliability
- How consistently a scale measures what it measures. (Chapters 5, 9)
- ReLU
- Rectified linear unit: an activation function that sets negative values to zero. (Chapter 15)
- remainder
- In a time series decomposition, what is left after removing trend and season. (Chapter 16)
- render
- Running the code in a Quarto document and producing the finished output. (Chapter 17)
- renv
- A package that records and restores the package versions used by a project. (Chapter 17)
- repeated measures
- Several measurements of the same cases, such as each student in four semesters. (Chapter 10)
- replicability
- Whether a new study, with new data, reaches the same conclusions. (Chapter 17)
- reporting sentence
- A sentence in a results section that states a finding with its numbers, such as an estimate, interval, and p-value. (Chapter 19)
- reporting standard
- A published list of the information a research report must contain, such as JARS for psychology, CONSORT for randomised trials, and STROBE for observational studies. (Chapter 19)
- repository
- A project folder whose history git keeps; also its online copy on GitHub. (Chapter 17)
- reproducibility
- Whether the same data and code give exactly the same results. (Chapter 17)
- reproducible research
- Research whose results can be recomputed from the shared data and code. (Chapter 1)
- research question
- A focused question that says precisely what a study will find out, narrow enough that data can answer it. (Chapter 5)
- research workflow
- The path of a research project from question and data to reported results. (Chapter 19)
- residual
- The difference between an observed value and the value a model predicts. Also: error. (Chapters 8, 13)
- reversed item
- A questionnaire item worded in the opposite direction to the others, which must be reversed before scoring. (Chapters 3, 9)
- ridge regression
- Regularised regression with a penalty on the squared coefficients, which shrinks them all towards zero. (Chapter 13)
- RMSE
- Root mean squared error: the square root of the average squared prediction error. (Chapter 13)
- ROC AUC
- The area under the ROC curve: how often a model ranks a positive case above a negative one; 0.5 is chance, 1 is perfect. (Chapter 11)
- ROC curve
- A plot of sensitivity against the false positive rate across all classification thresholds. (Chapter 12)
- rotation
- In factor analysis, turning the factors to make the loadings easier to interpret. (Chapter 9)
- R-squared
- The share of the variation in the outcome that a model explains, from 0 to 1. Also: R². (Chapters 8, 13)
- RStudio
- The most widely used editor for R, used throughout this book. (Chapter 1)
- RStudio Project
- A folder that RStudio treats as the home of one piece of work, setting the working directory. (Chapter 1)
S
- sample
- The part of a population that is observed. (Chapters 5, 7)
- sample description (“Table 1”)
- A table describing the participants, usually the first table of a results chapter. Also: sample description. (Chapter 19)
- sampling distribution
- The distribution of a statistic over many possible samples. (Chapter 7)
- sampling frame
- A list of the members of a population, from which a sample can be drawn. (Chapter 5)
- scale
- In ggplot2, the control of how data values are turned into positions, colours, or sizes. (Chapter 4)
- scale score
- A score combining several questionnaire items, usually their mean. (Chapter 3)
- scaling
- Converting variables to a common scale, usually z-scores, before calculating distances. (Chapter 9)
- scatter plot
- A plot of two numeric variables as points. (Chapter 4)
- scree plot
- A plot of eigenvalues, used to choose the number of components or factors. (Chapter 9)
- script
- A file of R code that can be saved and rerun. (Chapter 1)
- seasonal naive method
- A benchmark forecast that repeats the value from the same season a cycle earlier. (Chapter 16)
- seasonal period
- The length of a repeating pattern, such as 52 weeks or 12 months. (Chapter 16)
- seasonal plot
- A plot that overlays the cycles of a time series to show its seasonal pattern. (Chapter 16)
- seasonality
- A pattern in a time series that repeats at a fixed period. (Chapter 16)
- seasonally adjusted series
- A time series with the seasonal component removed. (Chapter 16)
- self-report
- A measure based on what participants say about themselves, such as their estimated hours of sleep. (Chapter 5)
- self-selection
- When people choose their own group, such as volunteering for a workshop, so that the groups may differ from the start. (Chapter 5)
- sequential palette
- Colours running from light to dark, used for quantities that go in one direction. (Chapter 4)
- server
- In Shiny, the function that computes the outputs from the inputs. (Chapter 17)
- setting
-
In ggplot2, giving a visual property a fixed value, such as
colour = "blue", rather than mapping it to a variable. (Chapter 4) - Shiny
- An R package for building interactive web applications and dashboards. (Chapter 17)
- shrinkage
- Pulling coefficients towards zero, as regularisation does. (Chapter 13)
- sigmoid
- The S-shaped function that turns any number into a value between 0 and 1. (Chapter 15)
- significance level
- The threshold (usually 0.05) below which a p-value counts as statistically significant. (Chapter 7)
- silhouette
- A measure of how much closer each case is to its own cluster than to the next nearest, from −1 to 1. (Chapters 9, 14)
- simple random sampling
- Sampling in which every member of the population has the same chance of being chosen. (Chapter 5)
- simulation
- Using random numbers to generate many possible outcomes, for example many possible futures of a time series. (Chapter 16)
- singular fit
- A warning that a mixed model estimates some variation as zero, often because the model is too complex for the data. (Chapter 10)
- skewness
- Asymmetry of a distribution; a long tail to the right is positive skew. (Chapter 6)
- slope
- The change in the predicted outcome for a one-unit change in a predictor. (Chapter 8)
- small cells
- Combinations of variables that describe very few people, who could be identified. (Chapter 17)
- SMOTE
- A method that balances classes by creating artificial cases of the rare class between existing ones. (Chapter 12)
- specificity
- The share of truly negative cases that a model predicts as negative. (Chapter 12)
- stability
- Whether clusters or results stay the same under small changes to the data or the method. (Chapter 14)
- standard deviation
- A measure of spread: roughly the typical distance of values from the mean. (Chapter 6)
- standard error
- The standard deviation of a sampling distribution; the uncertainty of an estimate. (Chapter 7)
- statistic
- A number calculated from a sample, such as a mean or a t value. (Chapter 7)
- statistically significant
- A result with a p-value below the significance level. (Chapter 7)
- STL
- Seasonal and trend decomposition using loess: a flexible method for decomposing a time series. (Chapter 16)
- stratified sampling
- Sampling in which the population is divided into groups (strata) and a random sample is drawn from each. (Chapter 5)
- stratified split
- A split of the data that keeps the same share of each outcome class in each part. (Chapter 11)
- structured output
- A language model’s answer in a fixed format, such as one category from a list, that code can use directly. (Chapter 18)
- study design
- The plan for who is measured, on what, when, and under which conditions. (Chapter 5)
- supervised learning
- Machine learning with a known outcome to predict. (Chapter 11)
- support vector
- A case close to the boundary of a support vector machine, which determines where the boundary lies. (Chapter 12)
- support vector machine
- A classifier that separates classes with the widest possible margin. (Chapter 12)
- symmetrical
- Having the same shape on both sides of the centre. (Chapter 6)
- synthetic data
- Artificial data with the structure of real data, used when the real data cannot be shared. (Chapter 17)
- system prompt
- Instructions given to a language model before the conversation, such as a task and a codebook. (Chapter 18)
T
- temperature
- A setting of a language model that controls how random its answers are; 0 gives the most likely answer. (Chapter 18)
- test set
- Data set aside to evaluate a final model once, at the end. (Chapter 11)
- test-retest reliability
- The extent to which a measure gives similar results for the same people on two occasions. (Chapter 5)
- theme
- In ggplot2, the non-data appearance of a plot, such as fonts and background. In qualitative coding, a category of meaning. (Chapter 4)
- threshold
- The probability above which a classifier predicts the positive class. (Chapter 12)
- tibble
- The tidyverse’s version of a data frame, which prints more neatly. (Chapters 2, 3)
- tidy data
- Data with one variable per column, one observation per row, and one value per cell. (Chapter 3)
- tidyverse
- A collection of R packages, such as dplyr and ggplot2, that share one way of working. (Chapter 3)
- time plot
- A plot of a time series against time. (Chapter 16)
- time series
- Observations of one quantity at regular intervals over time. (Chapter 16)
- token
- A word or part of a word, the unit of text a language model reads and writes. (Chapter 18)
- training
- Adjusting a model’s weights or parameters to fit the training data. (Chapter 15)
- training cut-off
- The date after which a language model’s training data contains no information. (Chapter 18)
- training set
- The data used to build and tune a model. (Chapter 11)
- transformer
- The neural network architecture behind large language models. (Chapter 15)
- tree depth
- The number of questions in a row a decision tree may ask. (Chapter 13)
- trend
- The long-term direction of a time series. (Chapter 16)
- trend line
- A line added to a scatter plot to show the overall relationship. (Chapter 4)
- true negative
- A negative case correctly predicted negative. (Chapter 12)
- true positive
- A positive case correctly predicted positive. (Chapter 12)
- tsibble
- A data frame for time series, which knows which column is time. (Chapter 16)
- Tukey’s test
- A post-hoc test comparing every pair of groups, with a correction for multiple comparisons. (Chapter 8)
- tuning
- Choosing hyperparameters by comparing their performance, usually with cross-validation. (Chapter 11)
- two-sample t-test
- A test comparing the means of two independent groups. (Chapter 7)
- two-tailed test
- A test that looks for a difference in either direction. (Chapter 7)
- two-way ANOVA
- ANOVA with two factors, including their interaction. (Chapter 8)
- type conversion
- Changing a value from one data type to another, such as text to numbers. (Chapter 2)
- Type I error
- Rejecting a true null hypothesis: a false alarm. (Chapter 7)
- Type II error
- Failing to reject a false null hypothesis: a missed effect. (Chapter 7)
U
- uncertainty
- In a mixture model, 1 minus a case’s largest membership probability. (Chapter 14)
- underfitting
- A model too simple to capture the real pattern in the data. (Chapter 11)
- unimodal
- Having one peak. (Chapter 6)
- unit of analysis
- What one row of the data represents, and what the question is about: a student, a semester, a supervisor. (Chapters 2, 5)
- unsupervised learning
- Machine learning without an outcome, looking for structure such as clusters. (Chapter 11)
- upsampling
- Balancing classes by repeating cases of the rare class in the training data. (Chapter 12)
- user interface
- In Shiny, what the user sees: inputs and places for outputs. (Chapter 17)
V
- validity
- Whether a measure measures what it claims to, or whether a study’s conclusions are justified. (Chapter 5)
- value label
- A label attached to a coded value, such as 1 = “Strongly disagree”, as in SPSS files. (Chapter 2)
- variable
- One characteristic measured on every case, such as age or faculty; one column of a data table. (Chapter 2)
- variable label
- A longer description attached to a variable, as in SPSS files. (Chapter 2)
- variable selection
- Choosing which predictors to keep in a model; the lasso does it automatically. (Chapter 13)
- variance
- The average squared distance from the mean; the square of the standard deviation. (Chapter 6)
- vector
- R’s basic structure: a sequence of values of the same type. (Chapters 1, 2)
- vector format
- An image format, such as PDF or SVG, that stays sharp at any size. (Chapter 4)
- verb
-
A dplyr function that does one thing to a data frame, such as
filter()ormutate(). (Chapter 3) - version control
- Keeping a history of every change to a project’s files. (Chapter 17)
W
- Ward’s method
- A way of merging clusters in hierarchical clustering that keeps them as compact as possible. (Chapter 9)
- weak learner
- A simple model, such as a small tree, that is only slightly better than chance on its own; boosting combines many. (Chapter 13)
- weight
- In a neural network, the number by which an input is multiplied; learned during training. (Chapter 15)
- weight decay
- A penalty on large weights in a neural network, which prevents overfitting. (Chapter 15)
- Welch test
- The version of the two-sample t-test that does not assume equal variances; R’s default. (Chapter 7)
- wide format
- Data with repeated measurements side by side in columns, one row per case. (Chapter 3)
- Wilcoxon signed-rank test
- A non-parametric test comparing two measurements on the same cases. (Chapter 7)
- within-group variation
- How much values vary around their own group mean; the denominator of the ANOVA F statistic. (Chapter 8)
- workflow
- In tidymodels, a recipe and a model specification combined. (Chapter 11)
- workflow set
- In tidymodels, several workflows fitted and compared together. (Chapter 12)
- working directory
- The folder R reads files from and saves files to by default. (Chapter 1)
X
- XGBoost
- Extreme gradient boosting: a fast, widely used boosting method. (Chapter 13)
Y
- YAML header
-
The settings at the top of a Quarto document, between two
---lines. (Chapter 17)
Z
- z-score
- A value expressed as the number of standard deviations from the mean. (Chapter 6)
R functions
Table 1.1 lists the functions used most often in the book, grouped by task, with the package they come from and the chapters that use them. base R functions are available without loading any package. For any function, ?name in the Console opens its help page.