| Area | Package | What it is for | Chapters | Version |
|---|---|---|---|---|
| Data | data2thesis | Elaf’s Graduate Wellbeing Study: the case-study data of this book | 1, 2, 3, 4, 5, and later | 1.1.0 |
| Data | modeldata | Example datasets for modelling, such as credit applications, concrete strength, and penguins | 11, 12, 13, 15 | 1.6.0 |
| Importing and exporting | readr | Reading and writing CSV and other text files, and parse_number() |
3, 19 | 2.1.5 |
| Importing and exporting | readxl | Reading Excel files | 2, 3, 19 | 1.4.5 |
| Importing and exporting | haven | Reading and writing SPSS, Stata, and SAS files, with their labels | 2, 19 | 2.5.5 |
| Importing and exporting | writexl | Writing Excel files | 2 | 1.5.4 |
| Importing and exporting | here | File paths relative to the project folder | 1, 2, 3, 19 | 1.0.1 |
| Data handling | tidyverse | Installs and loads the core tidyverse packages in one go | 3 | - |
| Data handling | dplyr | Filtering, selecting, creating, summarising, and joining data | 3, 4, 5, 6, 7, and later | 1.2.1 |
| Data handling | tidyr | Reshaping data between wide and long formats | 3, 4, 5, 6, 7, and later | 1.3.2 |
| Data handling | stringr | Working with text: detecting, replacing, and counting patterns | 3, 18, 19 | 1.5.1 |
| Data handling | purrr | Applying a function to each element of a list or vector | 12, 13 | 1.2.2 |
| Visualisation | ggplot2 | Plots built from data, mappings, and layers | 4, 5, 6, 7, 8, and later | 4.0.3 |
| Visualisation | scales | Formatting axis labels, such as percentages | 13 | 1.4.0 |
| Visualisation | patchwork | Combining several plots into one figure | 4, 19 | 1.3.2 |
| Visualisation | corrplot | Plotting correlation matrices | 9 | 0.95 |
| Statistics | psych | Descriptive statistics, factor analysis, and Cronbach’s alpha | 6, 9 | 2.6.5 |
| Statistics | GPArotation | Factor rotations, such as oblimin, used by psych | 9 | 2026.8.2 |
| Statistics | car | Regression tools, including Levene’s test | 8 | 3.1.3 |
| Statistics | broom | Model results as tidy data frames | 8, 19 | 1.0.13 |
| Statistics | lme4 | Linear and generalised linear mixed-effects models | 10, 19 | 1.1.37 |
| Statistics | lmerTest | p-values for mixed-effects models | 10 | 3.2.1 |
| Multivariate and clustering | factoextra | Plots for PCA and cluster analysis | 9, 14 | 1.0.7 |
| Multivariate and clustering | cluster | Clustering tools, including the silhouette | 14 | 2.1.8 |
| Multivariate and clustering | mclust | Gaussian mixture models | 14 | 6.1.2 |
| Multivariate and clustering | dbscan | Density-based clustering (DBSCAN) and related methods | 14 | 1.2.4 |
| Machine learning | tidymodels | Installs and loads the tidymodels packages: rsample, recipes, parsnip, workflows, tune, yardstick, and others | 11, 12, 13, 15 | 1.5.0 |
| Machine learning | yardstick | Measures of model performance, such as accuracy, ROC AUC, and kappa (part of tidymodels) | 11, 12, 13, 15, 18 | 1.4.0 |
| Machine learning | rpart | Decision trees | 12 | 4.1.24 |
| Machine learning | ranger | Fast random forests | 12 | 0.18.0 |
| Machine learning | kknn | k-nearest neighbours, the default engine for nearest_neighbor() |
11, 12 | 1.4.1 |
| Machine learning | kernlab | Support vector machines | 12 | 0.9.33 |
| Machine learning | themis | Resampling steps for imbalanced outcomes, such as upsampling and SMOTE | 12 | 1.1.0 |
| Machine learning | glmnet | Ridge, lasso, and elastic net regression | 13 | 4.1.10 |
| Machine learning | xgboost | Gradient boosting (XGBoost) | 13 | 3.2.1.1 |
| Machine learning | nnet | Neural networks with one hidden layer | 15 | 7.3.20 |
| Time series | tsibble | Data frames for time series | 16 | 1.2.0 |
| Time series | fable | Forecasting models: benchmarks, ETS, and ARIMA | 16 | 0.5.0 |
| Time series | feasts | Time series features, decomposition (STL), and autocorrelation | 16 | 0.5.0 |
| Time series | urca | Unit-root tests, needed by ARIMA() |
16 | 1.3.4 |
| Reporting and sharing | knitr | Running the code in Quarto documents, and kable() tables |
17, 19 | 1.50 |
| Reporting and sharing | rmarkdown | R Markdown documents, the predecessor of Quarto | 17 | 2.29 |
| Reporting and sharing | shiny | Interactive web applications and dashboards | 17 | 1.10.0 |
| Reporting and sharing | bslib | Modern page layouts for Shiny apps | 17 | 0.9.0 |
| Reporting and sharing | renv | Recording and restoring package versions | 17, 19 | 1.1.4 |
| Reporting and sharing | usethis | Project setup tasks, such as editing .Renviron and connecting to git |
17, 18 | - |
| AI | ellmer | Calling large language models from R | 18 | 0.4.0 |
Appendix C: R Packages
This appendix lists every R package used in the book: what it is for, the chapters that use it, and the version used when the book was built. Appendix A shows how to install them all with one command. A dash in the Version column marks a package that is recommended to readers but was not needed to build the book itself.
A package is a collection of functions, data, and documentation that someone has written and shared. R comes with a set of base packages, such as stats (t.test(), lm()) and utils (read.csv()), which are always available; everything else is installed once with install.packages() and loaded in each session with library() (Chapter 1).
The packages in this book
Two entries are collections rather than single packages. tidyverse installs and loads dplyr, tidyr, stringr, readr, ggplot2, purrr, and a few others; tidymodels does the same for the modelling packages of Chapters 11 to 15 (rsample, recipes, parsnip, workflows, tune, and yardstick, among others). Several model packages (ranger, kknn, kernlab, glmnet, xgboost, nnet) are rarely loaded by name: tidymodels calls them as engines, as in set_engine("ranger"), but they must be installed.
Finding and choosing packages
CRAN, R’s official package archive, holds more than 20,000 packages, and many more are shared on GitHub. A few ways to find the right one:
- CRAN Task Views (cran.r-project.org/web/views) are curated lists of packages by topic, such as psychometrics, survival analysis, or time series, maintained by experts in each field.
- A package’s vignettes, long-form tutorials that come with it, are the best introduction:
browseVignettes("dplyr")lists them. Many packages also have a website with examples. - Methods papers in journals such as the Journal of Statistical Software and The R Journal describe many packages in depth.
Before relying on a package for your thesis, a few signs show whether it is trustworthy: it is on CRAN (which checks that packages install and run); it was updated in the last year or two; it has documentation and examples; it is described in a published paper or widely used in your field; and its authors are known in the area. The packages in this book meet all or most of these.
Citing packages
Package authors are researchers too, and citing their work is how they receive credit. citation() gives the recommended reference for R itself, and citation("package") for a package:
citation("psych")To cite package 'psych' in publications use:
William Revelle (2026). _psych: Procedures for Psychological,
Psychometric, and Personality Research_. Northwestern University,
Evanston, Illinois. R package version 2.6.4,
<https://CRAN.R-project.org/package=psych>.
A BibTeX entry for LaTeX users is
@Manual{,
title = {psych: Procedures for Psychological, Psychometric, and Personality Research},
author = {{William Revelle}},
organization = {Northwestern University},
address = {Evanston, Illinois},
year = {2026},
note = {R package version 2.6.4},
url = {https://CRAN.R-project.org/package=psych},
}
Cite the packages that did substantial work in your analysis, and give their version numbers (Chapter 17). packageVersion("psych") shows the version installed on your computer.