Appendix C: R Packages

This appendix lists every R package used in the book: what it is for, the chapters that use it, and the version used when the book was built. Appendix A shows how to install them all with one command. A dash in the Version column marks a package that is recommended to readers but was not needed to build the book itself.

A package is a collection of functions, data, and documentation that someone has written and shared. R comes with a set of base packages, such as stats (t.test(), lm()) and utils (read.csv()), which are always available; everything else is installed once with install.packages() and loaded in each session with library() (Chapter 1).

The packages in this book

Table 1.1: The R packages used in this book.
Area Package What it is for Chapters Version
Data data2thesis Elaf’s Graduate Wellbeing Study: the case-study data of this book 1, 2, 3, 4, 5, and later 1.1.0
Data modeldata Example datasets for modelling, such as credit applications, concrete strength, and penguins 11, 12, 13, 15 1.6.0
Importing and exporting readr Reading and writing CSV and other text files, and parse_number() 3, 19 2.1.5
Importing and exporting readxl Reading Excel files 2, 3, 19 1.4.5
Importing and exporting haven Reading and writing SPSS, Stata, and SAS files, with their labels 2, 19 2.5.5
Importing and exporting writexl Writing Excel files 2 1.5.4
Importing and exporting here File paths relative to the project folder 1, 2, 3, 19 1.0.1
Data handling tidyverse Installs and loads the core tidyverse packages in one go 3 -
Data handling dplyr Filtering, selecting, creating, summarising, and joining data 3, 4, 5, 6, 7, and later 1.2.1
Data handling tidyr Reshaping data between wide and long formats 3, 4, 5, 6, 7, and later 1.3.2
Data handling stringr Working with text: detecting, replacing, and counting patterns 3, 18, 19 1.5.1
Data handling purrr Applying a function to each element of a list or vector 12, 13 1.2.2
Visualisation ggplot2 Plots built from data, mappings, and layers 4, 5, 6, 7, 8, and later 4.0.3
Visualisation scales Formatting axis labels, such as percentages 13 1.4.0
Visualisation patchwork Combining several plots into one figure 4, 19 1.3.2
Visualisation corrplot Plotting correlation matrices 9 0.95
Statistics psych Descriptive statistics, factor analysis, and Cronbach’s alpha 6, 9 2.6.5
Statistics GPArotation Factor rotations, such as oblimin, used by psych 9 2026.8.2
Statistics car Regression tools, including Levene’s test 8 3.1.3
Statistics broom Model results as tidy data frames 8, 19 1.0.13
Statistics lme4 Linear and generalised linear mixed-effects models 10, 19 1.1.37
Statistics lmerTest p-values for mixed-effects models 10 3.2.1
Multivariate and clustering factoextra Plots for PCA and cluster analysis 9, 14 1.0.7
Multivariate and clustering cluster Clustering tools, including the silhouette 14 2.1.8
Multivariate and clustering mclust Gaussian mixture models 14 6.1.2
Multivariate and clustering dbscan Density-based clustering (DBSCAN) and related methods 14 1.2.4
Machine learning tidymodels Installs and loads the tidymodels packages: rsample, recipes, parsnip, workflows, tune, yardstick, and others 11, 12, 13, 15 1.5.0
Machine learning yardstick Measures of model performance, such as accuracy, ROC AUC, and kappa (part of tidymodels) 11, 12, 13, 15, 18 1.4.0
Machine learning rpart Decision trees 12 4.1.24
Machine learning ranger Fast random forests 12 0.18.0
Machine learning kknn k-nearest neighbours, the default engine for nearest_neighbor() 11, 12 1.4.1
Machine learning kernlab Support vector machines 12 0.9.33
Machine learning themis Resampling steps for imbalanced outcomes, such as upsampling and SMOTE 12 1.1.0
Machine learning glmnet Ridge, lasso, and elastic net regression 13 4.1.10
Machine learning xgboost Gradient boosting (XGBoost) 13 3.2.1.1
Machine learning nnet Neural networks with one hidden layer 15 7.3.20
Time series tsibble Data frames for time series 16 1.2.0
Time series fable Forecasting models: benchmarks, ETS, and ARIMA 16 0.5.0
Time series feasts Time series features, decomposition (STL), and autocorrelation 16 0.5.0
Time series urca Unit-root tests, needed by ARIMA() 16 1.3.4
Reporting and sharing knitr Running the code in Quarto documents, and kable() tables 17, 19 1.50
Reporting and sharing rmarkdown R Markdown documents, the predecessor of Quarto 17 2.29
Reporting and sharing shiny Interactive web applications and dashboards 17 1.10.0
Reporting and sharing bslib Modern page layouts for Shiny apps 17 0.9.0
Reporting and sharing renv Recording and restoring package versions 17, 19 1.1.4
Reporting and sharing usethis Project setup tasks, such as editing .Renviron and connecting to git 17, 18 -
AI ellmer Calling large language models from R 18 0.4.0

Two entries are collections rather than single packages. tidyverse installs and loads dplyr, tidyr, stringr, readr, ggplot2, purrr, and a few others; tidymodels does the same for the modelling packages of Chapters 11 to 15 (rsample, recipes, parsnip, workflows, tune, and yardstick, among others). Several model packages (ranger, kknn, kernlab, glmnet, xgboost, nnet) are rarely loaded by name: tidymodels calls them as engines, as in set_engine("ranger"), but they must be installed.

Finding and choosing packages

CRAN, R’s official package archive, holds more than 20,000 packages, and many more are shared on GitHub. A few ways to find the right one:

  • CRAN Task Views (cran.r-project.org/web/views) are curated lists of packages by topic, such as psychometrics, survival analysis, or time series, maintained by experts in each field.
  • A package’s vignettes, long-form tutorials that come with it, are the best introduction: browseVignettes("dplyr") lists them. Many packages also have a website with examples.
  • Methods papers in journals such as the Journal of Statistical Software and The R Journal describe many packages in depth.

Before relying on a package for your thesis, a few signs show whether it is trustworthy: it is on CRAN (which checks that packages install and run); it was updated in the last year or two; it has documentation and examples; it is described in a published paper or widely used in your field; and its authors are known in the area. The packages in this book meet all or most of these.

Citing packages

Package authors are researchers too, and citing their work is how they receive credit. citation() gives the recommended reference for R itself, and citation("package") for a package:

citation("psych")
To cite package 'psych' in publications use:

  William Revelle (2026). _psych: Procedures for Psychological,
  Psychometric, and Personality Research_. Northwestern University,
  Evanston, Illinois. R package version 2.6.4,
  <https://CRAN.R-project.org/package=psych>.

A BibTeX entry for LaTeX users is

  @Manual{,
    title = {psych: Procedures for Psychological, Psychometric, and Personality Research},
    author = {{William Revelle}},
    organization = {Northwestern University},
    address = {Evanston, Illinois},
    year = {2026},
    note = {R package version 2.6.4},
    url = {https://CRAN.R-project.org/package=psych},
  }

Cite the packages that did substantial work in your analysis, and give their version numbers (Chapter 17). packageVersion("psych") shows the version installed on your computer.