Appendix D: Common Errors and How to Fix Them

Everyone who writes R code sees error messages every day, experts included. An error is not a sign that you are doing badly; it is R telling you, as precisely as it can, what it could not do. This appendix collects the messages that readers of this book are most likely to meet, what each one means, and how to fix it. The messages below are real: each example was run when the book was built, so the wording is what your R will show (it can differ slightly between versions).

Reading a message

R produces three kinds of messages:

  • An error stops the code: nothing after it runs. It starts with Error.
  • A warning lets the code finish but tells you something may be wrong. It starts with Warning. Never ignore a warning without understanding it: several below mean your results contain missing values or were calculated differently from what you intended.
  • A message is information only, such as a package telling you it has loaded.

When a long message appears, read it from the end: the last lines usually say what went wrong, and lines starting with ℹ point to where. Newer packages (dplyr, ggplot2, tidymodels) write especially helpful messages, often with a suggestion (Did you mean ...?).

Starting out

object ‘…’ not found

mean(sleep_hours)
Error in h(simpleError(msg, call)): error in evaluating the argument 'x' in selecting a method for function 'mean': object 'sleep_hours' not found

R does not know the name. The usual causes: a typing mistake (names are case-sensitive: Sleep_hours is not sleep_hours); the object was never created, because the line that creates it was not run; or the name is a column inside a data frame, which must be reached through the data frame, as in mean(semesters$sleep_hours) or inside a dplyr verb. If a document fails to render with this error although the code works in the Console, the object was created in the Console but not in the document (Chapter 17).

could not find function “…”

Mean(c(6.5, 7, 5.5))
Error in Mean(c(6.5, 7, 5.5)): could not find function "Mean"

Either the function name is misspelled (here, mean with a capital M), or it belongs to a package that is not loaded. Load the package with library(), or write package::function(). The pipe %>% from older code gives the same error until dplyr or magrittr is loaded; the native pipe |> needs no package.

there is no package called ‘…’

library(ggplot)
Error in library(ggplot): there is no package called 'ggplot'

The package is not installed, or its name is misspelled (the package is ggplot2). Install it once with install.packages("ggplot2"), then load it with library(ggplot2). Appendix A lists every package the book uses.

unexpected symbol, unexpected ‘)’, unexpected string constant

mean(c(6.5, 7) na.rm = TRUE)
Error in parse(text = input): <text>:1:16: unexpected symbol
1: mean(c(6.5, 7) na.rm
                   ^

A syntax error: R cannot read the line at all. The ^ marks where R got lost; the mistake is usually just before it. Here a comma is missing before na.rm. Other common causes are an extra or missing bracket (unexpected ')'), a missing comma between two pieces of text (unexpected string constant), and a missing quotation mark. RStudio highlights matching brackets and marks syntax errors with a red cross in the margin before you even run the code.

The Console shows + and nothing happens

R is waiting for the rest of an unfinished command, usually because a bracket or quotation mark is not closed. Press Esc to cancel, fix the line, and run it again.

argument “…” is missing, with no default

rnorm()
Error in rnorm(): argument "n" is missing, with no default

A required argument was not given. The help page (?rnorm) lists the arguments; those without a default value (here n, the number of values) must be supplied.

A misspelled argument, with no error at all

mean(c(6.5, NA, 7), na_rm = TRUE)
[1] NA

Not every mistake produces a message. The argument is na.rm, not na_rm, and because mean() accepts extra arguments, the misspelled one is silently ignored: the missing value is not removed, and the answer is NA. When a result looks wrong, check the spelling of every argument against the help page.

Files and folders

cannot open file …: No such file or directory

read.csv("studnets.csv")
Warning in file(file, "rt"): cannot open file 'studnets.csv': No such file or
directory
Error in file(file, "rt"): cannot open the connection

R looked for the file in the working directory and did not find it. Check the spelling of the file name (here studnets), check that the file is in the project folder, and build paths with here::here() inside an RStudio Project (Chapter 1), so they work on any computer. list.files() shows the files R can see.

cannot change working directory

setwd("C:/Users/Elaf/Documents/thesis")
Error in setwd("C:/Users/Elaf/Documents/thesis"): cannot change working directory

The folder does not exist on this computer, which is exactly why setwd() with a full path breaks as soon as a script is shared or moved. Use an RStudio Project and relative paths instead (Chapter 1).

Data

non-numeric argument to binary operator

"7" + 1
Error in "7" + 1: non-numeric argument to binary operator

Arithmetic on text. The quotation marks make "7" a piece of text, not a number. In real data, this happens when a numeric column was imported as text, often because of one stray value such as "7 hrs" or a missing-value code. Check with str() or glimpse(), and convert with as.numeric() or readr::parse_number() (Chapters 2 and 3).

argument is not numeric or logical: returning NA

mean(c("6.5", "7"))
Warning in mean.default(c("6.5", "7")): argument is not numeric or logical:
returning NA
[1] NA

A warning, not an error, and the result is NA: the same problem as above, a numeric column stored as text. Convert the column first.

NAs introduced by coercion

as.numeric(c("6.5", "7 hrs", "six"))
Warning: NAs introduced by coercion
[1] 6.5  NA  NA

Values that could not be converted to numbers became NA. Look at which ones ("7 hrs" and "six" here) before going on: parse_number() can rescue "7 hrs", but "six" needs a decision (Chapter 3). Losing data silently through this warning is one of the most common problems in real analyses.

The result is NA

mean(c(6.5, NA, 7))
[1] NA

Not an error: any calculation that includes a missing value gives NA, because the true answer is unknown. Add na.rm = TRUE to calculate from the available values, and report how many were missing (Chapter 6).

object ‘…’ not found, inside dplyr

students |> filter(facultty == "Education")
Error in `filter()`:
ℹ In argument: `facultty == "Education"`.
Caused by error:
! object 'facultty' not found

A column name is misspelled. dplyr says in which argument the problem is (ℹ In argument: ...) and which name it could not find. names(students) lists the correct names.

arguments imply differing number of rows; replacement has … rows

data.frame(student = 1:3, sleep = c(6.5, 7))
Error in data.frame(student = 1:3, sleep = c(6.5, 7)): arguments imply differing number of rows: 3, 2

Every column of a data frame must have the same length. The data behind a new column has more or fewer values than the data frame has rows; check the lengths with length() and nrow().

$ operator is invalid for atomic vectors

x <- c(sleep = 6.5, study = 30)
x$sleep
Error in x$sleep: $ operator is invalid for atomic vectors

$ works on data frames and lists, but x is a vector. Use x["sleep"] for a vector, and check what kind of object you have with class() or str().

subscript out of bounds

results <- list(6.5, 7)
results[[3]]
Error in results[[3]]: subscript out of bounds

You asked for an element that does not exist: the third element of a list with two. Check the length with length(), and the names with names().

invalid factor level, NA generated

answer <- factor(c("Yes", "No"))
answer[1] <- "Maybe"
Warning in `[<-.factor`(`*tmp*`, 1, value = "Maybe"): invalid factor level, NA
generated
answer
[1] <NA> No  
Levels: No Yes

A factor accepts only its existing levels (Chapter 2); any other value becomes NA. Add the level first with levels(), or work with the column as text and convert it to a factor at the end.

Detected an unexpected many-to-many relationship

students_small <- data.frame(id = c(1, 1), group = c("A", "B"))
scores_small <- data.frame(id = c(1, 1), score = c(3.2, 4.1))
left_join(students_small, scores_small, join_by(id))
Warning in left_join(students_small, scores_small, join_by(id)): Detected an unexpected many-to-many relationship between `x` and `y`.
ℹ Row 1 of `x` matches multiple rows in `y`.
ℹ Row 1 of `y` matches multiple rows in `x`.
ℹ If a many-to-many relationship is expected, set `relationship =
  "many-to-many"` to silence this warning.
  id group score
1  1     A   3.2
2  1     A   4.1
3  1     B   3.2
4  1     B   4.1

A join (Chapter 3) found identifiers that appear more than once in both tables, so every copy was matched with every other, multiplying rows. Usually one of the tables should have had one row per identifier: check for duplicates with count(id) |> filter(n > 1), and remove them or join by more columns.

Can’t combine … and …

tidyr::pivot_longer(students, cols = c(age, gender))
Error in `tidyr::pivot_longer()`:
! Can't combine `age` <integer> and `gender` <character>.

pivot_longer() puts the chosen columns into one column, so they must be of the same type: here a number and a text column. Reshape only columns of the same type, or convert them first.

Plots

Cannot use + with a single argument

plot <- ggplot(students, aes(x = age))
+ geom_histogram()
Error:
! Cannot use `+` with a single argument.
ℹ Did you accidentally put `+` on a new line?

In ggplot2, the + must come at the end of a line, not the start of the next one. Otherwise R thinks the first line is complete, and the second line starts with a stray +. The message even asks: “Did you accidentally put + on a new line?”

mapping must be created by aes() … Did you use %>% or |> instead of +?

ggplot(students, aes(x = age)) |> geom_histogram()
Error in `geom_histogram()`:
! `mapping` must be created by `aes()`.
✖ You've supplied a <ggplot2::ggplot> object.
ℹ Did you use `%>%` or `|>` instead of `+`?

Layers of a ggplot are added with +, not with the pipe. The pipe passes data into ggplot(); after that, it is + all the way.

object ‘…’ not found, in a plot

ggplot(students) + geom_point(x = age, y = financial_worry)
Error: object 'age' not found

Columns must be mapped inside aes(): geom_point(aes(x = age, y = financial_worry)). Outside aes(), R looks for objects called age and financial_worry and does not find them (Chapter 4 on mapping versus setting).

stat_count() must only have an x or y aesthetic

ggplot(students, aes(x = faculty, y = age)) + geom_bar()
Error in `geom_bar()`:
! Problem while computing stat.
ℹ Error occurred in the 1st layer.
Caused by error in `setup_params()`:
! `stat_count()` must only have an x or y aesthetic.

geom_bar() counts the rows in each category, so it takes only x. To plot a value you have calculated, such as a mean per faculty, use geom_col().

Statistical tests and models

grouping factor must have exactly 2 levels

t.test(age ~ faculty, data = students)
Error in t.test.formula(age ~ faculty, data = students): grouping factor must have exactly 2 levels

A two-sample t-test compares exactly two groups, but faculty has five. Use ANOVA for more than two groups (Chapter 8), or filter the data to the two groups you want to compare.

not enough ‘x’ observations

t.test(c(6.5))
Error in t.test.default(c(6.5)): not enough 'x' observations

The test needs more data than it was given, often because filtering or missing values left too few cases. Check the number of cases with nrow() or sum(!is.na(x)).

contrasts can be applied only to factors with 2 or more levels

lm(age ~ gender, data = filter(students, gender == "Female"))
Error in `contrasts<-`(`*tmp*`, value = contr.funs[1 + isOF[nn]]): contrasts can be applied only to factors with 2 or more levels

A categorical predictor has only one value in the data used, here because the data was filtered to women only, so there is nothing to compare. Remove the predictor, or check the filtering.

Chi-squared approximation may be incorrect

chisq.test(matrix(c(3, 1, 2, 4), nrow = 2))
Warning in chisq.test(matrix(c(3, 1, 2, 4), nrow = 2)): Chi-squared
approximation may be incorrect

    Pearson's Chi-squared test with Yates' continuity correction

data:  matrix(c(3, 1, 2, 4), nrow = 2)
X-squared = 0.41667, df = 1, p-value = 0.5186

Some expected counts are below 5, so the chi-square test’s p-value may be inaccurate (Chapter 7). Use Fisher’s exact test (fisher.test()), or combine small categories.

glm.fit: fitted probabilities numerically 0 or 1 occurred; algorithm did not converge

glm(passed ~ hours, family = binomial,
    data = data.frame(hours = 1:10, passed = c(0, 0, 0, 0, 0, 1, 1, 1, 1, 1)))
Warning: glm.fit: algorithm did not converge
Warning: glm.fit: fitted probabilities numerically 0 or 1 occurred

In a logistic regression, a predictor separates the outcomes perfectly (here, everyone above 5 hours passed), so the coefficients become enormous and meaningless, and the fitting algorithm may also report that it did not converge. With real data, this usually means a very small group or a predictor that is almost the outcome itself. Check the cross-table of the outcome and the predictor, and consider simplifying the model.

boundary (singular) fit: see help(‘isSingular’)

lmer(wellbeing ~ semester + (semester | supervisor_id),
     data = left_join(semesters, students, join_by(student_id)))
boundary (singular) fit: see help('isSingular')

A mixed-effects model estimated some random-effect variance as zero, or a correlation as exactly ±1: the random effects part is too complex for the data (Chapter 10). Simplify it, for example by removing the random slope: (1 | supervisor_id).

Observations deleted due to missingness

Not an error, but a line in the output of summary() for lm() and glm(): rows with a missing value in any variable of the model are left out. Check how many with nobs(model), report the number used, and think about whether the missing rows differ from the others (Chapters 6 and 8).

Machine learning

For a classification model, the outcome should be a

library(tidymodels)
fit(logistic_reg(), considering_dropout ~ age, data = students)
Error in `check_outcome()`:
! For a classification model, the outcome should be a <factor>, not a
  character vector.

tidymodels needs a categorical outcome to be a factor. Convert it first, and set the event of interest as the first level (Chapter 11): mutate(considering_dropout = factor(considering_dropout, levels = c("Yes", "No"))).

Can’t select columns that don’t exist … .pred_Yes

tibble(truth = factor(c("Yes", "No")), probability = c(0.2, 0.8)) |>
  roc_auc(truth, .pred_Yes)
Error in `roc_auc()`:
! Can't select columns that don't exist.
✖ Column `.pred_Yes` doesn't exist.

A column name does not exist in the data given to a yardstick function. The prediction columns created by augment() are named after the outcome’s levels (.pred_Yes, .pred_No), so check the names with names(), and check that the predictions were added to the data.

When you are stuck

  1. Read the message, from the end, and look at the line it points to.
  2. Restart R and run the script from the top (Session > Restart R). Many mysterious errors come from objects left over from earlier work.
  3. Check your data with str(), glimpse(), or summary(): wrong types and unexpected missing values cause most errors in real analyses.
  4. Make the problem small. Reproduce it with a few rows of data or a tiny example, like those at the start of every chapter. Often the cause becomes obvious; if not, a small example is what others need to help you. The reprex package formats such an example for sharing.
  5. Search for the message in quotation marks, or ask an AI assistant, including the code, the full message, and a description of your data (Chapter 18). Check the answer before trusting it.
  6. Ask a person: a colleague, your supervisor, the Posit Community forum, or Stack Overflow. A clear question with a small example usually gets a quick answer.