mean(sleep_hours)Error in h(simpleError(msg, call)): error in evaluating the argument 'x' in selecting a method for function 'mean': object 'sleep_hours' not found
Everyone who writes R code sees error messages every day, experts included. An error is not a sign that you are doing badly; it is R telling you, as precisely as it can, what it could not do. This appendix collects the messages that readers of this book are most likely to meet, what each one means, and how to fix it. The messages below are real: each example was run when the book was built, so the wording is what your R will show (it can differ slightly between versions).
R produces three kinds of messages:
Error.Warning. Never ignore a warning without understanding it: several below mean your results contain missing values or were calculated differently from what you intended.When a long message appears, read it from the end: the last lines usually say what went wrong, and lines starting with ℹ point to where. Newer packages (dplyr, ggplot2, tidymodels) write especially helpful messages, often with a suggestion (Did you mean ...?).
Error in h(simpleError(msg, call)): error in evaluating the argument 'x' in selecting a method for function 'mean': object 'sleep_hours' not found
R does not know the name. The usual causes: a typing mistake (names are case-sensitive: Sleep_hours is not sleep_hours); the object was never created, because the line that creates it was not run; or the name is a column inside a data frame, which must be reached through the data frame, as in mean(semesters$sleep_hours) or inside a dplyr verb. If a document fails to render with this error although the code works in the Console, the object was created in the Console but not in the document (Chapter 17).
Either the function name is misspelled (here, mean with a capital M), or it belongs to a package that is not loaded. Load the package with library(), or write package::function(). The pipe %>% from older code gives the same error until dplyr or magrittr is loaded; the native pipe |> needs no package.
The package is not installed, or its name is misspelled (the package is ggplot2). Install it once with install.packages("ggplot2"), then load it with library(ggplot2). Appendix A lists every package the book uses.
Error in parse(text = input): <text>:1:16: unexpected symbol
1: mean(c(6.5, 7) na.rm
^
A syntax error: R cannot read the line at all. The ^ marks where R got lost; the mistake is usually just before it. Here a comma is missing before na.rm. Other common causes are an extra or missing bracket (unexpected ')'), a missing comma between two pieces of text (unexpected string constant), and a missing quotation mark. RStudio highlights matching brackets and marks syntax errors with a red cross in the margin before you even run the code.
+ and nothing happensR is waiting for the rest of an unfinished command, usually because a bracket or quotation mark is not closed. Press Esc to cancel, fix the line, and run it again.
A required argument was not given. The help page (?rnorm) lists the arguments; those without a default value (here n, the number of values) must be supplied.
Not every mistake produces a message. The argument is na.rm, not na_rm, and because mean() accepts extra arguments, the misspelled one is silently ignored: the missing value is not removed, and the answer is NA. When a result looks wrong, check the spelling of every argument against the help page.
Warning in file(file, "rt"): cannot open file 'studnets.csv': No such file or
directory
Error in file(file, "rt"): cannot open the connection
R looked for the file in the working directory and did not find it. Check the spelling of the file name (here studnets), check that the file is in the project folder, and build paths with here::here() inside an RStudio Project (Chapter 1), so they work on any computer. list.files() shows the files R can see.
Error in setwd("C:/Users/Elaf/Documents/thesis"): cannot change working directory
The folder does not exist on this computer, which is exactly why setwd() with a full path breaks as soon as a script is shared or moved. Use an RStudio Project and relative paths instead (Chapter 1).
Arithmetic on text. The quotation marks make "7" a piece of text, not a number. In real data, this happens when a numeric column was imported as text, often because of one stray value such as "7 hrs" or a missing-value code. Check with str() or glimpse(), and convert with as.numeric() or readr::parse_number() (Chapters 2 and 3).
Warning in mean.default(c("6.5", "7")): argument is not numeric or logical:
returning NA
[1] NA
A warning, not an error, and the result is NA: the same problem as above, a numeric column stored as text. Convert the column first.
Values that could not be converted to numbers became NA. Look at which ones ("7 hrs" and "six" here) before going on: parse_number() can rescue "7 hrs", but "six" needs a decision (Chapter 3). Losing data silently through this warning is one of the most common problems in real analyses.
Not an error: any calculation that includes a missing value gives NA, because the true answer is unknown. Add na.rm = TRUE to calculate from the available values, and report how many were missing (Chapter 6).
Error in `filter()`:
ℹ In argument: `facultty == "Education"`.
Caused by error:
! object 'facultty' not found
A column name is misspelled. dplyr says in which argument the problem is (ℹ In argument: ...) and which name it could not find. names(students) lists the correct names.
Error in data.frame(student = 1:3, sleep = c(6.5, 7)): arguments imply differing number of rows: 3, 2
Every column of a data frame must have the same length. The data behind a new column has more or fewer values than the data frame has rows; check the lengths with length() and nrow().
$ works on data frames and lists, but x is a vector. Use x["sleep"] for a vector, and check what kind of object you have with class() or str().
You asked for an element that does not exist: the third element of a list with two. Check the length with length(), and the names with names().
Warning in `[<-.factor`(`*tmp*`, 1, value = "Maybe"): invalid factor level, NA
generated
[1] <NA> No
Levels: No Yes
A factor accepts only its existing levels (Chapter 2); any other value becomes NA. Add the level first with levels(), or work with the column as text and convert it to a factor at the end.
Warning in left_join(students_small, scores_small, join_by(id)): Detected an unexpected many-to-many relationship between `x` and `y`.
ℹ Row 1 of `x` matches multiple rows in `y`.
ℹ Row 1 of `y` matches multiple rows in `x`.
ℹ If a many-to-many relationship is expected, set `relationship =
"many-to-many"` to silence this warning.
id group score
1 1 A 3.2
2 1 A 4.1
3 1 B 3.2
4 1 B 4.1
A join (Chapter 3) found identifiers that appear more than once in both tables, so every copy was matched with every other, multiplying rows. Usually one of the tables should have had one row per identifier: check for duplicates with count(id) |> filter(n > 1), and remove them or join by more columns.
Error in `tidyr::pivot_longer()`:
! Can't combine `age` <integer> and `gender` <character>.
pivot_longer() puts the chosen columns into one column, so they must be of the same type: here a number and a text column. Reshape only columns of the same type, or convert them first.
+ with a single argumentError:
! Cannot use `+` with a single argument.
ℹ Did you accidentally put `+` on a new line?
In ggplot2, the + must come at the end of a line, not the start of the next one. Otherwise R thinks the first line is complete, and the second line starts with a stray +. The message even asks: “Did you accidentally put + on a new line?”
mapping must be created by aes() … Did you use %>% or |> instead of +?Error in `geom_histogram()`:
! `mapping` must be created by `aes()`.
✖ You've supplied a <ggplot2::ggplot> object.
ℹ Did you use `%>%` or `|>` instead of `+`?
Layers of a ggplot are added with +, not with the pipe. The pipe passes data into ggplot(); after that, it is + all the way.
Columns must be mapped inside aes(): geom_point(aes(x = age, y = financial_worry)). Outside aes(), R looks for objects called age and financial_worry and does not find them (Chapter 4 on mapping versus setting).
stat_count() must only have an x or y aestheticError in `geom_bar()`:
! Problem while computing stat.
ℹ Error occurred in the 1st layer.
Caused by error in `setup_params()`:
! `stat_count()` must only have an x or y aesthetic.
geom_bar() counts the rows in each category, so it takes only x. To plot a value you have calculated, such as a mean per faculty, use geom_col().
Error in t.test.formula(age ~ faculty, data = students): grouping factor must have exactly 2 levels
A two-sample t-test compares exactly two groups, but faculty has five. Use ANOVA for more than two groups (Chapter 8), or filter the data to the two groups you want to compare.
The test needs more data than it was given, often because filtering or missing values left too few cases. Check the number of cases with nrow() or sum(!is.na(x)).
Error in `contrasts<-`(`*tmp*`, value = contr.funs[1 + isOF[nn]]): contrasts can be applied only to factors with 2 or more levels
A categorical predictor has only one value in the data used, here because the data was filtered to women only, so there is nothing to compare. Remove the predictor, or check the filtering.
Warning in chisq.test(matrix(c(3, 1, 2, 4), nrow = 2)): Chi-squared
approximation may be incorrect
Pearson's Chi-squared test with Yates' continuity correction
data: matrix(c(3, 1, 2, 4), nrow = 2)
X-squared = 0.41667, df = 1, p-value = 0.5186
Some expected counts are below 5, so the chi-square test’s p-value may be inaccurate (Chapter 7). Use Fisher’s exact test (fisher.test()), or combine small categories.
Warning: glm.fit: algorithm did not converge
Warning: glm.fit: fitted probabilities numerically 0 or 1 occurred
In a logistic regression, a predictor separates the outcomes perfectly (here, everyone above 5 hours passed), so the coefficients become enormous and meaningless, and the fitting algorithm may also report that it did not converge. With real data, this usually means a very small group or a predictor that is almost the outcome itself. Check the cross-table of the outcome and the predictor, and consider simplifying the model.
boundary (singular) fit: see help('isSingular')
A mixed-effects model estimated some random-effect variance as zero, or a correlation as exactly ±1: the random effects part is too complex for the data (Chapter 10). Simplify it, for example by removing the random slope: (1 | supervisor_id).
Not an error, but a line in the output of summary() for lm() and glm(): rows with a missing value in any variable of the model are left out. Check how many with nobs(model), report the number used, and think about whether the missing rows differ from the others (Chapters 6 and 8).
Error in `check_outcome()`:
! For a classification model, the outcome should be a <factor>, not a
character vector.
tidymodels needs a categorical outcome to be a factor. Convert it first, and set the event of interest as the first level (Chapter 11): mutate(considering_dropout = factor(considering_dropout, levels = c("Yes", "No"))).
.pred_YesError in `roc_auc()`:
! Can't select columns that don't exist.
✖ Column `.pred_Yes` doesn't exist.
A column name does not exist in the data given to a yardstick function. The prediction columns created by augment() are named after the outcome’s levels (.pred_Yes, .pred_No), so check the names with names(), and check that the predictions were added to the data.
str(), glimpse(), or summary(): wrong types and unexpected missing values cause most errors in real analyses.