Data is the written result of measurement. Each value records one observation about one case: a student’s age, the faculty she belongs to, how many hours she slept in a semester, her answer to a questionnaire item. Before any analysis, a researcher has to know what each value stands for, what kind of measurement produced it, and how the values are arranged. The same questions arise whether the data comes from a questionnaire, a laboratory instrument, or an administrative register, and they decide which summaries and tests make sense later.
Research data rarely arrives in one neat file. In the wellbeing study, the survey tool produced an Excel export, the university’s records office sent the semester results as CSV files, and a colleague who helped with an earlier pilot study works in SPSS, so part of the data exists as an SPSS file too. Before any of these files can be opened, it helps to know how R holds data once it is inside. R has a small number of ways to hold data, called data structures. This chapter introduces them as the tools for recording measurements: vectors for the values of one variable, factors for categories, and data frames for whole tables of cases. It then shows how to bring files of every common format into R, and how to keep track of what each variable means.
TipBy the end of this chapter you will be able to
Describe data as a table of cases and variables, and identify the unit of analysis.
Create vectors and factors, and pick out and change their values.
Explain what a data frame is, and select rows and columns from one.
Recognise matrices and lists when you meet them.
Import data from CSV, Excel, and SPSS files, keeping SPSS labels, and export data again.
Build a codebook that links each variable to the question behind it.
Map familiar SPSS and Excel tasks to R, and read R’s most common error messages.
Almost all research data can be arranged as a table in which each row is a case and each column a variable. A case is the thing being measured: a student, a patient, a plot of land, a school. A variable is one characteristic measured on every case, such as age or faculty. Each cell then holds one value: the measurement of one variable on one case. This arrangement, sometimes called the data matrix, is the starting point of nearly every method in this book.
What counts as a case depends on the study, and the choice is called the unit of analysis. The wellbeing study has more than one. In the students table, each row is a student. In the semesters table, each row is one student in one semester, so each student appears up to four times:
The first four rows all belong to student S0001, one for each semester. Confusing the two units is a common source of error. A question about students should count each student once; a question about change over time needs all the semester rows, together with methods that know that rows from the same student belong together (Chapter 10). Chapter 5 returns to the unit of analysis when research questions are turned into data.
2.2 Variables and their values
In R, the values of one variable are held in a vector. You met vectors in Chapter 1: a set of values combined with c(). They are the basic building block of R, and almost everything else is built from them.
sleep <-c(6.5, 7, 5.5, 8, 6)sleep
[1] 6.5 7.0 5.5 8.0 6.0
A vector holds values of one type only. If you mix types, R quietly converts everything to the most flexible type, which is usually text:
mixed <-c(6.5, "seven", 5.5)mixed
[1] "6.5" "seven" "5.5"
class(mixed)
[1] "character"
The numbers are now text, in quotes, and they can no longer be averaged. This matters more than it seems. If one person in a survey types “seven” instead of 7, the whole column arrives in R as text. The Excel file later in this chapter shows exactly this.
2.2.1 Operations on a whole vector
Most things you do to a vector happen to every value at once. Converting the sleep values from hours into minutes needs only one multiplication:
sleep *60
[1] 390 420 330 480 360
Comparisons work the same way, and give one TRUE or FALSE for each value. The nights shorter than the recommended 7 hours are marked by:
sleep <7
[1] TRUE FALSE TRUE FALSE TRUE
Because R counts TRUE as 1 and FALSE as 0, sum() counts the short nights, and mean() gives their share:
sum(sleep <7)
[1] 3
mean(sleep <7)
[1] 0.6
Three of the five nights, or 60%, were short. The same line of code would work just as well on 2,000 nights.
2.2.2 Selecting values
Square brackets pick values from a vector by their position. Positions start at 1:
sleep[1] # the first night
[1] 6.5
sleep[c(1, 3)] # the first and third nights
[1] 6.5 5.5
sleep[-2] # every night except the second
[1] 6.5 5.5 8.0 6.0
Values can also be picked with a condition, by putting a TRUE/FALSE vector inside the brackets. R keeps the values where the condition is TRUE:
sleep[sleep <7]
[1] 6.5 5.5 6.0
The line reads aloud as “sleep, where sleep is less than 7”. This way of selecting, called logical indexing, is one of the most useful ideas in R, and you will use it constantly.
2.2.3 Changing values
Brackets on the left of the arrow change values. Suppose the second night turns out to have been 7.5 hours, not 7:
sleep[2] <-7.5sleep
[1] 6.5 7.5 5.5 8.0 6.0
2.3 Categories
Research data is full of categories: faculty, gender, treatment group, agreement on a scale. R stores categories as factors. A factor looks like text but knows the complete set of possible values, called its levels:
[1] Education Humanities Education Health Sciences
Levels: Education Health Sciences Humanities
levels(faculty)
[1] "Education" "Health Sciences" "Humanities"
table(faculty)
faculty
Education Health Sciences Humanities
2 1 1
By default, levels are in alphabetical order. When the categories have a natural order, set it yourself with the levels argument, so that tables and graphs show them in that order:
Factors matter most in statistics and graphs. When groups are compared in Chapter 7, or plotted in Chapter 4, R uses the factor’s levels to decide which groups exist and in which order to show them.
2.3.1 Kinds of measurement and R’s types
The choice between a factor and a number is not a matter of convenience; it follows from the kind of measurement that produced the values. Categories without an order, such as faculty, are nominal. Categories with an order but uneven or unknown distances between them, such as none, part-time, and full-time employment, are ordinal. Measurements on a scale with equal distances, such as hours of sleep or a wellbeing index, are numeric. Table 2.1 shows how each kind is stored in R. Chapter 5 explains these levels of measurement in full, and why they decide which summaries are meaningful.
Table 2.1: Kinds of measurement and how R stores them
Kind of measurement
Example
Stored in R as
Categories, no order (nominal)
faculty, gender
factor
Ordered categories (ordinal)
employment, agreement on a scale
factor with levels in order, or an ordered factor
Quantities with equal distances
sleep hours, age, wellbeing
numeric
Yes or no
invited to the workshop
logical, or a factor with two levels
2.4 A table of cases
Most research data is a table of this kind: one row for each case, one column for each variable. In R, such a table is a data frame. Each column is a vector, so each column has one type, but different columns can have different types.
A small data frame can be built with data.frame():
The last line combines a data frame with logical indexing: “pilot, the rows where sleep is under 7, all columns”. It is the R version of SPSS’s Select Cases.
In this listing, chr means character, int whole numbers, and num numbers with decimals. The categories, such as faculty, arrive as text, and they are converted into factors when they are needed as categories:
Two symbols are new here. The double equals sign, ==, means “is equal to”; a single = is used for arguments, so comparisons need two. The ampersand, &, means “and”: both conditions must be TRUE. Its partner | means “or”.
TipLooking at data like a spreadsheet
In RStudio, View(students) opens the data in a spreadsheet-style viewer, where you can scroll, sort, and filter. It is only for looking: changes you want to keep should be made with code, so they are recorded.
2.5 Matrices and lists
Two more structures appear regularly, even though they are rarely built by hand.
2.5.1 Matrices
A matrix is a grid of values of one type, usually numbers. Matrices appear most often as results. The function cor(), for example, calculates the correlations between several variables and returns them as a matrix. The correlations below use the first semester only, so that each student counts once:
Every variable correlates perfectly with itself, hence the 1s on the diagonal. Chapter 6 explains how to read correlations. For now, notice that values are picked from a matrix exactly as from a data frame, with [row, column]:
correlations["sleep_hours", "wellbeing"]
[1] 0.4602433
2.5.2 Lists
A list is a container that can hold anything: vectors of different lengths, data frames, even other lists. Its parts are usually named:
Lists are rarely built by hand, but R’s statistical functions return their results as lists. A preview of a test from Chapter 7, which examines whether students sleep 7 hours on average in their first semester, shows this:
result <-t.test(first_semester$sleep_hours, mu =7)names(result)
The function names() lists everything the test calculated, and $ picks out one part, here the students’ average sleep. This is how a number from a statistical test is retrieved for a report.
2.6 Importing data
Each file format has its own function for reading it, and every one of them gives back a data frame. The files used below come with the data2thesis package, and data2thesis_example() finds them on your computer. For your own data, you would put the file in your RStudio Project and give its name instead (see Chapter 1).
2.6.1 CSV files
A CSV file (comma-separated values) is a plain text table, where commas separate the columns. It is the most common format for sharing data, and R reads it with read.csv():
This is the export from the study’s survey tool, and it shows what real survey exports look like. The function dim() gives its size: 615 rows (more than the 600 students) and 64 columns. The first rows are:
head(raw[, 1:6])
# A tibble: 6 × 6
`Response ID` `Q1_Student ID` Q2_Age Q3_Gender Q4_Faculty Q5_Programme
<chr> <chr> <chr> <chr> <chr> <chr>
1 R0001 TEST 99 Female <NA> <NA>
2 R0002 test <NA> <NA> <NA> <NA>
3 R0003 TEST2 30 Male <NA> <NA>
4 R0004 S0191 27 male Humanities Masters
5 R0005 S0015 30 Male Health Sciences Master's
6 R0006 S0474 38 F health sciences Masters
The printout looks a little different from before, because read_excel() returns a tibble: a modern kind of data frame that prints more compactly and shows each column’s type under its name. It can be used exactly like a data frame.
Several problems are already visible. There are test responses (TEST), odd codes such as 99 for missing age, and the same answer spelled several ways (Male, male). Every column shown is <chr>, text, even the ages, because a few entries are not plain numbers. Chapter 3 is devoted to cleaning this file; for now, it is enough to be able to open it.
If a workbook has several sheets, choose one with the sheet argument, for example read_excel("file.xlsx", sheet = "Semester 2").
2.6.3 SPSS files
The haven package reads SPSS files (.sav), and also Stata and SAS files:
SPSS files carry labels, and haven keeps them. Each value is stored as a number, with its label shown in square brackets: 2 [Health Sciences]. The question behind each variable is kept too:
attr(spss$stress_1, "label")
[1] "I feel unable to control important things in my studies."
To analyse a labelled variable as categories, turn it into a factor with as_factor(), which uses the labels as levels:
table(as_factor(spss$faculty))
Education Health Sciences Humanities Natural Sciences
148 154 95 87
Social Sciences
116
This is one of the great conveniences for anyone moving from SPSS: the work put into labelling the data is not lost.
2.7 A codebook
A value such as 4 in the column stress_1 means nothing on its own. It becomes a measurement only when the reader knows the question behind it (“I feel unable to control important things in my studies”), the scale of the answers (1 = strongly disagree, 5 = strongly agree), and how missing answers are recorded. A codebook is the document that records this for every variable: its name, its meaning, its type, its possible values, and how it was measured. It is the bridge between the questionnaire that the participants saw and the table that the researcher analyses, and most theses include one in an appendix.
The labels in the SPSS file already contain much of a codebook, and R can collect them into a table. The code below takes the label of every variable, and the type in which it is stored:
variable label type
1 student_id Student identifier character
2 supervisor_id Supervisor identifier character
3 age Age in years at baseline numeric
4 gender Gender haven_labelled
5 faculty Faculty haven_labelled
6 programme Degree programme haven_labelled
7 study_mode Study mode haven_labelled
8 employment Paid work alongside study haven_labelled
9 has_children Has children haven_labelled
10 lives_away Moved away from family to study haven_labelled
11 financial_worry How worried are you about money? haven_labelled
12 workshop Randomly invited to the wellbeing workshop haven_labelled
The function sapply() applies a function to every column in turn and collects the results. A labelled categorical variable also records which number stands for which category, its value labels:
attr(spss$faculty, "labels")
Education Health Sciences Humanities Natural Sciences
1 2 3 4
Social Sciences
5
A codebook built in this way can be written to a file with the functions in the next section and completed by hand, adding the answer scales and any notes on how variables were coded. Keeping it next to the data means that anyone who opens the data, including the researcher a year later, can tell what every value means.
2.8 Exporting data
A data frame is saved with the matching write function. Each of these creates a file in the project folder:
The argument row.names = FALSE stops R from adding an extra column of row numbers to the CSV file.
Original data files should stay untouched. Cleaned or changed versions are saved under new names, and the script records how one became the other.
2.9 SPSS and Excel equivalents
Most tasks done in SPSS or Excel have a direct equivalent in R. Table 2.2 collects the ones from this chapter and the last.
Table 2.2: SPSS and Excel tasks and their R equivalents
In SPSS or Excel
In R
Open a data file
read_sav(), read_excel(), read.csv()
Variable View: names, types, labels
str(); attr(x, "label") for a variable label
Data View
View()
Select Cases
data[condition, ]
Compute Variable
data$new <- ...
Frequencies
table()
Descriptives
mean(), summary()
Save As
write_sav(), write_xlsx(), write.csv()
The biggest change is not a particular command. In SPSS or Excel, you change the data and the change is saved; the steps you took are forgotten. In R, the data file stays as it was, and the steps are saved in your script. Run the script again, and you get the same result again.
2.10 Error messages
Everyone who writes R sees error messages every day. They are not a sign of failure; they are R reporting precisely what it could not do. The message usually says what went wrong, and often where, so it should always be read. Four messages account for most of the errors a beginner meets.
Object not found. A name is misspelled, or the object was never created (perhaps the line that creates it was not run):
mean(slep)
Error: object 'slep' not found
Could not find function. A function name is misspelled, or its package has not been loaded with library():
reed_csv("students.csv")
Error in reed_csv("students.csv"): could not find function "reed_csv"
Non-numeric argument. A calculation was attempted on text, often because a column arrived as text:
mixed *2
Error in mixed * 2: non-numeric argument to binary operator
Undefined columns selected. A column name in brackets does not exist, often because of a typo or wrong capitals:
students[, "Faculty"]
Error in `[.data.frame`(students, , "Faculty"): undefined columns selected
When the problem is not obvious, check spelling and capitals first, then check that every line above was run, then read the help page of the function. Copying the error message into a search engine works surprisingly often, because someone has almost always met it before.
2.11 A first summary
With the data open, summary() gives a quick overview of every column: averages and ranges for numbers, counts for factors, and the number of missing values. Here it is for the first semester, one row per student:
gpa sleep_hours study_hours wellbeing
Min. :2.170 Min. : 3.500 Min. : 2.00 Min. :24.00
1st Qu.:2.888 1st Qu.: 5.800 1st Qu.:18.00 1st Qu.:52.00
Median :3.100 Median : 6.500 Median :26.00 Median :60.00
Mean :3.102 Mean : 6.484 Mean :27.64 Mean :60.46
3rd Qu.:3.340 3rd Qu.: 7.200 3rd Qu.:37.00 3rd Qu.:69.00
Max. :4.000 Max. :10.000 Max. :67.00 Max. :93.00
NA's :6 NA's :8
Already, the summary shows that the average student sleeps less than 7 hours, and that a few values are missing (NA's). Chapter 6 turns this quick look into a full description of the data.
NoteIn your field: health research
ToothGrowth is a real dataset that comes with R, from an experiment on the effect of vitamin C on tooth growth in 60 guinea pigs. Its structure is typical of experiments in health research: an outcome, a treatment group, and a dose.
The variable supp is already a factor with two levels, the two ways the vitamin was given (orange juice, OJ, or ascorbic acid, VC), and the table shows 10 animals for every combination of method and dose. The unit of analysis is the animal: each row is one guinea pig.
2.12 Chapter review
2.12.1 Summary
Data is the record of measurements: a table of cases (rows) and variables (columns). The unit of analysis is what one row represents, and the same study can have more than one.
A vector holds the values of one variable, all of one type. Mixing types turns everything into text.
Most operations work on every value of a vector at once. [ ] picks values by position or by a condition (logical indexing).
A factor holds categories with a fixed set of levels, whose order you can set. The kind of measurement decides whether a variable should be a factor or a number.
A data frame is a table: one row per case, one column (a vector) per variable. Select with $ or [rows, columns], and use str() to see its structure.
A matrix is a grid of one type, often a result such as a correlation matrix. A list can hold anything; statistical results are lists.
read.csv(), read_excel() (readxl), and read_sav() (haven) import data. haven keeps SPSS labels, and as_factor() turns them into factors. write.csv(), write_xlsx(), and write_sav() export data.
A codebook records what every variable means and how it was measured; SPSS labels provide much of it.
Error messages say what R could not do. Read them, and check spelling and capitals first.
2.12.2 Key terms
Case, variable, data matrix, unit of analysis, data structure, vector, type conversion, logical indexing, factor, level, data frame, tibble, matrix, list, CSV file, variable label, value label, codebook, error message.
2.13 Exercises
The playground has these and more, with hints and solutions, in your browser or as an RStudio project to download.
Create a vector with the ages 24, 31, 28, 45, and 26. Use logical indexing to show only the ages over 30, and count how many there are.
Create a factor from c("Agree", "Disagree", "Neutral", "Agree") whose levels run from "Disagree" through "Neutral" to "Agree". Check the order with table().
Using students, count the students who are part-time and have children. (Use &.)
Add a column to students called over_30 that is TRUE for students older than 30. Count those students, and explain why the count might need na.rm = TRUE.
Read the SPSS file with read_sav() and find the question behind support_3.
State the unit of analysis of each table in the package: students, semesters, questionnaire, and supervisors. For each, name one research question it could answer on its own.
2.14 Further reading
R for Data Science(Wickham et al. 2023): chapters “Data import”, “Spreadsheets”, and “Factors” go further with importing and with categories, and “A field guide to base R” covers the square-bracket style used in this chapter.
The documentation of the readxl and haven packages.
References
Wickham, Hadley, Mine Çetinkaya-Rundel, and Garrett Grolemund. 2023. R for Data Science: Import, Tidy, Transform, Visualize, and Model Data. 2nd ed. O’Reilly Media. https://r4ds.hadley.nz.