2 Data Structures in R

Data is the written result of measurement. Each value records one observation about one case: a student’s age, the faculty she belongs to, how many hours she slept in a semester, her answer to a questionnaire item. Before any analysis, a researcher has to know what each value stands for, what kind of measurement produced it, and how the values are arranged. The same questions arise whether the data comes from a questionnaire, a laboratory instrument, or an administrative register, and they decide which summaries and tests make sense later.

Research data rarely arrives in one neat file. In the wellbeing study, the survey tool produced an Excel export, the university’s records office sent the semester results as CSV files, and a colleague who helped with an earlier pilot study works in SPSS, so part of the data exists as an SPSS file too. Before any of these files can be opened, it helps to know how R holds data once it is inside. R has a small number of ways to hold data, called data structures. This chapter introduces them as the tools for recording measurements: vectors for the values of one variable, factors for categories, and data frames for whole tables of cases. It then shows how to bring files of every common format into R, and how to keep track of what each variable means.

TipBy the end of this chapter you will be able to
  • Describe data as a table of cases and variables, and identify the unit of analysis.
  • Create vectors and factors, and pick out and change their values.
  • Explain what a data frame is, and select rows and columns from one.
  • Recognise matrices and lists when you meet them.
  • Import data from CSV, Excel, and SPSS files, keeping SPSS labels, and export data again.
  • Build a codebook that links each variable to the question behind it.
  • Map familiar SPSS and Excel tasks to R, and read R’s most common error messages.

2.1 Data as measurement

Almost all research data can be arranged as a table in which each row is a case and each column a variable. A case is the thing being measured: a student, a patient, a plot of land, a school. A variable is one characteristic measured on every case, such as age or faculty. Each cell then holds one value: the measurement of one variable on one case. This arrangement, sometimes called the data matrix, is the starting point of nearly every method in this book.

What counts as a case depends on the study, and the choice is called the unit of analysis. The wellbeing study has more than one. In the students table, each row is a student. In the semesters table, each row is one student in one semester, so each student appears up to four times:

head(semesters, 5)
  student_id semester  gpa sleep_hours study_hours exercise_days caffeine_mg
1      S0001        1 3.02         7.2          26             5         130
2      S0001        2 3.27         8.4          17             3          80
3      S0001        3 3.17         7.8          25             7         105
4      S0001        4 3.00         7.2          28             4         145
5      S0002        1 3.25         6.4          32             2         210
  supervisor_meetings wellbeing
1                   0        74
2                   1        69
3                   3        71
4                   2        74
5                   3        60

The first four rows all belong to student S0001, one for each semester. Confusing the two units is a common source of error. A question about students should count each student once; a question about change over time needs all the semester rows, together with methods that know that rows from the same student belong together (Chapter 10). Chapter 5 returns to the unit of analysis when research questions are turned into data.

2.2 Variables and their values

In R, the values of one variable are held in a vector. You met vectors in Chapter 1: a set of values combined with c(). They are the basic building block of R, and almost everything else is built from them.

sleep <- c(6.5, 7, 5.5, 8, 6)
sleep
[1] 6.5 7.0 5.5 8.0 6.0

A vector holds values of one type only. If you mix types, R quietly converts everything to the most flexible type, which is usually text:

mixed <- c(6.5, "seven", 5.5)
mixed
[1] "6.5"   "seven" "5.5"  
class(mixed)
[1] "character"

The numbers are now text, in quotes, and they can no longer be averaged. This matters more than it seems. If one person in a survey types “seven” instead of 7, the whole column arrives in R as text. The Excel file later in this chapter shows exactly this.

2.2.1 Operations on a whole vector

Most things you do to a vector happen to every value at once. Converting the sleep values from hours into minutes needs only one multiplication:

sleep * 60
[1] 390 420 330 480 360

Comparisons work the same way, and give one TRUE or FALSE for each value. The nights shorter than the recommended 7 hours are marked by:

sleep < 7
[1]  TRUE FALSE  TRUE FALSE  TRUE

Because R counts TRUE as 1 and FALSE as 0, sum() counts the short nights, and mean() gives their share:

sum(sleep < 7)
[1] 3
mean(sleep < 7)
[1] 0.6

Three of the five nights, or 60%, were short. The same line of code would work just as well on 2,000 nights.

2.2.2 Selecting values

Square brackets pick values from a vector by their position. Positions start at 1:

sleep[1]         # the first night
[1] 6.5
sleep[c(1, 3)]   # the first and third nights
[1] 6.5 5.5
sleep[-2]        # every night except the second
[1] 6.5 5.5 8.0 6.0

Values can also be picked with a condition, by putting a TRUE/FALSE vector inside the brackets. R keeps the values where the condition is TRUE:

sleep[sleep < 7]
[1] 6.5 5.5 6.0

The line reads aloud as “sleep, where sleep is less than 7”. This way of selecting, called logical indexing, is one of the most useful ideas in R, and you will use it constantly.

2.2.3 Changing values

Brackets on the left of the arrow change values. Suppose the second night turns out to have been 7.5 hours, not 7:

sleep[2] <- 7.5
sleep
[1] 6.5 7.5 5.5 8.0 6.0

2.3 Categories

Research data is full of categories: faculty, gender, treatment group, agreement on a scale. R stores categories as factors. A factor looks like text but knows the complete set of possible values, called its levels:

faculty <- factor(c("Education", "Humanities", "Education", "Health Sciences"))
faculty
[1] Education       Humanities      Education       Health Sciences
Levels: Education Health Sciences Humanities
levels(faculty)
[1] "Education"       "Health Sciences" "Humanities"     
table(faculty)
faculty
      Education Health Sciences      Humanities 
              2               1               1 

By default, levels are in alphabetical order. When the categories have a natural order, set it yourself with the levels argument, so that tables and graphs show them in that order:

employment <- factor(
  c("None", "Full-time job", "Part-time job", "None"),
  levels = c("None", "Part-time job", "Full-time job")
)
table(employment)
employment
         None Part-time job Full-time job 
            2             1             1 

Factors matter most in statistics and graphs. When groups are compared in Chapter 7, or plotted in Chapter 4, R uses the factor’s levels to decide which groups exist and in which order to show them.

2.3.1 Kinds of measurement and R’s types

The choice between a factor and a number is not a matter of convenience; it follows from the kind of measurement that produced the values. Categories without an order, such as faculty, are nominal. Categories with an order but uneven or unknown distances between them, such as none, part-time, and full-time employment, are ordinal. Measurements on a scale with equal distances, such as hours of sleep or a wellbeing index, are numeric. Table 2.1 shows how each kind is stored in R. Chapter 5 explains these levels of measurement in full, and why they decide which summaries are meaningful.

Table 2.1: Kinds of measurement and how R stores them
Kind of measurement Example Stored in R as
Categories, no order (nominal) faculty, gender factor
Ordered categories (ordinal) employment, agreement on a scale factor with levels in order, or an ordered factor
Quantities with equal distances sleep hours, age, wellbeing numeric
Yes or no invited to the workshop logical, or a factor with two levels

2.4 A table of cases

Most research data is a table of this kind: one row for each case, one column for each variable. In R, such a table is a data frame. Each column is a vector, so each column has one type, but different columns can have different types.

A small data frame can be built with data.frame():

pilot <- data.frame(
  student = c("S1", "S2", "S3", "S4"),
  faculty = c("Education", "Humanities", "Education", "Health Sciences"),
  sleep   = c(6.5, 7.5, 5.5, 8),
  invited = c(TRUE, FALSE, TRUE, FALSE)
)
pilot
  student         faculty sleep invited
1      S1       Education   6.5    TRUE
2      S2      Humanities   7.5   FALSE
3      S3       Education   5.5    TRUE
4      S4 Health Sciences   8.0   FALSE

The quickest way to understand a data frame is str(), short for structure. It lists every column with its type and first few values:

str(pilot)
'data.frame':   4 obs. of  4 variables:
 $ student: chr  "S1" "S2" "S3" "S4"
 $ faculty: chr  "Education" "Humanities" "Education" "Health Sciences"
 $ sleep  : num  6.5 7.5 5.5 8
 $ invited: logi  TRUE FALSE TRUE FALSE

2.4.1 Selecting columns and rows

A single column is picked with $, as in Chapter 1:

pilot$sleep
[1] 6.5 7.5 5.5 8.0

Square brackets work on data frames too, with two positions separated by a comma: rows first, then columns. An empty position means “all”:

pilot[2, ]                        # the second row, all columns
  student    faculty sleep invited
2      S2 Humanities   7.5   FALSE
pilot[, c("student", "sleep")]    # all rows, two columns
  student sleep
1      S1   6.5
2      S2   7.5
3      S3   5.5
4      S4   8.0
pilot[pilot$sleep < 7, ]          # rows where sleep is under 7
  student   faculty sleep invited
1      S1 Education   6.5    TRUE
3      S3 Education   5.5    TRUE

The last line combines a data frame with logical indexing: “pilot, the rows where sleep is under 7, all columns”. It is the R version of SPSS’s Select Cases.

2.4.2 Adding columns

Assigning to a new column name adds a column:

pilot$short_sleep <- pilot$sleep < 7
pilot
  student         faculty sleep invited short_sleep
1      S1       Education   6.5    TRUE        TRUE
2      S2      Humanities   7.5   FALSE       FALSE
3      S3       Education   5.5    TRUE        TRUE
4      S4 Health Sciences   8.0   FALSE       FALSE

2.4.3 The student wellbeing data

The students data frame from the data2thesis package has the same structure as pilot, only larger:

str(students)
'data.frame':   600 obs. of  14 variables:
 $ student_id         : chr  "S0001" "S0002" "S0003" "S0004" ...
 $ supervisor_id      : chr  "SUP088" "SUP116" "SUP040" "SUP112" ...
 $ age                : int  31 28 33 34 35 36 NA 25 29 26 ...
 $ gender             : chr  "Female" "Female" "Male" "Female" ...
 $ faculty            : chr  "Health Sciences" "Education" "Education" "Health Sciences" ...
 $ programme          : chr  "Master's" "PhD" "PhD" "Master's" ...
 $ study_mode         : chr  "Full-time" "Full-time" "Full-time" "Part-time" ...
 $ employment         : chr  "None" "Part-time job" "None" "Full-time job" ...
 $ has_children       : chr  "No" "No" "No" "No" ...
 $ lives_away         : chr  "Yes" "Yes" "No" "No" ...
 $ financial_worry    : int  2 2 4 2 2 5 3 NA 2 2 ...
 $ workshop           : chr  "Not invited" "Invited" "Invited" "Invited" ...
 $ workshop_sessions  : int  0 0 5 6 2 0 4 0 0 0 ...
 $ considering_dropout: chr  "No" "No" "No" "No" ...

In this listing, chr means character, int whole numbers, and num numbers with decimals. The categories, such as faculty, arrive as text, and they are converted into factors when they are needed as categories:

students$faculty <- factor(students$faculty)
levels(students$faculty)
[1] "Education"        "Health Sciences"  "Humanities"       "Natural Sciences"
[5] "Social Sciences" 

Selection works as before. The number of PhD students in Health Sciences, for example, is found with two conditions:

hs_phd <- students[students$faculty == "Health Sciences" & students$programme == "PhD", ]
nrow(hs_phd)
[1] 44

Two symbols are new here. The double equals sign, ==, means “is equal to”; a single = is used for arguments, so comparisons need two. The ampersand, &, means “and”: both conditions must be TRUE. Its partner | means “or”.

TipLooking at data like a spreadsheet

In RStudio, View(students) opens the data in a spreadsheet-style viewer, where you can scroll, sort, and filter. It is only for looking: changes you want to keep should be made with code, so they are recorded.

2.5 Matrices and lists

Two more structures appear regularly, even though they are rarely built by hand.

2.5.1 Matrices

A matrix is a grid of values of one type, usually numbers. Matrices appear most often as results. The function cor(), for example, calculates the correlations between several variables and returns them as a matrix. The correlations below use the first semester only, so that each student counts once:

first_semester <- semesters[semesters$semester == 1, ]
vars <- first_semester[, c("gpa", "sleep_hours", "study_hours", "wellbeing")]
correlations <- cor(vars, use = "complete.obs")
round(correlations, 2)
             gpa sleep_hours study_hours wellbeing
gpa         1.00        0.25        0.07      0.31
sleep_hours 0.25        1.00       -0.56      0.46
study_hours 0.07       -0.56        1.00     -0.24
wellbeing   0.31        0.46       -0.24      1.00

Every variable correlates perfectly with itself, hence the 1s on the diagonal. Chapter 6 explains how to read correlations. For now, notice that values are picked from a matrix exactly as from a data frame, with [row, column]:

correlations["sleep_hours", "wellbeing"]
[1] 0.4602433

2.5.2 Lists

A list is a container that can hold anything: vectors of different lengths, data frames, even other lists. Its parts are usually named:

student <- list(
  id        = "S0001",
  programme = "Master's",
  sleep     = c(7.2, 8.4, 7.8, 7.2)
)
student$sleep
[1] 7.2 8.4 7.8 7.2

Lists are rarely built by hand, but R’s statistical functions return their results as lists. A preview of a test from Chapter 7, which examines whether students sleep 7 hours on average in their first semester, shows this:

result <- t.test(first_semester$sleep_hours, mu = 7)
names(result)
 [1] "statistic"   "parameter"   "p.value"     "conf.int"    "estimate"   
 [6] "null.value"  "stderr"      "alternative" "method"      "data.name"  
result$estimate
mean of x 
 6.483502 

The function names() lists everything the test calculated, and $ picks out one part, here the students’ average sleep. This is how a number from a statistical test is retrieved for a report.

2.6 Importing data

Each file format has its own function for reading it, and every one of them gives back a data frame. The files used below come with the data2thesis package, and data2thesis_example() finds them on your computer. For your own data, you would put the file in your RStudio Project and give its name instead (see Chapter 1).

2.6.1 CSV files

A CSV file (comma-separated values) is a plain text table, where commas separate the columns. It is the most common format for sharing data, and R reads it with read.csv():

path <- data2thesis_example("semesters.csv")
semester_file <- read.csv(path)
head(semester_file, 3)
  student_id semester  gpa sleep_hours study_hours exercise_days caffeine_mg
1      S0001        1 3.02         7.2          26             5         130
2      S0001        2 3.27         8.4          17             3          80
3      S0001        3 3.17         7.8          25             7         105
  supervisor_meetings wellbeing
1                   0        74
2                   1        69
3                   3        71

In your own project, this would simply be read.csv("semesters.csv"), or read.csv(here::here("data", "semesters.csv")) with a data folder.

2.6.2 Excel files

The readxl package reads Excel files. Install it once with install.packages("readxl"):

library(readxl)
raw <- read_excel(data2thesis_example("wellbeing_raw.xlsx"))
dim(raw)
[1] 615  64

This is the export from the study’s survey tool, and it shows what real survey exports look like. The function dim() gives its size: 615 rows (more than the 600 students) and 64 columns. The first rows are:

head(raw[, 1:6])
# A tibble: 6 × 6
  `Response ID` `Q1_Student ID` Q2_Age Q3_Gender Q4_Faculty      Q5_Programme
  <chr>         <chr>           <chr>  <chr>     <chr>           <chr>       
1 R0001         TEST            99     Female    <NA>            <NA>        
2 R0002         test            <NA>   <NA>      <NA>            <NA>        
3 R0003         TEST2           30     Male      <NA>            <NA>        
4 R0004         S0191           27     male      Humanities      Masters     
5 R0005         S0015           30     Male      Health Sciences Master's    
6 R0006         S0474           38     F         health sciences Masters     

The printout looks a little different from before, because read_excel() returns a tibble: a modern kind of data frame that prints more compactly and shows each column’s type under its name. It can be used exactly like a data frame.

Several problems are already visible. There are test responses (TEST), odd codes such as 99 for missing age, and the same answer spelled several ways (Male, male). Every column shown is <chr>, text, even the ages, because a few entries are not plain numbers. Chapter 3 is devoted to cleaning this file; for now, it is enough to be able to open it.

If a workbook has several sheets, choose one with the sheet argument, for example read_excel("file.xlsx", sheet = "Semester 2").

2.6.3 SPSS files

The haven package reads SPSS files (.sav), and also Stata and SAS files:

library(haven)
spss <- read_sav(data2thesis_example("wellbeing.sav"))
spss[1:3, c("student_id", "gender", "faculty", "stress_1")]
# A tibble: 3 × 4
  student_id gender     faculty             stress_1   
  <chr>      <dbl+lbl>  <dbl+lbl>           <dbl+lbl>  
1 S0001      1 [Female] 2 [Health Sciences] 3 [Neutral]
2 S0002      1 [Female] 1 [Education]       4 [Agree]  
3 S0003      2 [Male]   1 [Education]       4 [Agree]  

SPSS files carry labels, and haven keeps them. Each value is stored as a number, with its label shown in square brackets: 2 [Health Sciences]. The question behind each variable is kept too:

attr(spss$stress_1, "label")
[1] "I feel unable to control important things in my studies."

To analyse a labelled variable as categories, turn it into a factor with as_factor(), which uses the labels as levels:

table(as_factor(spss$faculty))

       Education  Health Sciences       Humanities Natural Sciences 
             148              154               95               87 
 Social Sciences 
             116 

This is one of the great conveniences for anyone moving from SPSS: the work put into labelling the data is not lost.

2.7 A codebook

A value such as 4 in the column stress_1 means nothing on its own. It becomes a measurement only when the reader knows the question behind it (“I feel unable to control important things in my studies”), the scale of the answers (1 = strongly disagree, 5 = strongly agree), and how missing answers are recorded. A codebook is the document that records this for every variable: its name, its meaning, its type, its possible values, and how it was measured. It is the bridge between the questionnaire that the participants saw and the table that the researcher analyses, and most theses include one in an appendix.

The labels in the SPSS file already contain much of a codebook, and R can collect them into a table. The code below takes the label of every variable, and the type in which it is stored:

codebook <- data.frame(
  variable = names(spss),
  label    = sapply(spss, function(x) attr(x, "label")),
  type     = sapply(spss, function(x) class(x)[1]),
  row.names = NULL
)
head(codebook, 12)
          variable                                      label           type
1       student_id                         Student identifier      character
2    supervisor_id                      Supervisor identifier      character
3              age                   Age in years at baseline        numeric
4           gender                                     Gender haven_labelled
5          faculty                                    Faculty haven_labelled
6        programme                           Degree programme haven_labelled
7       study_mode                                 Study mode haven_labelled
8       employment                  Paid work alongside study haven_labelled
9     has_children                               Has children haven_labelled
10      lives_away            Moved away from family to study haven_labelled
11 financial_worry           How worried are you about money? haven_labelled
12        workshop Randomly invited to the wellbeing workshop haven_labelled

The function sapply() applies a function to every column in turn and collects the results. A labelled categorical variable also records which number stands for which category, its value labels:

attr(spss$faculty, "labels")
       Education  Health Sciences       Humanities Natural Sciences 
               1                2                3                4 
 Social Sciences 
               5 

A codebook built in this way can be written to a file with the functions in the next section and completed by hand, adding the answer scales and any notes on how variables were coded. Keeping it next to the data means that anyone who opens the data, including the researcher a year later, can tell what every value means.

2.8 Exporting data

A data frame is saved with the matching write function. Each of these creates a file in the project folder:

write.csv(students, "students_clean.csv", row.names = FALSE)   # CSV
writexl::write_xlsx(students, "students_clean.xlsx")           # Excel (writexl package)
haven::write_sav(spss, "students_clean.sav")                   # SPSS

The argument row.names = FALSE stops R from adding an extra column of row numbers to the CSV file.

Original data files should stay untouched. Cleaned or changed versions are saved under new names, and the script records how one became the other.

2.9 SPSS and Excel equivalents

Most tasks done in SPSS or Excel have a direct equivalent in R. Table 2.2 collects the ones from this chapter and the last.

Table 2.2: SPSS and Excel tasks and their R equivalents
In SPSS or Excel In R
Open a data file read_sav(), read_excel(), read.csv()
Variable View: names, types, labels str(); attr(x, "label") for a variable label
Data View View()
Select Cases data[condition, ]
Compute Variable data$new <- ...
Frequencies table()
Descriptives mean(), summary()
Save As write_sav(), write_xlsx(), write.csv()

The biggest change is not a particular command. In SPSS or Excel, you change the data and the change is saved; the steps you took are forgotten. In R, the data file stays as it was, and the steps are saved in your script. Run the script again, and you get the same result again.

2.10 Error messages

Everyone who writes R sees error messages every day. They are not a sign of failure; they are R reporting precisely what it could not do. The message usually says what went wrong, and often where, so it should always be read. Four messages account for most of the errors a beginner meets.

Object not found. A name is misspelled, or the object was never created (perhaps the line that creates it was not run):

mean(slep)
Error: object 'slep' not found

Could not find function. A function name is misspelled, or its package has not been loaded with library():

reed_csv("students.csv")
Error in reed_csv("students.csv"): could not find function "reed_csv"

Non-numeric argument. A calculation was attempted on text, often because a column arrived as text:

mixed * 2
Error in mixed * 2: non-numeric argument to binary operator

Undefined columns selected. A column name in brackets does not exist, often because of a typo or wrong capitals:

students[, "Faculty"]
Error in `[.data.frame`(students, , "Faculty"): undefined columns selected

When the problem is not obvious, check spelling and capitals first, then check that every line above was run, then read the help page of the function. Copying the error message into a search engine works surprisingly often, because someone has almost always met it before.

2.11 A first summary

With the data open, summary() gives a quick overview of every column: averages and ranges for numbers, counts for factors, and the number of missing values. Here it is for the first semester, one row per student:

summary(first_semester[, c("gpa", "sleep_hours", "study_hours", "wellbeing")])
      gpa         sleep_hours      study_hours      wellbeing    
 Min.   :2.170   Min.   : 3.500   Min.   : 2.00   Min.   :24.00  
 1st Qu.:2.888   1st Qu.: 5.800   1st Qu.:18.00   1st Qu.:52.00  
 Median :3.100   Median : 6.500   Median :26.00   Median :60.00  
 Mean   :3.102   Mean   : 6.484   Mean   :27.64   Mean   :60.46  
 3rd Qu.:3.340   3rd Qu.: 7.200   3rd Qu.:37.00   3rd Qu.:69.00  
 Max.   :4.000   Max.   :10.000   Max.   :67.00   Max.   :93.00  
                 NA's   :6        NA's   :8                      

Already, the summary shows that the average student sleeps less than 7 hours, and that a few values are missing (NA's). Chapter 6 turns this quick look into a full description of the data.

NoteIn your field: health research

ToothGrowth is a real dataset that comes with R, from an experiment on the effect of vitamin C on tooth growth in 60 guinea pigs. Its structure is typical of experiments in health research: an outcome, a treatment group, and a dose.

str(ToothGrowth)
'data.frame':   60 obs. of  3 variables:
 $ len : num  4.2 11.5 7.3 5.8 6.4 10 11.2 11.2 5.2 7 ...
 $ supp: Factor w/ 2 levels "OJ","VC": 2 2 2 2 2 2 2 2 2 2 ...
 $ dose: num  0.5 0.5 0.5 0.5 0.5 0.5 0.5 0.5 0.5 0.5 ...
table(ToothGrowth$supp, ToothGrowth$dose)
    
     0.5  1  2
  OJ  10 10 10
  VC  10 10 10

The variable supp is already a factor with two levels, the two ways the vitamin was given (orange juice, OJ, or ascorbic acid, VC), and the table shows 10 animals for every combination of method and dose. The unit of analysis is the animal: each row is one guinea pig.

2.12 Chapter review

2.12.1 Summary

  • Data is the record of measurements: a table of cases (rows) and variables (columns). The unit of analysis is what one row represents, and the same study can have more than one.
  • A vector holds the values of one variable, all of one type. Mixing types turns everything into text.
  • Most operations work on every value of a vector at once. [ ] picks values by position or by a condition (logical indexing).
  • A factor holds categories with a fixed set of levels, whose order you can set. The kind of measurement decides whether a variable should be a factor or a number.
  • A data frame is a table: one row per case, one column (a vector) per variable. Select with $ or [rows, columns], and use str() to see its structure.
  • A matrix is a grid of one type, often a result such as a correlation matrix. A list can hold anything; statistical results are lists.
  • read.csv(), read_excel() (readxl), and read_sav() (haven) import data. haven keeps SPSS labels, and as_factor() turns them into factors. write.csv(), write_xlsx(), and write_sav() export data.
  • A codebook records what every variable means and how it was measured; SPSS labels provide much of it.
  • Error messages say what R could not do. Read them, and check spelling and capitals first.

2.12.2 Key terms

Case, variable, data matrix, unit of analysis, data structure, vector, type conversion, logical indexing, factor, level, data frame, tibble, matrix, list, CSV file, variable label, value label, codebook, error message.

2.13 Exercises

The playground has these and more, with hints and solutions, in your browser or as an RStudio project to download.

  1. Create a vector with the ages 24, 31, 28, 45, and 26. Use logical indexing to show only the ages over 30, and count how many there are.
  2. Create a factor from c("Agree", "Disagree", "Neutral", "Agree") whose levels run from "Disagree" through "Neutral" to "Agree". Check the order with table().
  3. Using students, count the students who are part-time and have children. (Use &.)
  4. Add a column to students called over_30 that is TRUE for students older than 30. Count those students, and explain why the count might need na.rm = TRUE.
  5. Read the SPSS file with read_sav() and find the question behind support_3.
  6. State the unit of analysis of each table in the package: students, semesters, questionnaire, and supervisors. For each, name one research question it could answer on its own.

2.14 Further reading

  • R for Data Science (Wickham et al. 2023): chapters “Data import”, “Spreadsheets”, and “Factors” go further with importing and with categories, and “A field guide to base R” covers the square-bracket style used in this chapter.
  • The documentation of the readxl and haven packages.

References

Wickham, Hadley, Mine Çetinkaya-Rundel, and Garrett Grolemund. 2023. R for Data Science: Import, Tidy, Transform, Visualize, and Model Data. 2nd ed. O’Reilly Media. https://r4ds.hadley.nz.