Lecture slides

Data Structures in R

Data Structures in R

Data as the written result of measurement

Chapter 2

Polla Fattah

By the end of today you can

  • describe data as a table of cases and variables;
  • identify the unit of analysis;
  • create vectors and factors, and select and change their values;
  • select rows and columns from a data frame;
  • recognise matrices and lists;
  • import CSV, Excel, and SPSS files, keeping SPSS labels;
  • build a codebook and export data;
  • read R’s most common error messages.

Data is a record of measurement

Each value records one observation about one case.

Before any analysis, know:

  • what each value stands for;
  • what kind of measurement produced it;
  • how the values are arranged.

These answers decide which summaries and tests make sense later.

Cases and variables

Term Meaning Example
case the thing being measured a student, a patient, a plot of land
variable one characteristic measured on every case age, faculty
value one variable on one case 31 years

Rows are cases, columns are variables. This data matrix is the starting point of nearly every method.

The unit of analysis

What one row represents is the unit of analysis, and one study can have several.

Table One row is
students one student
semesters one student in one semester

The same student, four rows

  student_id semester  gpa sleep_hours study_hours
1      S0001        1 3.02         7.2          26
2      S0001        2 3.27         8.4          17
3      S0001        3 3.17         7.8          25
4      S0001        4 3.00         7.2          28
5      S0002        1 3.25         6.4          32

A question about students counts each student once.

A question about change over time needs all the rows, and methods that know they belong together (Chapter 10).

Data structures hold the measurements

flowchart LR
    V["Vector<br>one variable"] --> F["Factor<br>categories"]
    V --> D["Data frame<br>a table of cases"]
    D --> M["Matrix and list<br>results"]

R has a small number of ways to hold data. Everything is built from vectors.

A vector holds one variable

sleep <- c(6.5, 7, 5.5, 8, 6)
sleep
#> [1] 6.5 7.0 5.5 8.0 6.0

A set of values combined with c(), the basic building block of R.

One vector, one type

mixed <- c(6.5, "seven", 5.5)
mixed
#> [1] "6.5"   "seven" "5.5"
class(mixed)
#> [1] "character"

Mixing types converts everything to the most flexible type, usually text.

The numbers can no longer be averaged.

One typo can turn a column into text

If one person types “seven” instead of 7, the whole column arrives as text.

Survey exports do exactly this, and the Excel file later today shows it.

Operations apply to every value

sleep * 60
#> [1] 390 420 330 480 360

Hours into minutes with one multiplication.

The same line works on five nights or on two thousand.

Comparisons give TRUE or FALSE for each value

sleep < 7
#> [1]  TRUE FALSE  TRUE FALSE  TRUE

Each night is checked against the recommended 7 hours.

Counting with TRUE and FALSE

sum(sleep < 7)
#> [1] 3
mean(sleep < 7)
#> [1] 0.6

R counts TRUE as 1 and FALSE as 0.

sum() counts the short nights; mean() gives their share.

Square brackets select by position

sleep[1]         # the first night
sleep[c(1, 3)]   # the first and third nights
sleep[-2]        # every night except the second

Positions start at 1. A minus sign leaves a position out.

Logical indexing selects by a condition

sleep[sleep < 7]
#> [1] 6.5 5.5 6.0

Read it aloud: “sleep, where sleep is less than 7”.

One of the most useful ideas in R, used constantly.

Brackets on the left change values

sleep[2] <- 7.5
sleep
#> [1] 6.5 7.5 5.5 8.0 6.0

The second night turns out to have been 7.5 hours, not 7.

Factors hold categories

faculty <- factor(c("Education", "Humanities",
                    "Education", "Health Sciences"))
levels(faculty)
#> [1] "Education"  "Health Sciences"  "Humanities"

A factor looks like text but knows the complete set of possible values, its levels.

Set the order when categories have one

employment <- factor(
  c("None", "Full-time job", "Part-time job", "None"),
  levels = c("None", "Part-time job", "Full-time job")
)
table(employment)

By default levels are alphabetical.

Tables and graphs follow the level order, so set it to match the meaning.

Levels decide groups and their order

When groups are compared (Chapter 7) or plotted (Chapter 4), R uses the factor’s levels:

  • which groups exist;
  • in which order they appear.

A wrongly ordered factor gives a correct but confusing table.

The kind of measurement decides the type

Kind of measurement Example Stored in R as
nominal faculty, gender factor
ordinal employment, agreement factor with ordered levels
equal distances sleep hours, wellbeing numeric
yes or no invited to the workshop logical, or two-level factor

Chapter 5 explains the levels of measurement in full.

A data frame is a table of cases

pilot <- data.frame(
  student = c("S1", "S2", "S3", "S4"),
  faculty = c("Education", "Humanities", "Education", "Health Sciences"),
  sleep   = c(6.5, 7.5, 5.5, 8),
  invited = c(TRUE, FALSE, TRUE, FALSE)
)

Each column is a vector with one type. Different columns can have different types.

str() shows the structure

str(pilot)
#> 'data.frame':    4 obs. of  4 variables:
#>  $ student: chr  "S1" "S2" "S3" "S4"
#>  $ faculty: chr  "Education" "Humanities" ...
#>  $ sleep  : num  6.5 7.5 5.5 8
#>  $ invited: logi  TRUE FALSE TRUE FALSE

Every column, its type, and its first values: the quickest way to understand a data frame.

Rows first, then columns

pilot[2, ]                        # the second row, all columns
pilot[, c("student", "sleep")]    # all rows, two columns
pilot[pilot$sleep < 7, ]          # rows where sleep is under 7

An empty position means “all”.

The last line is the R version of SPSS’s Select Cases.

Adding a column

pilot$short_sleep <- pilot$sleep < 7

Assigning to a new name adds a column.

Here: TRUE for every student who slept under 7 hours.

The student wellbeing data

str(students)
#> 'data.frame':    600 obs. of  14 variables:
#>  $ student_id   : chr  "S0001" "S0002" "S0003" ...
#>  $ age          : int  31 28 33 34 35 36 NA 25 ...
#>  $ faculty      : chr  "Health Sciences" "Education" ...

chr is text, int whole numbers, num numbers with decimals.

Categories arrive as text and become factors when needed as categories.

Two conditions at once

hs_phd <- students[students$faculty == "Health Sciences" &
                   students$programme == "PhD", ]
nrow(hs_phd)
#> [1] 44
Symbol Meaning
== is equal to (a single = is for arguments)
& and: both must be TRUE
| or: at least one must be TRUE

Looking at data like a spreadsheet

View(students)

A spreadsheet-style viewer: scroll, sort, and filter.

It is only for looking. Changes you want to keep are made with code, so they are recorded.

A matrix is a grid of one type

first_semester <- semesters[semesters$semester == 1, ]
vars <- first_semester[, c("gpa", "sleep_hours",
                           "study_hours", "wellbeing")]
correlations <- cor(vars, use = "complete.obs")

Matrices usually appear as results, such as a correlation matrix.

The first semester only, so each student counts once.

Selecting from a matrix

correlations["sleep_hours", "wellbeing"]
#> [1] 0.46

The same [row, column] as a data frame.

Every variable correlates perfectly with itself, so the diagonal is all 1s (Chapter 6).

A list can hold anything

student <- list(
  id        = "S0001",
  programme = "Master's",
  sleep     = c(7.2, 8.4, 7.8, 7.2)
)
student$sleep

Vectors of different lengths, data frames, even other lists, usually with names.

Statistical results are lists

result <- t.test(first_semester$sleep_hours, mu = 7)
names(result)
#>  [1] "statistic" "parameter" "p.value" "conf.int" "estimate" ...
result$estimate
#> mean of x
#>  6.483502

names() lists what the test calculated; $ retrieves one part for a report.

Every import gives a data frame

Format Package Function
CSV built in read.csv()
Excel readxl read_excel()
SPSS, Stata, SAS haven read_sav(), read_dta(), read_sas()

The example files come with data2thesis; data2thesis_example() finds them.

CSV: a plain text table

path <- data2thesis_example("semesters.csv")
semester_file <- read.csv(path)

Commas separate the columns. The most common format for sharing data.

In your own project: read.csv("semesters.csv"), or read.csv(here::here("data", "semesters.csv")).

Excel: the survey export

library(readxl)
raw <- read_excel(data2thesis_example("wellbeing_raw.xlsx"))
dim(raw)
#> [1] 615  64

615 rows for 600 students, and 64 columns.

Install readxl once with install.packages("readxl").

What a real export looks like

  `Response ID` `Q1_Student ID` Q2_Age Q3_Gender Q4_Faculty
  <chr>         <chr>           <chr>  <chr>     <chr>
1 R0001         TEST            99     Female    <NA>
2 R0002         test            <NA>   <NA>      <NA>
3 R0003         TEST2           30     Male      <NA>
4 R0004         S0191           27     male      Humanities

Test responses, 99 for a missing age, Male and male, and every column is text.

Chapter 3 cleans this file. Today it is enough to open it.

A tibble is a modern data frame

read_excel() returns a tibble.

  • it prints more compactly;
  • it shows each column’s type under its name;
  • it is used exactly like a data frame.

For a workbook with several sheets: read_excel("file.xlsx", sheet = "Semester 2").

SPSS files keep their labels

library(haven)
spss <- read_sav(data2thesis_example("wellbeing.sav"))

Each value is stored as a number with its label: 2 [Health Sciences].

haven also reads Stata and SAS files.

The question behind a variable

attr(spss$stress_1, "label")
#> [1] "I feel unable to control important things in my studies."

The variable label travels with the data.

The work put into labelling in SPSS is not lost.

Labels become factor levels

table(as_factor(spss$faculty))

as_factor() turns a labelled variable into a factor, using the labels as levels.

Ready to analyse as categories.

A value needs a codebook

A 4 in stress_1 means nothing on its own. The reader needs:

  • the question behind it;
  • the answer scale (1 = strongly disagree, 5 = strongly agree);
  • how missing answers are recorded.

A codebook records this for every variable, usually in a thesis appendix.

Building a codebook from SPSS labels

codebook <- data.frame(
  variable = names(spss),
  label    = sapply(spss, function(x) attr(x, "label")),
  type     = sapply(spss, function(x) class(x)[1]),
  row.names = NULL
)

sapply() applies a function to every column and collects the results.

Value labels record the categories

attr(spss$faculty, "labels")
#>       Education  Health Sciences       Humanities
#>               1                2                3

Which number stands for which category.

Write the codebook to a file and complete it by hand: answer scales and coding notes.

Exporting data

write.csv(students, "students_clean.csv", row.names = FALSE)
writexl::write_xlsx(students, "students_clean.xlsx")
haven::write_sav(spss, "students_clean.sav")

row.names = FALSE stops R adding a column of row numbers.

Originals stay untouched

flowchart LR
    A["Original file<br>never edited"] --> B["Script"] --> C["Cleaned file<br>new name"]

The script records how one became the other.

SPSS and Excel equivalents

In SPSS or Excel In R
open a data file read_sav(), read_excel(), read.csv()
Variable View str(), attr(x, "label")
Data View View()
Select Cases data[condition, ]
Compute Variable data$new <- ...
Frequencies and Descriptives table(), mean(), summary()

The real change is where the steps live

SPSS / Excel:  the data changes, the steps are forgotten
R:             the data stays, the steps are saved in the script

Run the script again, and you get the same result again.

Error messages report what R could not do

They are not a sign of failure. The message usually says what went wrong, and often where.

Always read it. Four messages cover most beginner errors.

Object not found

mean(slep)
#> Error: object 'slep' not found

A misspelled name, or the line that creates the object was never run.

Could not find function

reed_csv("students.csv")
#> Error in reed_csv("students.csv"): could not find function "reed_csv"

A misspelled function name, or its package is not loaded with library().

Non-numeric argument

mixed * 2
#> Error in mixed * 2: non-numeric argument to binary operator

A calculation on text, often because a column arrived as text.

Undefined columns selected

students[, "Faculty"]
#> Error: undefined columns selected

The column is faculty. Names are case-sensitive.

Check spelling and capitals first, then that every line above was run, then the help page.

A first summary

summary(first_semester[, c("gpa", "sleep_hours",
                           "study_hours", "wellbeing")])

Averages and ranges for numbers, counts for factors, and the number of missing values.

The average student sleeps less than 7 hours, and a few values are missing. Chapter 6 goes further.

In your field: health research

str(ToothGrowth)
table(ToothGrowth$supp, ToothGrowth$dose)

Vitamin C and tooth growth in 60 guinea pigs: an outcome, a treatment, and a dose.

supp is a factor (orange juice or ascorbic acid), with 10 animals per combination. One row is one animal.

Practical lab: the Chapter 2 playground

Work through the playground exercises in your browser, with hints and solutions.

Every exercise also runs in RStudio, from the downloadable chapter project.

Practical exercises 1–3: vectors, factors, selection

  1. Show only the ages over 30 in a vector, and count them.
  2. Build an agreement factor ordered from "Disagree" to "Agree".
  3. Count the students who are part-time and have children.

Practical exercises 4–6: columns, labels, units

  1. Add an over_30 column and explain why its count may need na.rm = TRUE.
  2. Find the question behind support_3 in the SPSS file.
  3. State the unit of analysis of students, semesters, questionnaire, and supervisors.

Try this yourself

Take a small dataset from your own field, or invent one with five cases.

  • build it as a data frame with at least one factor;
  • check it with str();
  • select the rows that meet one condition;
  • write a three-line codebook for it.

Then say what one row represents.

Troubleshooting guide (Part 1)

Symptom Likely cause
a number column cannot be averaged one entry is text, so the whole column is text
categories appear in the wrong order factor levels are alphabetical by default
a count of students is too large rows are student-semesters, not students
= in a condition gives an error comparisons need ==

Troubleshooting guide (Part 2)

Symptom Likely cause
SPSS values show as numbers labels not converted with as_factor()
an extra numbered column in the CSV row.names = FALSE was left out
undefined columns selected a typo or wrong capitals in a column name
edits in View() are lost the viewer is only for looking

Completion checklist

Misconceptions to leave behind (Part 1)

Misconception Better mental model
a factor is just text it knows its levels and their order
a vector can mix numbers and text mixing converts everything to text
every row is a student the unit of analysis depends on the table

Misconceptions to leave behind (Part 2)

Misconception Better mental model
error messages mean failure they say precisely what R could not do
SPSS labels are lost in R haven keeps them
the cleaned file replaces the original the script links the original to the cleaned copy

The chapter in one sentence

Know what one row represents, store each variable in the type its measurement deserves, and record what every value means.

Next: Chapter 3

The next chapter cleans the survey export:

  • what clean and tidy data mean;
  • the tidyverse and the pipe;
  • selecting, filtering, creating, and summarising;
  • reshaping and joining tables;
  • a cleaning script that records every change.

Questions

In your own research data, what is one row?

Is there more than one unit of analysis hiding in it?