flowchart LR
V["Vector<br>one variable"] --> F["Factor<br>categories"]
V --> D["Data frame<br>a table of cases"]
D --> M["Matrix and list<br>results"]
Data Structures in R
Data Structures in R
Data as the written result of measurement
Chapter 2
Polla Fattah
By the end of today you can
- describe data as a table of cases and variables;
- identify the unit of analysis;
- create vectors and factors, and select and change their values;
- select rows and columns from a data frame;
- recognise matrices and lists;
- import CSV, Excel, and SPSS files, keeping SPSS labels;
- build a codebook and export data;
- read R’s most common error messages.
Data is a record of measurement
Each value records one observation about one case.
Before any analysis, know:
- what each value stands for;
- what kind of measurement produced it;
- how the values are arranged.
These answers decide which summaries and tests make sense later.
Cases and variables
| Term | Meaning | Example |
|---|---|---|
| case | the thing being measured | a student, a patient, a plot of land |
| variable | one characteristic measured on every case | age, faculty |
| value | one variable on one case | 31 years |
Rows are cases, columns are variables. This data matrix is the starting point of nearly every method.
The unit of analysis
What one row represents is the unit of analysis, and one study can have several.
| Table | One row is |
|---|---|
students |
one student |
semesters |
one student in one semester |
The same student, four rows
student_id semester gpa sleep_hours study_hours
1 S0001 1 3.02 7.2 26
2 S0001 2 3.27 8.4 17
3 S0001 3 3.17 7.8 25
4 S0001 4 3.00 7.2 28
5 S0002 1 3.25 6.4 32
A question about students counts each student once.
A question about change over time needs all the rows, and methods that know they belong together (Chapter 10).
Data structures hold the measurements
R has a small number of ways to hold data. Everything is built from vectors.
A vector holds one variable
A set of values combined with c(), the basic building block of R.
One vector, one type
Mixing types converts everything to the most flexible type, usually text.
The numbers can no longer be averaged.
One typo can turn a column into text
If one person types “seven” instead of 7, the whole column arrives as text.
Survey exports do exactly this, and the Excel file later today shows it.
Operations apply to every value
Hours into minutes with one multiplication.
The same line works on five nights or on two thousand.
Comparisons give TRUE or FALSE for each value
Each night is checked against the recommended 7 hours.
Counting with TRUE and FALSE
R counts TRUE as 1 and FALSE as 0.
sum() counts the short nights; mean() gives their share.
Square brackets select by position
sleep[1] # the first night
sleep[c(1, 3)] # the first and third nights
sleep[-2] # every night except the secondPositions start at 1. A minus sign leaves a position out.
Logical indexing selects by a condition
Read it aloud: “sleep, where sleep is less than 7”.
One of the most useful ideas in R, used constantly.
Brackets on the left change values
The second night turns out to have been 7.5 hours, not 7.
Factors hold categories
faculty <- factor(c("Education", "Humanities",
"Education", "Health Sciences"))
levels(faculty)
#> [1] "Education" "Health Sciences" "Humanities"A factor looks like text but knows the complete set of possible values, its levels.
Set the order when categories have one
employment <- factor(
c("None", "Full-time job", "Part-time job", "None"),
levels = c("None", "Part-time job", "Full-time job")
)
table(employment)By default levels are alphabetical.
Tables and graphs follow the level order, so set it to match the meaning.
Levels decide groups and their order
When groups are compared (Chapter 7) or plotted (Chapter 4), R uses the factor’s levels:
- which groups exist;
- in which order they appear.
A wrongly ordered factor gives a correct but confusing table.
The kind of measurement decides the type
| Kind of measurement | Example | Stored in R as |
|---|---|---|
| nominal | faculty, gender | factor |
| ordinal | employment, agreement | factor with ordered levels |
| equal distances | sleep hours, wellbeing | numeric |
| yes or no | invited to the workshop | logical, or two-level factor |
Chapter 5 explains the levels of measurement in full.
A data frame is a table of cases
pilot <- data.frame(
student = c("S1", "S2", "S3", "S4"),
faculty = c("Education", "Humanities", "Education", "Health Sciences"),
sleep = c(6.5, 7.5, 5.5, 8),
invited = c(TRUE, FALSE, TRUE, FALSE)
)Each column is a vector with one type. Different columns can have different types.
str() shows the structure
str(pilot)
#> 'data.frame': 4 obs. of 4 variables:
#> $ student: chr "S1" "S2" "S3" "S4"
#> $ faculty: chr "Education" "Humanities" ...
#> $ sleep : num 6.5 7.5 5.5 8
#> $ invited: logi TRUE FALSE TRUE FALSEEvery column, its type, and its first values: the quickest way to understand a data frame.
Rows first, then columns
pilot[2, ] # the second row, all columns
pilot[, c("student", "sleep")] # all rows, two columns
pilot[pilot$sleep < 7, ] # rows where sleep is under 7An empty position means “all”.
The last line is the R version of SPSS’s Select Cases.
Adding a column
Assigning to a new name adds a column.
Here: TRUE for every student who slept under 7 hours.
The student wellbeing data
str(students)
#> 'data.frame': 600 obs. of 14 variables:
#> $ student_id : chr "S0001" "S0002" "S0003" ...
#> $ age : int 31 28 33 34 35 36 NA 25 ...
#> $ faculty : chr "Health Sciences" "Education" ...chr is text, int whole numbers, num numbers with decimals.
Categories arrive as text and become factors when needed as categories.
Two conditions at once
hs_phd <- students[students$faculty == "Health Sciences" &
students$programme == "PhD", ]
nrow(hs_phd)
#> [1] 44| Symbol | Meaning |
|---|---|
== |
is equal to (a single = is for arguments) |
& |
and: both must be TRUE |
| |
or: at least one must be TRUE |
Looking at data like a spreadsheet
A spreadsheet-style viewer: scroll, sort, and filter.
It is only for looking. Changes you want to keep are made with code, so they are recorded.
A matrix is a grid of one type
first_semester <- semesters[semesters$semester == 1, ]
vars <- first_semester[, c("gpa", "sleep_hours",
"study_hours", "wellbeing")]
correlations <- cor(vars, use = "complete.obs")Matrices usually appear as results, such as a correlation matrix.
The first semester only, so each student counts once.
Selecting from a matrix
The same [row, column] as a data frame.
Every variable correlates perfectly with itself, so the diagonal is all 1s (Chapter 6).
A list can hold anything
student <- list(
id = "S0001",
programme = "Master's",
sleep = c(7.2, 8.4, 7.8, 7.2)
)
student$sleepVectors of different lengths, data frames, even other lists, usually with names.
Statistical results are lists
result <- t.test(first_semester$sleep_hours, mu = 7)
names(result)
#> [1] "statistic" "parameter" "p.value" "conf.int" "estimate" ...
result$estimate
#> mean of x
#> 6.483502names() lists what the test calculated; $ retrieves one part for a report.
Every import gives a data frame
| Format | Package | Function |
|---|---|---|
| CSV | built in | read.csv() |
| Excel | readxl | read_excel() |
| SPSS, Stata, SAS | haven | read_sav(), read_dta(), read_sas() |
The example files come with data2thesis; data2thesis_example() finds them.
CSV: a plain text table
Commas separate the columns. The most common format for sharing data.
In your own project: read.csv("semesters.csv"), or read.csv(here::here("data", "semesters.csv")).
Excel: the survey export
615 rows for 600 students, and 64 columns.
Install readxl once with install.packages("readxl").
What a real export looks like
`Response ID` `Q1_Student ID` Q2_Age Q3_Gender Q4_Faculty
<chr> <chr> <chr> <chr> <chr>
1 R0001 TEST 99 Female <NA>
2 R0002 test <NA> <NA> <NA>
3 R0003 TEST2 30 Male <NA>
4 R0004 S0191 27 male Humanities
Test responses, 99 for a missing age, Male and male, and every column is text.
Chapter 3 cleans this file. Today it is enough to open it.
A tibble is a modern data frame
read_excel() returns a tibble.
- it prints more compactly;
- it shows each column’s type under its name;
- it is used exactly like a data frame.
For a workbook with several sheets: read_excel("file.xlsx", sheet = "Semester 2").
SPSS files keep their labels
Each value is stored as a number with its label: 2 [Health Sciences].
haven also reads Stata and SAS files.
The question behind a variable
The variable label travels with the data.
The work put into labelling in SPSS is not lost.
Labels become factor levels
as_factor() turns a labelled variable into a factor, using the labels as levels.
Ready to analyse as categories.
A value needs a codebook
A 4 in stress_1 means nothing on its own. The reader needs:
- the question behind it;
- the answer scale (1 = strongly disagree, 5 = strongly agree);
- how missing answers are recorded.
A codebook records this for every variable, usually in a thesis appendix.
Building a codebook from SPSS labels
codebook <- data.frame(
variable = names(spss),
label = sapply(spss, function(x) attr(x, "label")),
type = sapply(spss, function(x) class(x)[1]),
row.names = NULL
)sapply() applies a function to every column and collects the results.
Value labels record the categories
Which number stands for which category.
Write the codebook to a file and complete it by hand: answer scales and coding notes.
Exporting data
write.csv(students, "students_clean.csv", row.names = FALSE)
writexl::write_xlsx(students, "students_clean.xlsx")
haven::write_sav(spss, "students_clean.sav")row.names = FALSE stops R adding a column of row numbers.
Originals stay untouched
flowchart LR
A["Original file<br>never edited"] --> B["Script"] --> C["Cleaned file<br>new name"]
The script records how one became the other.
SPSS and Excel equivalents
| In SPSS or Excel | In R |
|---|---|
| open a data file | read_sav(), read_excel(), read.csv() |
| Variable View | str(), attr(x, "label") |
| Data View | View() |
| Select Cases | data[condition, ] |
| Compute Variable | data$new <- ... |
| Frequencies and Descriptives | table(), mean(), summary() |
The real change is where the steps live
SPSS / Excel: the data changes, the steps are forgotten
R: the data stays, the steps are saved in the script
Run the script again, and you get the same result again.
Error messages report what R could not do
They are not a sign of failure. The message usually says what went wrong, and often where.
Always read it. Four messages cover most beginner errors.
Object not found
A misspelled name, or the line that creates the object was never run.
Could not find function
A misspelled function name, or its package is not loaded with library().
Non-numeric argument
A calculation on text, often because a column arrived as text.
Undefined columns selected
The column is faculty. Names are case-sensitive.
Check spelling and capitals first, then that every line above was run, then the help page.
A first summary
Averages and ranges for numbers, counts for factors, and the number of missing values.
The average student sleeps less than 7 hours, and a few values are missing. Chapter 6 goes further.
In your field: health research
Vitamin C and tooth growth in 60 guinea pigs: an outcome, a treatment, and a dose.
supp is a factor (orange juice or ascorbic acid), with 10 animals per combination. One row is one animal.
Practical lab: the Chapter 2 playground
Work through the playground exercises in your browser, with hints and solutions.
Every exercise also runs in RStudio, from the downloadable chapter project.
Practical exercises 1–3: vectors, factors, selection
- Show only the ages over 30 in a vector, and count them.
- Build an agreement factor ordered from
"Disagree"to"Agree". - Count the students who are part-time and have children.
Practical exercises 4–6: columns, labels, units
- Add an
over_30column and explain why its count may needna.rm = TRUE. - Find the question behind
support_3in the SPSS file. - State the unit of analysis of
students,semesters,questionnaire, andsupervisors.
Try this yourself
Take a small dataset from your own field, or invent one with five cases.
- build it as a data frame with at least one factor;
- check it with
str(); - select the rows that meet one condition;
- write a three-line codebook for it.
Then say what one row represents.
Troubleshooting guide (Part 1)
| Symptom | Likely cause |
|---|---|
| a number column cannot be averaged | one entry is text, so the whole column is text |
| categories appear in the wrong order | factor levels are alphabetical by default |
| a count of students is too large | rows are student-semesters, not students |
= in a condition gives an error |
comparisons need == |
Troubleshooting guide (Part 2)
| Symptom | Likely cause |
|---|---|
| SPSS values show as numbers | labels not converted with as_factor() |
| an extra numbered column in the CSV | row.names = FALSE was left out |
undefined columns selected |
a typo or wrong capitals in a column name |
edits in View() are lost |
the viewer is only for looking |
Completion checklist
Misconceptions to leave behind (Part 1)
| Misconception | Better mental model |
|---|---|
| a factor is just text | it knows its levels and their order |
| a vector can mix numbers and text | mixing converts everything to text |
| every row is a student | the unit of analysis depends on the table |
Misconceptions to leave behind (Part 2)
| Misconception | Better mental model |
|---|---|
| error messages mean failure | they say precisely what R could not do |
| SPSS labels are lost in R | haven keeps them |
| the cleaned file replaces the original | the script links the original to the cleaned copy |
The chapter in one sentence
Know what one row represents, store each variable in the type its measurement deserves, and record what every value means.
Next: Chapter 3
The next chapter cleans the survey export:
- what clean and tidy data mean;
- the tidyverse and the pipe;
- selecting, filtering, creating, and summarising;
- reshaping and joining tables;
- a cleaning script that records every change.
Questions
In your own research data, what is one row?
Is there more than one unit of analysis hiding in it?