(6.5 + 7 + 5.5) / 3[1] 6.333333
Every data analysis is a long chain of decisions. Which cases are kept and which are excluded, how a messy answer is corrected, which variables are combined into a score, which test is used, and how its result is rounded: each choice shapes the final numbers, and each should be open to inspection. When an analysis is done by clicking through menus, most of these decisions leave no trace. Weeks later, not even the researcher can say exactly how a number in the thesis was produced, and an examiner who asks cannot be given a complete answer.
Writing the analysis as code changes this. A script records every step in order, in a form that a supervisor can read, an examiner can check, and the researcher can run again after correcting a mistake or receiving new data. This is the main reason research is increasingly done in a programming language, and it is the reason this book teaches R. The chapter starts from nothing: installing R, finding your way around the program used to write it, and learning the handful of ideas on which everything else rests.
The book follows one research project from its first day to its last. Elaf has just begun a Master’s degree in Educational Psychology. During her first year of graduate study she noticed how tired many of her fellow students were: they slept little, lived on coffee, worried about their supervisors and their money, and a few spoke quietly about leaving. With her supervisor, she turned that observation into a thesis. It asks how sleep, stress, and supervisor support relate to graduate students’ wellbeing, their grades, and their thoughts of dropping out, and whether a short wellbeing workshop can help. Her university agreed to support a two-year study of 600 graduate students from five faculties: a questionnaire at the start, a record of each student’s sleep, study, and wellbeing at the end of every semester, and a six-week workshop to which half of the students would be invited at random.
At their first meeting, her supervisor set out what the thesis would demand. The research questions had to become precise hypotheses before any results were seen. The raw survey export, full of test entries, typing errors, and impossible values, had to be cleaned without losing track of a single change. The sample had to be described honestly, including the students who stopped taking part. Every analysis had to be chosen for a reason and its assumptions checked, and every result reported with its size and its uncertainty, not only with a p-value. At the end, examiners would read the results chapter and could ask how any number in it had been produced. “You should be able to show them,” the supervisor said, “step by step.”
Elaf had used Excel for coursework and had once clicked through an SPSS menu in a statistics class, but she had never written a line of code. Her supervisor suggested R, for exactly the reason given above: an analysis written as code can be shown step by step. The first data has now arrived, and this chapter is Elaf’s first day with R, and yours. Each later chapter takes her one step further, from cleaning the data to the finished thesis. By the end of this one, you will have R running on your computer, you will know your way around the program used to write it, and you will have opened her data and answered a first, simple question about it.
Teaching & Practice: Interactive Lecture Slides · Browser Playground Exercises
R is a free program for working with data: cleaning it, analysing it, and turning it into tables and graphs. It began in the early 1990s at the University of Auckland, where two statisticians, Ross Ihaka and Robert Gentleman, wrote it for teaching (Ihaka and Gentleman 1996). Its first stable version appeared in 2000. Today it is used in universities, hospitals, governments, and companies around the world.
Four properties make R worth learning for a researcher. It is made for data: statistical tests, models, and graphs are part of the language rather than add-ons. It is free and open source, so you, your students, and anyone who wants to check your work can use it without buying a licence. Its instructions form a written record of the analysis, which anyone can run again to obtain exactly the same results; research that can be repeated in this way is called reproducible, and Chapter 17 returns to the idea. Finally, it is shared by a large community. Thousands of researchers publish free packages, add-ons that give R new abilities, so whatever method your field uses, someone has probably written a package for it.
The price is that R asks you to type instead of click. That feels slow at first. It becomes fast surprisingly quickly, and it pays you back every time you need to repeat, correct, or extend an analysis.
In SPSS or Excel, you click, and the program changes your data or produces output. In R, you write an instruction, and R carries it out. The instructions are saved in a file, so the next time you need the same analysis, you run the file instead of clicking through the menus again. Chapter 2 shows how familiar SPSS and Excel tasks map to R.
Working with R involves two programs. R is the engine that does the work. RStudio is the program in which you drive it: you write R code there, run it, and see the results. You will almost never open R itself; you open RStudio, and RStudio uses R for you.
R must be installed first. Download it from cran.r-project.org, choosing the version for your operating system (Windows, macOS, or Linux), and install it like any other program. Then download the free RStudio Desktop from posit.co and install it. Appendix A walks through both installations step by step, with solutions to common problems.
RStudio is the most widely used editor for R, and this book uses it throughout. Positron, a newer editor from the same company, works with R and Python in one window and is a good alternative if you plan to use both languages. Everything in this book works in either.
The exercises for this chapter run in your web browser. Try them in the playground first, and install R when you are ready.
When you open RStudio, the window is divided into panes. With a script open (you will open one shortly), there are four, described in Table 1.1.
| Where | Pane | What it is for |
|---|---|---|
| Top left | Source | Where you write and save your code, in files called scripts. |
| Bottom left | Console | Where R runs code and shows the results. |
| Top right | Environment | The objects you have created, such as your data. |
| Bottom right | Files, Plots, Packages, Help | Your files, your graphs, your installed packages, and R’s help pages. |
You can type code directly into the Console. Click in it, type 2 + 2, and press Enter. R answers straight away. The [1] at the start of the answer simply means “this is the first value of the result”; you can ignore it for now.
The simplest use of R is as a calculator. Suppose you slept 6.5, 7, and 5.5 hours on three nights last week. Your average is:
R follows the usual order of operations: multiplication and division before addition and subtraction, and brackets first of all. Without the brackets, R would divide only the last number by 3:
The usual arithmetic symbols work as you would expect: +, -, * (multiply), / (divide), and ^ (power).
Typing the same numbers again and again is tiresome and invites mistakes. Instead, you can store a value under a name. The stored value is called an object, and you create it with the assignment arrow <-, which you can read as “gets”:
Nothing is printed, but R now remembers that nights is 3. The object appears in the Environment pane, and you can use it by name:
To store several values in one object, combine them with c(), which stands for combine:
Chapter 2 explains these collections, called vectors, in detail. For now, notice how much clearer the code becomes when the numbers have a name:
Names follow a few rules. A name starts with a letter and can contain letters, numbers, dots, and underscores, such as sleep, sleep_week1, or avg.sleep. It cannot contain spaces, so an underscore takes their place: sleep_hours, not sleep hours. R is case-sensitive, which makes Sleep and sleep two different objects. Beyond these rules, the best names say what the object holds: sleep_hours is far better than x, and your future self will be grateful for it.
If you assign a new value to an existing name, the old value is replaced without warning:
Research data is not all numbers. A survey records numbers (hours of sleep), words (the student’s faculty), and yes-or-no answers (whether the student was invited to a workshop). R keeps track of the type of every value, because the type decides what can be done with it: hours of sleep can be averaged, but faculty names cannot. Table 1.2 lists the types you will meet most often.
| Type | Holds | Example |
|---|---|---|
| numeric | numbers, with or without decimals | 6.5, 600 |
| character | text, written inside quotes | "Education" |
| logical | TRUE or FALSE |
TRUE |
The function class() tells you the type of an object:
[1] "numeric"
[1] "character"
[1] "logical"
Text must be inside quotes. Without them, R thinks you mean an object with that name:
This is one of the most common errors for beginners. The message says exactly what went wrong: R looked for an object called Education and did not find one.
Real data always has gaps. A student skips a question, or a value is lost. R marks a missing value with NA, short for not available. It is neither zero nor empty text; it means “we do not know”:
If one value is unknown, the average is unknown too, so R returns NA rather than guessing. You will see shortly how to tell R to leave missing values out.
You have already used several functions: c(), sum(), mean(), and class(). A function takes some input, does something with it, and returns a result. You call a function by writing its name followed by brackets, with the input inside:
Many functions accept more than one input. The inputs are called arguments, and they are separated by commas. The function round(), for example, takes a number and the number of decimal places to keep:
Arguments have names. You can leave the names out if you give the arguments in the expected order, so round(6.333333, 1) gives the same result, but writing the names makes your code easier to read.
The missing value from before can now be handled. The function mean() has an argument called na.rm, short for “NA remove”. Set it to TRUE, and R calculates the average of the values it does know:
You can also put one function inside another. R works from the inside out:
Every function has a help page. Type a question mark before its name in the Console:
The page opens in the Help pane. Help pages are written for experienced users, so they can look dense at first. Start with three parts: Usage (how to call the function), Arguments (what each input means), and Examples at the bottom, which you can copy and run.
When a search engine or an AI assistant suggests code, treat it like advice from a knowledgeable stranger: often right, sometimes wrong, and always worth checking. Chapter 18 shows how to use AI tools well.
Code typed in the Console is gone once you close RStudio. For real work, write your code in a script: a plain text file, ending in .R, that holds your instructions in order. A new script is created with File > New File > R Script, and it opens in the Source pane. You type one instruction per line, and run the current line, or the lines you have selected, with Ctrl+Enter (Cmd+Enter on a Mac): the code is sent to the Console, and the result appears there. Ctrl+S (Cmd+S) saves the script.
Lines that start with # are comments. R ignores them; they are notes for people. Use them to explain why you did something:
[1] 6.3
A script is the analysis written down. Months later, when a supervisor asks how a number was obtained, the script is the answer.
Research involves many files: data, scripts, graphs, and drafts. An RStudio Project keeps everything for one piece of work together in one folder, and makes sure R looks for files in that folder.
Create one with File > New Project > New Directory > New Project, give it a name (for example, wellbeing-thesis), and choose where to put it. RStudio creates the folder, with a file ending in .Rproj inside. From then on, double-click that file to open the project, and RStudio starts in the right place with the right files.
The folder R looks in for files is called the working directory. In a project, it is the project folder, so a file stored there can be opened by its name alone:
As a project grows, it needs subfolders, such as data/ for data and figures/ for graphs. The here package builds file paths that start from the project folder, so the same code works on any computer:
setwd()
Older tutorials start scripts with setwd("C:/Users/Elaf/Documents/thesis"). That line works only on the computer where it was written. Use a project, and your code works for your supervisor too.
R comes with a lot built in, but much of its power comes from packages. Using a package takes two steps. It is installed once on your computer with install.packages(), which downloads it from CRAN, R’s official collection of packages:
It is then loaded, with library(), in every session in which you want to use it:
A useful comparison is a library book: installing a package is like buying a book and putting it on your shelf, and loading it is taking it off the shelf to read. You buy it once, but you take it down every time you need it.
The data for this book comes as a package too, called data2thesis. It is not on CRAN, so you install it from this book’s website:
With R, RStudio, and the package installed, the data can be loaded:
Elaf, her study, and her data are fictional. The data was generated by a computer program to look like a realistic survey of graduate students, so that every method in this book has something to find. It describes no real people, and its patterns (for example, that the workshop raised wellbeing, or that students who sleep more have higher grades) were built into the program by the author, not discovered. Nothing in this book is evidence about the wellbeing of real students, and the data and results must not be cited or used as findings about students, universities, or any real situation. The same applies to the counselling service’s records and to the students’ written answers.
The package contains several data frames: tables with one row per case and one column per variable, like a spreadsheet. The main one is students, with one row per student. Two functions give its size:
There are 600 students and 14 variables. The function head() shows the first six rows:
student_id supervisor_id age gender faculty programme study_mode
1 S0001 SUP088 31 Female Health Sciences Master's Full-time
2 S0002 SUP116 28 Female Education PhD Full-time
3 S0003 SUP040 33 Male Education PhD Full-time
4 S0004 SUP112 34 Female Health Sciences Master's Part-time
5 S0005 SUP118 35 Female Social Sciences Master's Part-time
6 S0006 SUP028 36 Male Humanities PhD Full-time
employment has_children lives_away financial_worry workshop
1 None No Yes 2 Not invited
2 Part-time job No Yes 2 Invited
3 None No No 4 Invited
4 Full-time job No No 2 Invited
5 None No No 2 Invited
6 Part-time job No Yes 5 Not invited
workshop_sessions considering_dropout
1 0 No
2 0 No
3 5 No
4 6 No
5 2 No
6 0 No
The names of the variables are listed by names(), and the help page, ?students, describes each one.
[1] "student_id" "supervisor_id" "age"
[4] "gender" "faculty" "programme"
[7] "study_mode" "employment" "has_children"
[10] "lives_away" "financial_worry" "workshop"
[13] "workshop_sessions" "considering_dropout"
To pick out one variable, write the data frame’s name, a dollar sign, and the variable’s name: students$age is the age of every student. A natural first question concerns the age of the students in the study:
The answer is NA. One student’s age is missing, and R will not guess, but the solution is already familiar:
The students are 29.7 years old on average. The number of missing values can be counted too: is.na() marks each missing value as TRUE, and sum() counts them, because R counts every TRUE as 1:
Only 1 age is missing. It is worth knowing: Chapter 3 shows why it is missing, and what to do about it.
Finally, table() counts how many times each value appears. It shows how many students come from each faculty, and how many were invited to the wellbeing workshop:
Education Health Sciences Humanities Natural Sciences
148 154 95 87
Social Sciences
116
Invited Not invited
300 300
Exactly 300 students, half of the study, were invited. That is not a coincidence: they were chosen at random, and Chapter 5 explains why random assignment matters so much for the conclusions a study can draw.
R comes with dozens of real datasets for practice. PlantGrowth records the dried weight of plants grown under a control condition and two different treatments, from a classic experiment comparing crop yields. The same functions work on it:
weight group
1 4.17 ctrl
2 5.58 ctrl
3 5.18 ctrl
4 6.11 ctrl
5 4.50 ctrl
6 4.61 ctrl
ctrl trt1 trt2
10 10 10
[1] 5.073
Type data() in the Console to see the full list of built-in datasets, and ?PlantGrowth to read about this one.
<- stores a value in an object. c() combines several values.TRUE or FALSE). NA marks a missing value.? opens a function’s help page.install.packages(), and load them each session with library().R, RStudio, Console, script, comment, object, assignment, vector, data type, missing value (NA), function, argument, package, data frame, working directory, RStudio Project, reproducible research.
These short exercises check your understanding. The playground has more, with hints and solutions, which you can run in your browser or download as an RStudio project.
supervisor_sleep and find the average, rounded to one decimal place."600" and 600 with class(), and explain the difference.mean(c(4, NA, 6)) returns NA, and show how to get the average of the two known values.students, find the youngest and the oldest student. (Hint: min() and max() also have an na.rm argument.)table(), find out how many students are studying part-time (the variable is study_mode).