Lecture slides

Using AI with R

Using AI with R

Draft with AI, check everything, and stay responsible

Chapter 18

Polla Fattah

By the end of today you can

  • explain how the reliability of coding is measured, and why chance matters;
  • explain how large language models work, and why they make confident mistakes;
  • use an AI assistant to write, explain, and debug R code, and check its answers;
  • call a language model from R with ellmer, with structured answers;
  • validate AI coding against human coding with accuracy and Cohen’s kappa;
  • use AI responsibly: protect data, keep work reproducible, and disclose it.

Text as research data

“What has been your biggest challenge during your studies?”

535 students answered in their own words (RQ11).

200 answers were coded by hand into themes, which took most of two weeks. Could an AI tool code the rest reliably?

Qualitative coding

Reading each text and assigning it to a theme defined in a codebook.

Coding is judgement, and judgements differ.

Whether the coder is a person, a keyword list, or a program, the question is the same: does it code reliably?

Raw agreement can mislead

coder_a <- c(rep("Workload", 16), rep("Family", 4))
coder_b <- c(rep("Workload", 14), rep("Family", 2),   # the first 16
             rep("Workload", 2),  rep("Family", 2))   # the last 4
Observed agreement Chance agreement Cohen’s kappa
80% 68% 0.37

Each coder puts 80% under Workload, so random coders would agree 68% of the time.

Cohen’s kappa

\[ \kappa = \frac{\text{observed} - \text{chance}}{1 - \text{chance}} \]

Kappa Agreement
0 no better than chance
0.41 to 0.60 moderate
0.61 to 0.80 substantial
above 0.80 almost perfect

kap() from yardstick. The standard measure of inter-rater reliability.

How large language models work

  • a neural network (Chapter 15) with billions of weights;
  • trained on enormous amounts of text: books, websites, code;
  • learns to predict the next token, a word or part of one;
  • answers by generating one likely token at a time;
  • further training turns it into an assistant that follows instructions.

Why they make confident mistakes

Strength or weakness Because
fluent in R much R was in the training text
hallucination: plausible inventions it produces what is plausible, not verified
out of date knowledge stops at the training cut-off
knows nothing about your data unless you tell it

Use AI to draft, and check everything it produces.

Asking good questions

Include Example
the goal the percentage per faculty who considered dropping out
the data column names and types, e.g. from glimpse(); never identifiable data
the tools dplyr with the native pipe
for errors the exact message and the code

A plausible answer that is wrong

students |>
  summarise(percent = 100 * sum(considering_dropout == "Yes") / nrow(students),
            .by = faculty)

It runs, and the numbers look reasonable.

But they add up to 15: nrow(students) counts every student, not those in each faculty.

The corrected code

students |>
  summarise(percent = 100 * mean(considering_dropout == "Yes"),
            .by = faculty)

The mean of TRUE/FALSE is the proportion TRUE, within each faculty (Chapter 6).

Not a hallucination but a plausible misunderstanding: the most dangerous kind, because nothing looks wrong.

Four checks

  • run the code on a tiny example with a known answer;
  • check a number by another route: count(faculty, considering_dropout);
  • read the help page: does the function exist and do that?
  • ask the assistant to explain each line.

Assistants are excellent teachers: “explain this line by line”, “why does this give NA?”.

AI inside the editor

  • GitHub Copilot in RStudio suggests code as you type, even from a comment;
  • Positron Assistant and similar chats can see open files and, with permission, your data’s structure.

Suggestions arrive fast and look authoritative. Read each one before accepting it.

The answers to be coded

open_responses |> slice(c(5, 9, 15, 55)) |> pull(biggest_challenge)

They vary in length and style, often mention more than one challenge, and contain spelling mistakes.

Seven themes. The hand-coded sample is the benchmark, like a test set (Chapter 11).

The hand-coded themes

theme n
Supervision 68
Workload 45
Isolation 31
Health 20
Family 14
Finances 14
Other 8

Coding by keywords

keywords <- c(
  Supervision = "\\b(supervis|feedback|guidance|meeting)",
  Finances    = "\\b(money|fee|scholarship|rent|afford|salary|income|funding)",
  Family      = "\\b(famil|child|kid|son\\b|daughter|baby|parent|husband|wife)"
)  # plus Workload, Health, Isolation, and Other in the full list
code_by_keywords <- function(answer) {
  hits <- sapply(keywords, \(pattern) str_count(str_to_lower(answer), pattern))
  if (all(hits == 0)) "Other" else names(keywords)[which.max(hits)]
}

A regular expression: | means “or”, \\b marks the start of a word.

Words are not meaning

code_by_keywords("I miss my family and I cannot afford the fees.")
#> [1] "Finances"

One Family word, one Isolation phrase, two Finances words: Finances wins, where a reader might disagree.

Keywords against hand coding Accuracy Kappa
61% 0.51

“My supervisor” in an answer about letting down one’s family is about family.

Calling a language model from R

Kind Examples Needs
online chat_anthropic(), chat_openai(), chat_google_gemini() an API key; small cost per request
local chat_ollama() through Ollama a download; no key, no cost, text stays on your computer

Store an API key in .Renviron (usethis::edit_r_environ()), never in a shared script.

A chat from R

library(ellmer)
chat <- chat_ollama(model = "gemma4:e4b")            # local
# chat <- chat_anthropic(model = "claude-sonnet-5") # online, needs a key
chat$chat("In one sentence, what does Cohen's kappa measure?")

Always set the model explicitly and record it: models are updated and retired.

The study used a local model, so answers never left the computer and anyone can repeat it for free.

The codebook as a system prompt

codebook <- "
You are helping a researcher code open-ended survey answers ...
Assign each answer to exactly ONE theme: the main challenge.
Themes:
- Supervision: the supervisor; feedback, guidance, meetings.
- Workload: too much work or too little time; deadlines, reading.
- Finances: money, fees, scholarships, living costs.
- Family: children, partners, parents, caring duties.
- Health: physical or mental health, sleep, stress, burnout.
- Isolation: loneliness, far from home, no colleagues.
- Other: anything else, such as ethics approval or statistics.
"

The same instructions for every answer.

Structured answers for many texts

theme_type <- type_object(
  theme = type_enum(themes, "The single main theme of the answer."))

chat <- chat_anthropic(system_prompt = codebook, model = "claude-sonnet-5",
                       params = params(temperature = 0))

ai_codes <- parallel_chat_structured(
  chat, prompts = as.list(open_responses$biggest_challenge),
  type = theme_type)

type_enum() restricts replies to the themes. temperature = 0 makes answers as repeatable as possible.

Checking a local model’s replies in code

chat <- chat_ollama(system_prompt = codebook, model = "gemma4:e4b",
                    params = params(temperature = 0))
replies  <- parallel_chat_text(chat, as.list(open_responses$biggest_challenge))
ai_codes <- themes[match(tolower(gsub("[^A-Za-z]", "", replies)),
                         tolower(themes))]

Smaller models may ignore the structured format. Ask for the theme name only, then let R check it.

A reply that is not a theme becomes NA: counted, not silently accepted.

Run once, save, analyse the file

The script records the provider, model, date, software versions, and time taken, and codes the hand-coded sample twice.

Value
model gemma4:e4b
requests 735
time about 3 minutes
replies outside the themes 0

The model against the hand coding

ai_check |> accuracy(truth = theme, estimate = ai_theme)
ai_check |> kap(truth = theme, estimate = ai_theme)
Coder Accuracy Kappa
keywords 61% 0.51
language model 90% 0.87

The model reaches almost perfect agreement with the hand coding.

Read the disagreements

Answer Hand coding Model
Getting feedback on my draft took over a month, and by then I had l… Supervision Isolation
Combining coursework, my job at a school and my thesis leaves me no… Workload Health
I stopped exercising and I can feel the difference. The doctor told… Health Other
Writing my literature review took far longer than I planned. And I … Workload Isolation

Is the model wrong, the codebook unclear, or the answer ambiguous? A clearer rule helps human and machine alike.

Is the model repeatable?

Asked a second time, the model gave the same theme for 100% of the answers.

That is not guaranteed for every model or setting; online models can change between runs.

So the outputs are saved and analysed from the file.

What the students said

All 535 answers, coded by the model once its agreement was checked.

Text and numbers tell the same story

theme students mean support
Supervision 177 2.82
Isolation 61 3.30
Other 23 3.31
Health 56 3.33
Family 47 3.43
Finances 49 3.45
Workload 122 3.46

Students whose main challenge is supervision report the least supervisor support: a check of validity.

Check, and stay responsible

  • run code on known cases;
  • check facts and references against sources: assistants invent plausible references;
  • check coding against human coding.

Whatever appears in your thesis is your responsibility, whoever or whatever drafted it.

Protect your data

  • text sent online leaves your computer, and may be stored;
  • check the ethics approval and consent forms first;
  • check the provider’s terms: free accounts often allow training on your data;
  • remove names, places, and identifying details;
  • for sensitive data, use a local model, and validate it too.

Keep it reproducible

Record Why
provider, exact model, date models change and retire
settings such as temperature they change answers
system prompt and codebook the instructions are part of the method
the model’s outputs analyse the saved file, not a new call

Code a sample twice to check consistency.

Disclose it

An AI tool cannot be an author: it cannot take responsibility (COPE).

Answers were coded into seven themes by a large language model (Gemma 4, e4b version, run locally with Ollama and the ellmer R package, temperature 0) using the codebook in Appendix X. Agreement with the author’s hand coding of a random sample of 200 answers was assessed with Cohen’s kappa.

Be aware of bias

Models learn from human text and can reproduce its biases:

  • in how they describe groups of people;
  • in which answers they find typical;
  • in how well they understand non-standard English.

Check whether accuracy differs between groups, as for any model about people.

In your field: medicine and public health

medication_type <- type_object(
  drug          = type_string("Name of the medication"),
  dose_mg       = type_number("Dose in milligrams"),
  times_per_day = type_integer("How many times a day it is taken"))

chat$chat_structured(
  "Patient to continue metformin 500 mg twice daily with meals.",
  type = medication_type)

Free text into fixed fields. Validate against clinician-checked records, and never send patient data to public services.

Practical lab: the Chapter 18 playground

Work through the playground exercises in your browser, with hints and solutions.

The saved model coding is included, so no API key is needed.

Practical exercises 1–3: assistants and keywords

  1. Ask an assistant for average sleep by faculty and semester, and check it two ways.
  2. Add three keywords: does kappa improve, or are you overfitting?
  3. Read ten keyword disagreements: do you agree with the hand coding?

Practical exercises 4–6: models and reliability

  1. Code 20 answers with a language model; change one theme definition and recode.
  2. Write the methods paragraph for AI coding in your own study.
  3. A coder who always says Workload: raw agreement and kappa?

Try this yourself

Plan AI use in your own research.

  • what would you ask it to do, and what would you check by hand?
  • does your consent form allow sending data to a company?
  • which local model could you use instead?
  • what would you record, and how would you disclose it?

Troubleshooting guide (Part 1)

Symptom Likely cause
high agreement, low kappa one theme dominates; chance inflates agreement
a function that does not exist hallucination; check the help page
code runs but answers the wrong question a plausible misunderstanding; test it
an outdated way of doing something the training cut-off

Troubleshooting guide (Part 2)

Symptom Likely cause
replies in the wrong format a small model ignored structure; check in code
different answers on another day model or settings changed; save outputs
an API key in a shared script store it in .Renviron
the model disagrees on two-challenge answers the codebook’s rule is unclear

Completion checklist

Misconceptions to leave behind (Part 1)

Misconception Better mental model
high agreement means reliable coding kappa shows agreement beyond chance
the AI knows the answer it produces plausible text, to be validated

Misconceptions to leave behind (Part 2)

Misconception Better mental model
a model’s answers never change versions, settings, and dates change them
AI use must be hidden most allow it with disclosure

The chapter in one sentence

Treat AI output as a draft: validate it against human judgement, protect participants, record what you did, and say so.

Next: Chapter 19

The final chapter puts everything together:

  • the shape of a research project;
  • from raw export to clean data, and describing the sample;
  • answering each research question;
  • the chain of reasoning and the results chapter;
  • choosing a method, and where to go next.

Questions

Which part of your own analysis would you trust an AI to draft?

How would you check it?