Playground: Chapter 18

Using AI with R

This page practises the ideas of Chapter 18 in four parts. Practise the chapter works through the book’s exercises, with starter code, hints, and solutions. Go further goes beyond the book, with new questions about agreement and about checking code. Check your understanding asks short questions whose answers open with a click. Do it yourself gives open tasks with no answers: you decide how to use and check AI in your own work, as you will in your thesis.

Calling a language model needs a local model on your own computer (Ollama) or an API key, so that step is practised in the download project. Everything around it runs here: checking code an assistant suggests, the keyword method, Cohen’s kappa, and the language model’s saved results from the chapter, which the book’s script produced once with a local model (Gemma 4, e4b version, through Ollama). In the exercises, replace each ______ with your own code and press Run Code.

Practise the chapter

These are the exercises at the end of Chapter 18, with the same numbers. In the setup, open_responses_coded holds the 200 hand-coded answers (theme), ai_coding the language model’s saved theme for every answer, keywords and code_by_keywords() the chapter’s dictionary method, and kappa_of() calculates Cohen’s kappa.

Exercise 1: Checking an assistant’s code

An assistant, asked for the average sleep in each faculty and semester, suggested the first block of code below. Run it, check it with a tiny example and by another route, and fix it.

NoteHint

The tiny example shows what one missing value does to a mean. The argument that removes missing values before calculating was met in Chapter 1.

TipSolution
summarise(sleep = mean(sleep_hours, na.rm = TRUE), ...)

The suggestion runs without an error, but 13 of the 20 faculty-semester averages are NA: a single missing sleep value in a group makes its mean missing, as the tiny example shows. With na.rm = TRUE, every group has an average (Education students slept 6.64 hours in semester 1), and counting the missing values shows how many answers each average leaves out, which a thesis should report. The check by another route, a sum divided by a count, gives the same 6.64. An assistant’s code that runs is not necessarily code that answers the question.

Exercise 2: Three more keywords

Add three keywords to the dictionary that you think would fix some of its mistakes, and report whether the kappa improves. Discuss whether you are now fitting the dictionary to these answers.

NoteHint

Add a word of your choice inside the Workload pattern, for example course. Any word you add must appear in the answers to change anything.

TipSolution
new_keywords["Workload"] <- "\\b(time|deadline|workload|too much|assignment|reading|busy|thesis|work|course)"

The original dictionary agrees with the hand coding on 61% of the answers (kappa 0.51). With thesis, work, and course added to Workload, and isolat and on my own added to Isolation, the agreement rises only to 61.5%, and the kappa does not improve (0.51): every new word fixes some answers and breaks others (“work” also appears in answers about jobs and money). Other choices will do a little better or worse. Either way, adjusting the dictionary while looking at these 200 answers is the overfitting of Chapter 11: the improvement would be smaller, or vanish, on the 335 answers the dictionary was not tuned on. A fair test needs answers that were not used to build it.

Exercise 3: Reading the disagreements

Read ten answers where the keyword coding disagrees with the hand coding, and say whether you would agree with the hand coding in every case.

NoteHint

Keep the answers where the two themes are not equal.

TipSolution
filter(theme != keyword_theme)

The keywords disagree with the hand coding on 78 of the 200 answers. In most of the ten, the hand coding is clearly right, and the keyword failure is instructive: “the comments I get are very brief” is about supervision without a single supervision keyword; “the other students already knew each other when I arrived” is isolation without “lonely” or “alone”; “I drink far too much coffee just to stay awake. Plus there are too many deadlines” has keywords for both health and workload, and the dictionary picks the wrong one. A few are genuinely ambiguous, mentioning two challenges equally, and a second human coder might have chosen differently. Reading disagreements shows whether the benchmark itself is trustworthy, and where the codebook needs clearer rules.

Exercise 4: The language model’s coding

Code 20 of the hand-coded answers with a language model, using the chapter’s codebook, and count how many agree with the hand coding; then change one theme definition and code them again. Running a model needs Ollama or an API key, so that is done in the download project. Here, look at what the chapter’s model did with the same 200 answers: its saved themes are in ai_coding.

NoteHint

Kappa compares the hand coding with the model’s theme, in the column ai_theme.

TipSolution
kappa = kappa_of(model_check$theme, model_check$ai_theme)

The model agrees with the hand coding on 18 of the first 20 answers, and on 90% of all 200, with a kappa of 0.87, almost perfect agreement, against 0.51 for the keywords. Its disagreements are spread thinly, the largest being answers the hand coding called Supervision (8 of 68 went elsewhere). Asked the same 200 answers a second time, at temperature 0, it gave exactly the same themes. In the download project, coding 20 answers with your own model and changing one theme definition shows how much the results depend on the wording of the codebook, which is why the codebook used must be reported.

Exercise 5: A methods paragraph for AI coding

Write the paragraph for your own methods section describing how you would use and validate AI coding in a study of your own. Write your answer first, then open the model answer.

“Open-ended answers were coded into [n] themes defined in a codebook (Appendix X). A random sample of 200 answers was first coded by the author. All answers were then coded by a large language model ([model and version], run locally with [software], temperature 0), given the codebook and one answer at a time. Agreement between the model and the author on the hand-coded sample was assessed with Cohen’s kappa, and disagreements were read and discussed; the model’s coding was used only because agreement was almost perfect (κ = [value]). The model was run a second time on the sample to check that its coding was stable. No identifiable data was sent to an online service. The prompts, the model’s replies, and the date of the run are available in the supplementary material.” The paragraph names the model, its settings, how it was validated against a human, and how the data was protected.

Exercise 6: A coder who always says Workload

In the two-coder example at the start of the chapter, change the second coder so that it puts every answer under Workload. Calculate the raw agreement and kappa, and explain the result.

NoteHint

The second coder gives the same theme to all 20 answers.

TipSolution
coder_b <- rep("Workload", 20)

The raw agreement is 80%, higher than many real coders manage, and kappa is exactly 0. The second coder has done no coding at all; it agrees on the 16 Workload answers only because it says Workload to everything, which is exactly the agreement expected by chance given the two coders’ proportions. Kappa removes that chance agreement and reveals that nothing is left. A lazy coder, or a model that always predicts the most common theme, can look good on raw agreement; kappa cannot be fooled this way.

Go further

These exercises go beyond the book.

Exercise 7: Kappa with three themes

Two coders read 20 answers and agree on 16. Calculate the agreement and kappa, then add more answers on which both coders say Workload (10, 20, 40, and 80 of them in all), keeping the same four disagreements, and watch how the two measures change.

NoteHint

Kappa compares the two coders, coder_a and coder_b.

TipSolution
kappa = kappa_of(coder_a, coder_b)

With 10 Workload answers, the coders agree on 80% of the answers, and kappa is 0.68. As Workload becomes more common, the raw agreement climbs to 87%, 92%, and 96%, but kappa rises much less, to 0.73, 0.76, and 0.78. The four disagreements stay the same, and the extra agreement comes from a theme that both coders would often pick by chance, so kappa gives it little credit. When one category dominates, raw agreement flatters the coders; kappa keeps the rarer themes, where coding is hard, in view.

Exercise 8: The dictionary’s most common confusion

Make the confusion matrix of the keyword coding against the hand coding, and list the most common kinds of disagreement.

NoteHint

The disagreements are the cells where the hand-coded theme and the keyword theme are different.

TipSolution
filter(hand != keywords)

The most common confusion is Workload answers coded as Other (10): answers about workload that contain none of the Workload keywords, such as “I underestimated how much work a thesis is”, fall through to Other. Next come Isolation answers coded as Supervision (8) and as Other (6): isolation is described in many different words (“nobody in my group works on anything close to my topic”), and mentions of a supervisor pull answers to Supervision. Isolation is the theme the dictionary handles worst, with only 8 of 31 correct. A confusion matrix tells you where a method fails, which a single agreement figure hides.

Exercise 9: Long answers and short answers

Count the words in each hand-coded answer with str_count(), and compare the average length by theme, and between answers the keywords coded correctly and incorrectly.

NoteHint

A word is a run of characters that are not spaces; the regular expression \\S+ matches one.

TipSolution
mutate(words = str_count(biggest_challenge, "\\S+"))

Supervision answers are the longest (21 words on average) and Isolation answers the shortest (15). The keywords coded correctly answers averaging 17 words, and incorrectly answers averaging 20: longer answers tend to mention more than one challenge, and a dictionary that counts keywords is easily pulled towards the wrong one. Short answers are hard in a different way, since they give so few words to go on. Knowing which answers are hard helps to decide where a human should check the automatic coding.

Exercise 10: Three mistakes an assistant might make

An assistant wrote the three snippets below. Each runs, or almost runs, and each has a different kind of mistake. Run them, find each mistake, and fix it.

NoteHint

For each result, ask whether it is possible: an average of NA, a faculty with no students, and percentages that do not describe each faculty. Check the exact spelling of the categories with unique(students$faculty).

TipSolution
mean(students$age, na.rm = TRUE)                              # 29.7
students |> filter(faculty == "Education") |> nrow()          # 148
students |>
  count(faculty, study_mode) |>
  mutate(percent = 100 * n / sum(n), .by = faculty) |>
  filter(study_mode == "Part-time")

The first returns NA, because one age is missing: a missing-value mistake, easy to spot. The second returns 0, because the faculty is spelled “Education” with a capital letter: a mistake about the data, which an assistant cannot know unless it is told the exact values. The third runs and gives plausible numbers, but divides by all 600 students, so the percentages are shares of the whole university, not of each faculty; adding .by = faculty makes each faculty’s rows add up to 100. This last kind, a silent logic error like the chapter’s example, is the most dangerous, because nothing looks wrong.

Check your understanding

Answer each question in your own words first, then click to see a model answer.

1. Why can raw agreement between two coders be misleading?

Because some agreement happens by chance, and when one theme is common, chance agreement is high: two coders who both choose the common theme most of the time will agree often even if they are guessing. Cohen’s kappa subtracts the agreement expected by chance, so a coder who always chooses the most common theme gets a kappa of 0, however high the raw agreement.

2. What is a hallucination?

A confident, fluent answer from a language model that is false: an invented function, argument, reference, or fact. It happens because a language model generates plausible text rather than looking things up. Hallucinations are found by checking: running code, reading help pages, and verifying every reference.

3. Why is a temperature of 0 used when a model codes research data?

Temperature controls how much randomness the model uses when choosing each word. At 0, it always chooses the most likely answer, so the same answer gets the same theme when asked again, as the chapter’s rerun showed. Coding needs consistency and reproducibility, not variety; a higher temperature is for creative text.

4. What must be recorded so that AI coding can be reproduced and checked?

The provider and the exact model and version, the date, the settings (such as temperature), the full prompt and codebook, the software versions, and the model’s replies themselves, saved to a file. Models change and are withdrawn, so the saved replies are often the only way to check the results later.

5. When must data not be sent to an online model?

When it could identify participants or is otherwise confidential, unless the ethics approval, the consent form, and the university’s rules explicitly allow it, and the provider’s terms protect the data. Open answers can identify people through the details they contain. A local model, which keeps the data on your own computer, avoids the problem, as the chapter’s run did.

Do it yourself

These tasks have no starter code and no answers. The code space is empty and ready to run; the data and the functions code_by_keywords() and kappa_of() are already loaded.

1. Invent a new open question, such as “What would most improve your supervision?”, and write ten plausible answers of your own. Write a codebook of four or five themes, code the answers yourself, then ask a fellow student (or a language model, in the download project) to code them with your codebook, and calculate the kappa between the two codings. Which themes caused disagreement?

2. In the download project, change one theme definition in the chapter’s codebook, code 20 hand-coded answers with a local model using the new codebook, and compare the agreement with the hand coding before and after the change.

3. Write the paragraph that discloses the use of AI in a thesis of your own: what tools you used, for what (code, language, coding of data), how you checked their output, and where the details can be found. No code is needed; write it in a document.

Work on your own computer

NoteDownload the Chapter 18 project

The project contains the data, the language model’s saved results, and all four parts of this page as an R script, with the ellmer code to code answers with a language model: a free local model through Ollama, or an online model with your own API key (see README.txt). Answers to the questions are in solutions.R, and the open tasks have space to write your code.

  • Download chapter18.zip, unzip it, and double-click chapter18.Rproj.
  • Or type this one line in RStudio’s Console:
usethis::use_course("https://polla-fattah.github.io/data2thesis_r/playground/chapter18.zip")