Playground: Chapter 14

Advanced Clustering

This page practises the ideas of Chapter 14 in four parts. Practise the chapter works through the book’s exercises, with starter code, hints, and solutions. Go further goes beyond the book, with simulations and some of R’s built-in datasets. Check your understanding asks short questions whose answers open with a click. Do it yourself gives open tasks with an empty code space and no answers: you decide what counts as a group, as you will in your own thesis.

The chapter’s mclust package is not available in the browser, so this page works with functions that come with R and two short functions defined in the setup. dbscan_by_hand() runs DBSCAN exactly as the dbscan package does (it gives identical clusters and noise points on this data), and ari() calculates the adjusted Rand index, as mclust’s adjustedRandIndex() does. Exercises that need the mixture model are marked as tasks for the download project, which uses mclust and dbscan exactly as the chapter does; where a mixture model’s profiles are needed on this page, k-means stands in for it.

Practise the chapter

These are the exercises at the end of Chapter 14, with the same numbers. In the setup, profiles holds each student’s averages and scale scores and profile_data the same values as z-scores, exactly as in the chapter.

Exercise 1: Four mixture profiles

Fit Mclust(profile_data, G = 4) and describe the four profiles. Identify which of the three-cluster profiles has been split, and compare the new solution with Chapter 9’s k-means.

This exercise needs mclust, so it is done in the download project.

set.seed(123)
gmm4 <- Mclust(profile_data, G = 4, verbose = FALSE)
profiles |>
  mutate(cluster = gmm4$classification) |>
  summarise(students = n(), across(sleep_hours:satisfaction, ~ round(mean(.x), 1)),
            .by = cluster)
table(three = profile_gmm$classification, four = gmm4$classification)

The four profiles are the disengaged students (201: few study hours, low support and satisfaction), the balanced (200), the overloaded (166: 41 hours of study, 5.7 hours of sleep), and a small extreme group (32) who study 43 hours, sleep 4.9 hours, and take over 560 mg of caffeine a day. The cross-table shows that the extreme group was split off the three-cluster overloaded profile (31 of its 32 students came from it). This is close to Chapter 9’s four k-means clusters, which also separated a very-high-caffeine group; the two four-cluster solutions agree well (adjusted Rand index 0.73).

Exercise 2: Clear members

Count the students with a membership probability above 0.95 for their profile, and report the share of the sample.

This exercise needs mclust, so it is done in the download project.

certainty <- apply(profile_gmm$z, 1, max)
sum(certainty > 0.95)
mean(certainty > 0.95)

432 of the 599 students (72%) belong to their profile with a probability above 0.95; the other 28% sit between profiles, including 79 whose most likely profile has less than an 80% probability. A thesis can report exactly this: “72% of students clearly matched one profile; the rest combined features of two.” That is more honest than a table that assigns every student to a profile with equal confidence.

Exercise 3: DBSCAN with different eps

Run DBSCAN on the profile data with eps of 1.2, 1.6, and 2.4 (and minPts = 8), describe how the number of clusters and noise points change, and explain why a small eps labels so many students as noise.

NoteHint

minPts is the number of students a neighbourhood must hold for its centre to be a core point; the chapter used 8. Noise points are labelled 0.

TipSolution
clusters <- dbscan_by_hand(profile_data, eps = eps, minPts = 8)

With eps = 1.2, DBSCAN finds 4 small clusters and labels 290 students, almost half, as noise. With 1.6, one cluster and 49 noise points; with 2, one cluster and 13; with 2.4, one cluster and 9. In seven dimensions, points are far apart: a radius of 1.2 standard deviations holds few neighbours, so most students fail to be core points, and only the densest pockets form clusters. As eps grows, the pockets join into the one continuous cloud the chapter described. There is no eps at which the profiles appear as separate dense regions, which is itself the finding: the profiles overlap without gaps.

Exercise 4: The profiles without the noise students

Leave out the DBSCAN noise students, fit the mixture model again, and describe whether the profiles change.

This exercise needs mclust, so it is done in the download project.

keep <- dbscan(profile_data, eps = 2, minPts = 8)$cluster != 0
gmm_kept <- Mclust(profile_data[keep, ], G = 3, verbose = FALSE)
profiles[keep, ] |>
  mutate(cluster = gmm_kept$classification) |>
  summarise(students = n(), across(sleep_hours:satisfaction, ~ round(mean(.x), 1)),
            .by = cluster)

Without the 13 noise students, the three profiles are almost unchanged: the same balanced, disengaged, and overloaded groups, with averages that differ by a few tenths at most. The overloaded profile’s average caffeine falls from 284 to 268 mg, because the most extreme caffeine users were among the noise students. The two solutions agree closely on the students they share (adjusted Rand index 0.85). The profiles are robust: they do not depend on a handful of unusual students.

Exercise 5: Mixture model against hierarchical clustering

Use adjustedRandIndex() to compare the three-cluster mixture model with hierarchical clustering (Ward’s method, cut into three clusters) from Chapter 9.

This exercise needs the mixture model, so it is done in the download project. Go further compares clusterings with ari() in the browser.

hierarchical <- cutree(hclust(dist(profile_data), method = "ward.D2"), k = 3)
adjustedRandIndex(profile_gmm$classification, hierarchical)
table(profile_gmm$classification, hierarchical)

The adjusted Rand index is 0.63: clear agreement, far above chance, but far from identical. The cross-table shows where they differ: most disagreements are students whom the mixture model calls disengaged and Ward’s method places with the balanced students (72 of them). k-means agrees with each of the other two about as well (0.70 and 0.64). The three methods find the same broad structure, and disagree about students at the boundaries, which is where the membership probabilities were lowest.

Exercise 6: Profiles and final GPA

Compare the three profiles on final GPA, and say whether the pattern is what their descriptions would lead you to expect. On this page, k-means with three clusters stands in for the mixture model.

NoteHint

The profiles are compared by their average final GPA; students who left before semester 4 have none, hence na.rm = TRUE.

TipSolution
final_gpa = round(mean(final_gpa, na.rm = TRUE), 2)

The balanced profile has the highest final GPA (3.25), the overloaded profile comes next (3.08), and the disengaged profile is lowest (2.95); the mixture model’s profiles in the download project give almost the same (3.27, 3.07, 2.99). Mostly as expected, with one instructive twist: the overloaded students study far more than anyone (42 hours a week) yet finish below the balanced students, who study 24 hours. Chapter 8 found that grades level off beyond about 32 hours of study, and the overloaded students also sleep least and are most stressed. Final GPA was not used to form the profiles, so these differences are evidence that the profiles capture something real. Note that the comparison covers only students who stayed: 36 of the 599 had no final GPA.

Exercise 7: A gap between the groups

In the simulation at the start of the chapter, move the long group upwards (a mean of 3 for y) so that a gap separates it from the round group. Describe how each method’s result changes, and explain why DBSCAN now behaves differently. On this page, k-means and DBSCAN run; the mixture model’s results come from the download project.

NoteHint

The loop runs the simulation twice: with the chapter’s value of 1.4, and with the new mean for the long group.

TipSolution
for (long_y in c(1.4, 3)) {

With the chapter’s layout, k-means scores 0.53 and DBSCAN 0.00: the end of the long group touches the round group, so DBSCAN sees one dense region. With a gap, DBSCAN jumps to 0.98: it finds exactly two clusters, and labels 9 of the 15 scattered points as noise. k-means improves only to 0.79, because it still draws a straight boundary halfway between two centres, cutting off part of the long group. The mixture model scores 0.98 and 1.00 (download project). DBSCAN’s definition of a group, a dense region separated by sparse space, fits data with gaps perfectly and fails completely without them; the moved group changed nothing about the groups themselves, only whether a gap separates them.

Go further

These exercises go beyond the book.

Exercise 8: Membership probabilities by hand

The chapter’s mixture model for Old Faithful found two groups of waiting times: short waits (36% of eruptions, mean 54.6 minutes) and long waits (64%, mean 80.1 minutes), both with a standard deviation of 5.9 minutes. The probability that a wait belongs to the short group is the short group’s share of the total height of the two weighted curves at that wait. Calculate it for waits of 50, 67, and 85 minutes.

NoteHint

dnorm() gives the height of a normal curve; multiplying by each group’s share makes the larger group’s curve taller. The denominator is the total of the two weighted heights.

TipSolution
long <- 0.64 * dnorm(wait, mean = 80.1, sd = 5.9)
round(short / (short + long), 3)

A 50-minute wait belongs to the short group with a probability of 1.000, an 85-minute wait with 0.000, and a 67-minute wait with 0.42, exactly what mclust reports in the chapter. The 67-minute wait lies nearer the short group’s mean (12.4 minutes away, against 13.1 for the long group), yet it is slightly more likely to be long, because long waits are nearly twice as common. That is the whole of a mixture model’s soft assignment: the height of each group’s curve, weighted by its size.

Exercise 9: Groups of different sizes and densities

Simulate three groups: a large, spread-out group of 300, a small, tight group of 30 close to it, and a medium group of 60 further away. Compare k-means and DBSCAN with the true groups.

NoteHint

Inside the loop, eps takes each value in turn; pass it on to dbscan_by_hand().

TipSolution
db <- dbscan_by_hand(groups[, c("x", "y")], eps = eps, minPts = 5)

k-means scores only 0.38: it gives 106 members of the large group to the small group’s cluster, because it places its boundaries halfway between centres and so tends to make clusters of similar size and spread. DBSCAN does not do much better, at best 0.66: with a small eps, the large group’s thin outskirts break into fragments and noise; with a larger eps, the small group merges with the large group it touches. A single density threshold cannot suit groups of different densities. A mixture model, which lets each group have its own size and spread, scores 0.94 on the same data (download project). Groups of unequal size and spread are common in real data, and they are where the choice of method matters most.

Exercise 10: DBSCAN on Old Faithful

Run DBSCAN on both variables of the built-in faithful data (eruption length and waiting time, scaled) with several values of eps, and describe the clusters it finds.

NoteHint

tapply(x, group, f) applies f to the values of x in each group; describe each cluster by its average waiting time.

TipSolution
tapply(faithful$waiting, db, mean)

With eps = 0.1, DBSCAN breaks the data into 8 fragments and 126 noise points; with 0.2 and 0.3, it finds the two kinds of eruption, short eruptions after short waits (about 54 minutes) and long eruptions after long waits (about 80 minutes), with 25 and then only 8 points as noise; with 0.5, the two groups merge into one. Unlike the student profiles, the geyser’s two groups are separated by a sparse gap, so there is a wide range of eps over which DBSCAN finds them clearly. Where the range of good values is wide, the clusters are a feature of the data; where every value gives a different answer, they are a feature of the setting.

Exercise 11: Earthquakes near Fiji

R’s built-in quakes data gives the location of 1,000 earthquakes near Fiji since 1964. Cluster their positions (longitude and latitude, scaled) with k-means into two groups and with DBSCAN, and compare the two.

NoteHint

ari() compares any two clusterings of the same cases: here the k-means clusters and the DBSCAN clusters.

TipSolution
ari(km, db)

The two methods largely agree (adjusted Rand index 0.94): both separate the long band of earthquakes to the east, along the Tonga trench (785 earthquakes), from the smaller group to the west (about 190). DBSCAN adds two things k-means cannot: it labels 15 isolated earthquakes as noise, and it finds a third, tiny cluster of 12 earthquakes at the northern end of the western group, which k-means had to merge into a larger cluster. The plot shows that the eastern cluster is a long, curved band, a shape DBSCAN follows easily; k-means gets it right here only because the two groups are far apart. Geography, not the method, decides what the clusters mean: the bands trace where tectonic plates meet.

Check your understanding

Answer each question in your own words first, then click to see a model answer.

1. What are the three definitions of a group used by k-means, mixture models, and DBSCAN?

k-means defines a group as the cases nearest to a centre, which gives round groups of similar size, with every case in exactly one group. A mixture model defines a group as a probability distribution with its own centre, spread, shape, and size, and gives each case a probability of belonging to each group. DBSCAN defines a group as a dense region separated from others by sparse space, of any shape, and allows cases in sparse regions to belong to no group at all.

2. What is the difference between hard and soft membership?

Hard membership assigns each case to exactly one cluster, with complete certainty, as k-means does. Soft membership gives each case a probability of belonging to each cluster, as a mixture model does, so a case halfway between two clusters is shown as uncertain rather than forced into one. Soft membership lets a thesis report how many cases clearly fit a profile.

3. What does BIC compare when choosing the number of clusters in a mixture model?

It balances how well each model fits the data against how many parameters it uses. More clusters, or more flexible cluster shapes, always fit better, so BIC adds a penalty for every extra parameter, and the model with the best balance wins. It is a guide, not a verdict: models with nearly the same BIC are nearly equally supported, and interpretability should help to choose between them.

4. Why did DBSCAN find only one cluster of students in the chapter?

Because the students form one continuous cloud, with no sparse gaps between the profiles. DBSCAN looks for dense regions separated by sparser space; where the profiles meet, students are as densely packed as anywhere else, so the regions join into one. The result is informative: the profiles describe where students sit within one cloud, not separate islands of students.

5. How can a clustering be judged when there is no answer key?

By several kinds of evidence together: whether the clusters are clearly separated (for example, the silhouette or the membership probabilities); whether they are stable, appearing again in bootstrap samples and with other reasonable methods (the adjusted Rand index measures agreement); whether they are interpretable and useful; and whether they differ on outcomes they were not built from, such as final GPA. No single measure decides; a convincing clustering passes several of these checks.

Do it yourself

These tasks have no starter code and no answers. Each code space below is empty and ready to run: write your own code, as you would for a thesis. The data (profiles and profile_data), the functions dbscan_by_hand() and ari(), and the dplyr package are already loaded, and R’s built-in datasets are always available.

1. Choose a different set of variables for student profiles (for example, first-semester sleep, caffeine, wellbeing, and support). Cluster the students with k-means and with DBSCAN, choose the number of clusters and eps with a reason for each, describe the clusters, and say how many students fit no cluster. (In the download project, fit a mixture model too, and report the share of clear members.)

2. Check the stability of three k-means profiles: draw 20 bootstrap samples of the students, cluster each, and use ari() to compare each bootstrap clustering with the clustering of the full data on the students they share. Report how stable the profiles are. (In the download project, do the same for the mixture model.)

3. Write the clustering paragraph of a methods section for your analysis in Task 1, reporting every choice: the variables and their scaling, the method or methods, how the number of clusters and any settings were chosen, and how the clusters were validated. No code is needed; write it on paper or in a document.

Work on your own computer

NoteDownload the Chapter 14 project

The project contains the data and all four parts of this page as an R script. Its Practise part uses mclust and dbscan exactly as the chapter does, including the exercises that cannot run in a browser; the other parts use base R, as on this page. Answers to the questions are in solutions.R, and the open tasks have space to write your code.

  • Download chapter14.zip, unzip it, and double-click chapter14.Rproj.
  • Or type this one line in RStudio’s Console:
usethis::use_course("https://polla-fattah.github.io/data2thesis_r/playground/chapter14.zip")