Today’s goal
Design a test strategy that finds important failures without coupling every test to implementation details.
We will connect:
- risk, evidence, and the testing pyramid;
- static analysis, unit, component, integration, and E2E tests;
- semantic queries, accessible names, keyboard, and focus behavior;
- async UI, network mocking, cancellation, and races;
- Playwright journeys, visual regression, contracts, and browser matrices;
- offline, cache, routing, permissions, security, and performance tests;
- flakiness, CI selection, coverage, maintenance, and test architecture.
By the end of today you can
- choose a test level from the risk and boundary involved;
- test user-visible behavior without inspecting private state unnecessarily;
- use semantic queries and accessible names effectively;
- distinguish simulated DOM tests from real-browser tests;
- test loading, empty, error, retry, cancellation, and optimistic rollback;
- mock at stable boundaries rather than mocking every internal module;
- design realistic browser journeys and failure artifacts;
- detect flakiness instead of normalizing it;
- use coverage as a map rather than a grade;
- maintain a layered suite that remains useful as the UI evolves.
Testing pyramid, trophy, and reality
The shape is less important than the reasoning:
- fast checks should catch cheap mistakes early;
- integration tests should cover important boundaries;
- E2E tests should protect critical journeys;
- production signals should reveal what test environments missed.
Choose the mix from risk, not from a diagram’s proportions.
Why role-based queries are valuable
screen.getByRole("button", { name: /save/i });Role-based queries:
- reflect the accessibility tree;
- survive many visual refactors;
- encourage meaningful semantics;
- match how assistive technologies identify controls.
They are a quality signal, not complete accessibility certification.
Role queries do not replace accessibility audits
A control can be findable by role and still have:
- poor focus behavior;
- incorrect state announcements;
- bad color contrast;
- keyboard traps;
- confusing reading order;
- missing error association.
Automated checks supplement human and assistive-technology evaluation.
Component test example
it("shows an error when the email is invalid", async () => {
await user.type(screen.getByLabelText("Email"), "not-an-email");
await user.click(screen.getByRole("button", { name: "Save" }));
expect(screen.getByRole("alert")).toHaveTextContent("valid email");
});The test follows interaction and outcome rather than implementation.
Browser simulation versus a real browser
| Simulated environment | Real browser |
|---|---|
| fast and focused | realistic layout and platform behavior |
| easy unit/integration loop | network, focus, CSS, storage, workers |
| incomplete browser APIs | higher cost and setup |
| good for most logic | needed for critical browser behavior |
Use both intentionally.
Demonstration: Failing race condition
// User types rapidly: "erb" then "erbil"
// Without AbortController, Request 1 resolves after Request 2:
test("demonstrates race condition failure", async () => {
render(<LiveSearch />);
await user.type(screen.getByRole("searchbox"), "erb");
await user.type(screen.getByRole("searchbox"), "il");
// If the component lacks cancellation, delayed response for "erb"
// overwrites the newer "erbil" results in the DOM!
expect(screen.getByRole("searchbox")).toHaveValue("erbil");
// FAILS: DOM displays items for "erb" instead of "erbil"
expect(await screen.findByText("Erbil International Airport")).toBeInTheDocument();
});A test that does not control network timing will never detect this intermittent race.
Demonstration: Setting up the race test
Simulate network latency on the first query using MSW:
const heldResolvers: Array<() => void> = [];
server.use(
http.get("/api/search", ({ request }) => {
const q = new URL(request.url).searchParams.get("q");
if (q === "erb") {
// Hold Request 1 until manually released
return new Promise(r => heldResolvers.push(() =>
r(HttpResponse.json([{ id: 1, name: "Old Erbil Entry" }]))
));
}
return HttpResponse.json([{ id: 2, name: "Erbil International Airport" }]);
})
);Demonstration: Asserting race resilience
Trigger rapid typing and verify late responses are ignored:
render(<LiveSearch />);
// User types "erbil" - Request 2 resolves quickly
await user.type(screen.getByRole("searchbox"), "erbil");
expect(await screen.findByText("Erbil International Airport")).toBeInTheDocument();
// Resolve delayed Request 1: must NOT clobber current UI
heldResolvers[0]?.();
expect(screen.queryByText("Old Erbil Entry")).not.toBeInTheDocument();The test controls network timing at the transport boundary.
What the passing race test proves
Passing the mocked component race test validates:
- Order-independence: delayed earlier requests do not overwrite newer state;
- DOM accuracy: rendered search results reflect the active query;
- Cancellation:
AbortControllerdispatches signal on new keystrokes.
The contract between input and view is verified.
What the passing test still cannot prove
Even with the component test green, it cannot prove:
- Backend capacity: server rate-limiting under concurrent queries;
- Screen-reader cadence:
aria-livespeech queue congestion; - Device rendering: frame drops on low-tier mobile hardware;
- Input methods: IME Arabic/CJK composition events.
Confidence requires complementary evidence across layers.
Test architecture smells
Watch for:
- every test queries a test ID;
- every internal function is mocked;
- refactoring markup breaks hundreds of tests;
- E2E suite takes hours;
- failures disappear on retry;
- 100% coverage but critical bugs escape;
- snapshots are approved without review;
- production bugs cannot be reproduced.
These are signals to redesign the evidence strategy.
Practical stages 1–2: pure logic & accessible semantics
- Stage 1 (Pure Logic & Parser Unit Testing):
- Test price formatting, pagination math, and schema parsers against valid, boundary, and corrupt input.
- Pure fast Node runtime without DOM overhead.
- Stage 2 (Component Semantics & Accessible Names):
- Query by
getByRoleandgetByLabelText. - Test keyboard navigation (
Tab,Escape) and focus retention.
- Query by
Verification: Tests depend on platform accessibility contracts, not CSS classes or private component state.
Practical stages 3–4: boundary mocking & fault injection
- Stage 3 (Network Interception with MSW):
- Intercept requests at the HTTP transport boundary.
- Simulate network delays, 500 server crashes, and offline states.
- Stage 4 (Asynchronous Resilience & Fault Injection):
- Inject deliberate faults (dropped
AbortController, broken optimistic rollback, missingaria-invalid). - Confirm tests fail immediately and pinpoint the exact failure mechanism.
- Inject deliberate faults (dropped
Verification: Tests prove loading, empty, error, retry, and cancellation behaviors without arbitrary sleep() delays.
Practical stage 5: Playwright critical-path journey
- Stage 5 (Playwright Critical-Path Browser Journey):
- Run a real headless Chromium/Firefox/WebKit test.
- Exercise the full user journey: search, filter, optimistic order drafting, and error recovery.
- Capture trace artifacts, network waterfalls, and screenshots on failure.
Verification: Fast feedback in CI with zero flakiness; high confidence across real browser layout engines.
Troubleshooting guide (Part 1)
| Symptom | Likely cause |
|---|---|
| Tests break after harmless markup refactor | Assertions depend on private structure |
| Everything uses test IDs | User-facing semantics are missing or ignored |
| Async tests need sleeps | Tests wait for time instead of conditions |
| Mocks hide integration failures | Mock boundary is too deep |
| E2E suite is slow and flaky | Too much setup and shared mutable data |
Troubleshooting guide (Part 2)
| Symptom | Likely cause |
|---|---|
| Retry makes CI green | Flakiness is being normalized |
| Coverage is high but bugs escape | Assertions and risk mapping are weak |
| Visual diffs are always approved | Review policy and baseline ownership are weak |
| Offline test passes only once | Storage and cleanup are not isolated |
Completion checklist
- risks determine the test layers;
- static, unit, component, integration, and E2E roles are clear;
- tests use semantic user-facing queries where appropriate;
- keyboard and focus behavior are covered;
- async states and races are deterministic;
- network mocks sit at stable boundaries;
- critical browser journeys have failure artifacts;
- accessibility and visual checks supplement behavior tests;
- flakiness has ownership and a removal policy;
- coverage and CI selection support, rather than replace, judgment.
Misconceptions to leave behind (Part 1)
| Misconception | Better mental model |
|---|---|
| 100% coverage means well tested | Coverage is a map, not a confidence grade |
| Unit tests are always better | Use the nearest boundary that proves the risk |
| Everything should be E2E | Broad tests are costly and should protect journeys |
| Component tests should inspect state | Test visible behavior and semantics |
| CSS selectors are forbidden | Use the selector that matches the contract |
| Test IDs are bad | They are a legitimate fallback when semantics are unsuitable |
Misconceptions to leave behind (Part 2)
| Misconception | Better mental model |
|---|---|
getByRole certifies accessibility | Queries are one quality signal, not an audit |
| Simulated DOM equals a browser | Real browser behavior needs targeted tests |
| Mocks make tests reliable | Boundary choice and realistic failures matter |
| Sleeps fix async tests | Wait for meaningful conditions |
| Retries solve flaky tests | They can hide nondeterminism |
| More tests always mean higher quality | Better evidence and maintainability matter |