Skip to content

1 - The Modern Web Platform & Browser Internals

Chapter 1: follow a page from resource discovery to rendering and interaction, then test the model in DevTools.

The Modern Web Platform & Browser Internals

From the first HTML bytes to the next interaction

Chapter 1 · Modern Front-End Engineering

Polla Fattah


The files loaded. Why does the page still feel slow?

A catalogue heading appears before its image.

Clicking Load products appears to do nothing - then everything changes together.

What would you inspect first: the requests, the handler, or the rendering work?


What we will explain

  • How the browser discovers resources.
  • Why downloading and executing a script are different events.
  • How document state becomes a frame.
  • Why synchronous code, microtasks, and timers behave differently.
  • What evidence DevTools can - and cannot - provide.

One small page

<link rel="stylesheet" href="styles.css">
<script src="legacy.js"></script>
<script defer src="app.js"></script>
<h1>Products</h1>
<button id="load-products">Load products</button>
<p id="status" role="status"></p>
<img src="hero.webp" alt="Featured products"
     width="800" height="400">
<ul id="products"></ul>

Use the full document in the chapter for the practical.


A browser provides more than a JavaScript engine

The engine executes language code.

The browser also provides networking, the DOM, events, timers, storage, and rendering.

Work can happen concurrently, but ordinary page JavaScript and much rendering work compete for main-thread time.


Navigation has setup costs

For https://shop.example.com/products?page=2:

  1. Resolve the destination if necessary.
  2. Establish or reuse a connection.
  3. Request the document.
  4. Begin processing available HTML.

Caches and restored pages can change this path. A small resource on a new origin may still incur setup cost.


Parsing and downloading overlap

HTML bytes arrive → parsing begins
                          ↓
                 stylesheet discovered → request
                          ↓
                 more markup arrives → continue

The browser does not generally wait for the complete HTML before discovering other resources.

The arrows show dependencies, not measured durations.


Discovery can be the delay

<img src="hero.webp" alt="Featured products">

versus a URL assigned only when JavaScript runs:

const image = new Image();
image.src = "hero.webp";

Unless another reference reveals it earlier, the second request depends on script execution. Assigning src can start the request before insertion into the DOM.


A stylesheet can hide another dependency

HTML → stylesheet → needed background image

The browser needs the CSS before discovering the image through that rule.

Discovery, resource use, and request scheduling are different steps. A declared font that is never used need not be downloaded.


Speculative discovery can look ahead

A preload scanner may inspect available markup while the main parser waits for a script.

It can find a later image or script reference.

It cannot execute arbitrary JavaScript to discover its future results or inspect bytes that have not arrived.

Do not assume a universal browser thread arrangement.


Resource hints have different jobs

HintPurpose
PreloadFetch a current-page resource earlier
PrefetchSpeculate about likely future use
PreconnectBegin connection setup

A hint does not make bandwidth unlimited. Measure whether it addresses an actual delay.


Source HTML is not the live DOM

document.querySelector("#status").textContent = "Loaded";

The DOM changes; the original response does not.

Demonstration: compare the document response in Network with the status paragraph in Elements/Inspector.


Classic scripts can stop the parser

Place this in legacy.js in the head:

console.log(Boolean(document.querySelector("h1")));

The heading has not been parsed: false.

Move the script after the heading: it can find the node.

Node availability does not prove that it has painted.


Choose script timing

For external scripts declared in the initial HTML:

ModeMain execution condition
ClassicParser waits at the script
DeferAfter parsing; ordered with deferred classic scripts
AsyncWhen ready; document order not guaranteed
ModuleDeferred by default; dependency graph applies

async does not mean execution on a background thread.


CSS can delay a script too

<link rel="stylesheet" href="styles.css">
<script src="legacy.js"></script>

A previously encountered stylesheet can block this script.

The parser is waiting for the script, so CSS can indirectly prolong the pause.

CSS blocking presentation and a script blocking parsing are different dependencies.


Readiness is not responsiveness

  • DOMContentLoaded: parsing and relevant deferred script processing have progressed to this milestone.
  • Window load: additional load-delaying resources have completed.
  • Neither guarantees a particular paint time or fast future interactions.

Module experiments here omit imports and top-level await to keep the comparison focused.


From state to a frame

DOM + stylesheet information
              ↓
      Style calculation
              ↓
           Layout
              ↓
    Paint and raster work
              ↓
         Compositing

An update may reuse work or skip stages. This is a conceptual model, not an engine’s thread diagram.


Width and transform do different things

list.style.width = "320px";
list.style.transform = "translateX(20px)";

Width can change wrapping and neighboring geometry.

A transform changes visual placement without reallocating normal-flow space.

Some transforms can reuse painted content. That is not a universal performance guarantee.


A read can force layout

for (const row of rows) {
  row.style.width = "300px";
  widths.push(row.getBoundingClientRect().width);
}

Compare with writing all widths first, then reading them.

Reset the starting state. Batching may reduce repeated layout; the first read can still require it.


Predict the output

console.log("start");
setTimeout(() => console.log("timeout"), 0);
Promise.resolve().then(() => console.log("promise"));
console.log("end");

Write your prediction before running it.


Explain the output

start
end
promise
timeout

Synchronous work finishes first. The already-fulfilled Promise’s reaction runs at a microtask checkpoint before the timer task.

This is not a claim that all asynchronous operations finish before timers.


A task is not followed by a guaranteed paint

Selected task → microtask checkpoint → scheduling continues
                                            ↓
                       rendering when opportunity and scheduling permit

requestAnimationFrame() schedules a callback for a rendering update, not an after-paint notification.

An expensive callback remains expensive.


Why might “Working…” never appear?

Inside the click handler:

status.textContent = "Working…";
const start = performance.now();
while (performance.now() - start < 200) {
  // Bounded demonstration only.
}
status.textContent = "Finished";

Both DOM changes occur. The intermediate state may never reach the screen.


Microtasks are not background work

A microtask can queue another microtask.

The checkpoint keeps processing them, delaying later tasks if the chain grows.

Use the bounded chain in the chapter. Do not freeze the page with an endless loop.

Moving expensive work into a Promise callback does not make it free.


DevTools demonstration: discovery

  1. Serve the chapter’s small page over local HTTP.
  2. Record browser, cache, and throttling settings.
  3. Capture the image request from initial markup.
  4. Remove that reference; assign the same URL after a timer.
  5. Compare initiators and request start times in separate runs.

Explain differences; do not expect identical millisecond values.


DevTools demonstration: execution and rendering

  1. Record a click on the bounded slow handler.
  2. Locate its execution in a performance trace.
  3. Compare the DOM update with the visible result.
  4. Reset the page before comparing layout experiments.

A breakpoint changes timing. A console log proves execution, not paint.


Evidence before optimization

ObservationNext question
Image request starts lateWhat revealed its URL?
Click response is delayedWhat occupied the main thread?
Layout repeatsAre writes and geometry reads interleaved?
Trace looks differentDid cache, viewport, or throttling change?

An unexplained trace is a reason to investigate, not to invent a guaranteed sequence.


Practical 01

Produce predictions, comparable observations, and an explanation of one surprise.

The core exercise covers discovery, script timing, scheduling, and rendering work.

Fixed-height virtualization is optional; deeper profiling belongs with Chapter 15.

Open the practical


Check your understanding

  • Why can a downloaded script still be waiting?
  • What does finding a DOM node tell you about paint?
  • Why does a zero-delay timer run later?
  • Which evidence would distinguish loading delay from execution cost?

Next: what should the document mean?

Browser behavior explains how a page runs.

Chapter 2 examines how semantic structure, accessible controls, language, and direction describe the interface people use.

Read Chapter 1 · Read Chapter 2

2 - Semantic HTML, Accessibility, Internationalization & the DOM

Chapter 2: choose meaningful HTML, build accessible interfaces, respect language and direction, and work with the DOM safely.

Semantic HTML, Accessibility, Internationalization & the DOM

Meaningful structure, inclusive interaction, and the browser’s live document tree

Chapter 2

Polla Fattah


Today’s goal

Build a mental model in which one document is shared by:

  • the browser;
  • keyboard users;
  • assistive technologies;
  • JavaScript;
  • search and other software.

The goal is not to memorise tags. It is to choose structure, behaviour, language, and DOM operations deliberately.


By the end of today you can

  • choose HTML for meaning and behaviour, not default appearance;
  • structure a page with headings and landmarks;
  • explain accessible names, focus order, labels, and errors;
  • use native controls before inventing ARIA widgets;
  • distinguish language from writing direction;
  • query, traverse, create, and update DOM nodes safely;
  • explain event propagation and delegation;
  • describe where Web Components fit without replacing semantic HTML.

The chapter’s single model

flowchart TD
    A[Semantic HTML] --> B[Native browser behaviour]
    B --> C[Accessibility representation]
    C --> D[Language and direction]
    D --> E[DOM tree and events]
    E --> F[Web Components]

These are not separate tricks. They are different views of the same interface model.


A running example: a public-service portal

We will imagine a multilingual service interface with:

  • a site header and primary navigation;
  • a page explaining one service;
  • a request form;
  • a table of existing requests;
  • a status filter;
  • Arabic and English content;
  • a small custom element where a reusable boundary is useful.

The interface will remain understandable before JavaScript enhances it.


HTML describes meaning, not appearance

These fragments may look similar after CSS:

<div class="heading">Service Requests</div>
<div class="navigation">
  <div class="link">Home</div>
  <div class="link">Requests</div>
</div>

But the browser cannot infer the intended roles reliably.


Prefer the meaningful structure

<h1>Service Requests</h1>

<nav aria-label="Primary">
  <a href="/">Home</a>
  <a href="/requests">Requests</a>
</nav>

<main>
  <!-- the page's primary content -->
</main>

The second version identifies headings, navigation, links, and main content without requiring a stylesheet or a JavaScript guess.


One durable rule

Choose HTML according to meaning and behaviour before styling it according to appearance.

HTML answers:

What is this?

CSS answers:

How should it look?

JavaScript answers only the behaviour that needs application logic.


Build a document hierarchy

<body>
  <header>...</header>
  <nav aria-label="Primary">...</nav>
  <main>
    <h1>Citizen Services</h1>
    <section>
      <h2>Request a document</h2>
    </section>
  </main>
  <footer>...</footer>
</body>

The hierarchy should make sense if all styling disappears.


Headings are content structure

Use heading levels to represent relationships:

<h1>Citizen Services</h1>
  <h2>Request a document</h2>
    <h3>Required information</h3>
  <h2>Existing requests</h2>

Do not choose <h3> because its default font size looks convenient.


There is no automatic outline rescue

Do not assume that nested sections automatically create a perfect heading outline for every tool.

Make the heading hierarchy explicit:

  • one clear page-level <h1>;
  • meaningful <h2> regions;
  • <h3> only when a subsection belongs to the preceding <h2>;
  • no skipped levels just for visual size.

Use CSS for size. Use headings for structure.


Landmarks give the page a map

ElementMeaning
<header>Introductory content for the page or section
<nav>A group of navigation links
<main>The page’s primary content
<footer>Closing information for the page or section
<aside>Related content, not the main flow

Landmarks help users and tools move through a large interface.


<main> and <nav>

There should normally be one primary <main> for the page.

Give multiple navigation regions useful names:

<nav aria-label="Primary">...</nav>
<nav aria-label="Service sections">...</nav>
<nav aria-label="Breadcrumb">...</nav>

The label distinguishes regions that would otherwise all be announced as navigation.


These elements are contextual:

<article>
  <header><h2>Request status</h2></header>
  <p>...</p>
  <aside>Last updated yesterday</aside>
  <footer>Source: service desk</footer>
</article>

They are meaningful when their content has the corresponding relationship; they are not decorative wrappers to use everywhere.


<section> and <article>

Use <section> for a thematic region that normally has a heading.

Use <article> for content that could stand on its own or be reused:

<section aria-labelledby="requests-title">
  <h2 id="requests-title">Existing requests</h2>
  <article>Request #1042</article>
  <article>Request #1043</article>
</section>

If a generic container has no semantic purpose, <div> is the honest choice.


Lists carry relationships

Use list elements when the content is a list:

<ul>
  <li>Proof of identity</li>
  <li>Proof of address</li>
</ul>

Do not build a list from paragraphs and line breaks merely because CSS will make it look like one.


Tables carry data relationships

<table>
  <caption>Existing service requests</caption>
  <thead>
    <tr><th scope="col">Reference</th><th scope="col">Status</th></tr>
  </thead>
  <tbody>
    <tr><td>1042</td><td>In review</td></tr>
  </tbody>
</table>

The caption and header relationships make the data understandable outside the visual grid.


Figures, captions, and media

<figure>
  <img src="map.png" alt="Service centres in Baghdad and Basra">
  <figcaption>Available service centres.</figcaption>
</figure>

An image’s alternative text describes its purpose in context. Decorative images may use an empty alt=""; omitting alt is not the same decision.


Native controls are behaviour with a contract

Native elements bring established behaviour:

<button type="button">Show request details</button>
<a href="/requests/1042">Open request 1042</a>
<input type="checkbox" name="urgent" id="urgent">
<details><summary>More information</summary>...</details>

The browser, keyboard, accessibility APIs, and form system already understand these controls.


<!-- Changes state on this page -->
<button type="button" data-action="filter">Filter</button>

<!-- Takes the user to another resource -->
<a href="/requests/1042">View request</a>

Do not use a clickable <div> when a button or link expresses the real job.


The clickable <div> problem

This creates a visual imitation, not a complete control:

<div class="button" onclick="openDialog()">Open</div>

It lacks, unless you recreate them correctly:

  • keyboard activation;
  • focusability;
  • the button role;
  • disabled semantics;
  • expected browser behaviour;
  • a reliable accessible name.

Start with <button> and style it.


The accessibility tree is not the DOM tree

The browser uses the DOM and other information to expose an accessibility representation.

Assistive technology may receive:

button: Submit request
heading level 1: Citizen Services
navigation: Primary
textbox: Email address

Good structure gives the browser useful information to expose.


Accessible names answer “what is this?”

Controls need names users can perceive through their chosen interface:

<label for="email">Email address</label>
<input id="email" name="email" type="email">

Visible labels are usually the strongest starting point. Placeholder text is a hint and should not replace a label.


Name icon-only controls deliberately

<button type="button" aria-label="Close request details">
  <span aria-hidden="true">×</span>
</button>

The icon is visual decoration. The accessible name communicates the action.

Prefer visible text when space allows; it helps every user, not only screen reader users.


Keyboard interaction is part of the interface

Ask these questions for every interactive feature:

  • Can the user reach it with the keyboard?
  • Is the focus order logical?
  • Is the focus indicator visible?
  • Can the user operate it without a pointer?
  • Does focus move somewhere sensible after a state change?

Accessibility is interaction design, not a final inspection step.


Focus order follows the interface’s meaning

Avoid using positive tabindex values to force an artificial order.

Prefer:

  • meaningful document order;
  • native controls;
  • a visible focus style;
  • a small, deliberate focus-management rule only when a component requires it.

If the DOM order is wrong, CSS reordering does not repair the reading or keyboard model.


Forms are an accessibility system

<form>
  <label for="reference">Reference number</label>
  <input id="reference" name="reference" required>
  <button type="submit">Find request</button>
</form>

The form needs a relationship between label, control, name, value, validation, and error feedback.


Native validation is useful information

<label for="email">Email address</label>
<input id="email" name="email" type="email" required>

Native constraints can provide a baseline. If JavaScript adds custom errors, it must preserve a clear message, association, and recovery path.

Do not make invalid state visible only through colour.


<fieldset>
  <legend>Preferred contact method</legend>
  <label><input type="radio" name="contact" value="email"> Email</label>
  <label><input type="radio" name="contact" value="phone"> Phone</label>
</fieldset>

fieldset and legend communicate the group relationship that a visual box alone cannot provide.


Associate errors with the control

<label for="phone">Phone number</label>
<input id="phone" aria-describedby="phone-error" aria-invalid="true">
<p id="phone-error">Enter a phone number including the area code.</p>

The error should explain what is wrong and how to fix it. Do not only add a red border and expect the user to infer the problem.


ARIA adds semantics only when HTML cannot

ARIA can describe a custom interaction, but it does not automatically create the keyboard behaviour, focus management, state changes, or event handling.

flowchart TD
    A["First choice: Native HTML"] --> B["Second choice: Native HTML + enhancement"]
    B --> C["Last choice: Custom ARIA widget with full contract"]

ARIA can make good HTML worse

Do not add roles that contradict the element’s real meaning:

<button role="heading">Submit</button>

The role does not repair a confused design. Use the element that expresses the job, then add only the missing semantic information.


Language begins in the document

<html lang="en">

For a passage in another language:

<p lang="ar">مرحباً بكم في خدمات المواطنين</p>

Language metadata supports pronunciation, hyphenation, spell checking, translation, search, and assistive technology decisions.


Language and direction are different

lang identifies the language. dir identifies the writing direction.

<p lang="ar" dir="rtl">خدمات المواطنين</p>
<p lang="en" dir="ltr">Citizen services</p>

Arabic does not mean every surrounding interface region must be permanently right-aligned. Direction belongs to the text and layout context.


Direction can be automatic

<p dir="auto">User-entered text appears here.</p>

dir="auto" lets the browser infer direction from the first strong character. It is useful for unknown user content, but explicit direction is better when the application knows the language and layout contract.


Bidirectional text needs isolation

User-generated text can contain characters with different directionality.

<p>
  Ticket reference: <bdi>{{ reference }}</bdi>
</p>

Use <bdi> when embedded user text should not disturb the surrounding order. Use <bdo dir="rtl">...</bdo> only when deliberately overriding direction.


RTL is not “mirror everything”

RTL-aware design considers:

  • text direction;
  • logical properties such as margin-inline-start;
  • icon meaning;
  • navigation order;
  • tables and numbers;
  • mixed-language content;
  • focus and reading order.

Do not replace every left value with right and call the interface localized.


The DOM is a live tree

HTML is the source representation. The DOM is the browser’s live object tree.

flowchart TD
    Doc[Document] --> HTML[html]
    HTML --> Head[head]
    HTML --> Body[body]
    Body --> Header[header]
    Body --> Main[main]

JavaScript reads and changes this live tree. It does not edit the original Markdown or HTML source file.


The DOM contains different node types:

  • document nodes;
  • element nodes;
  • text nodes;
  • comment nodes.

An element is a node, but not every node is an element. APIs such as children, childNodes, parentElement, and textContent expose different views of that tree.


Query the DOM precisely

const form = document.querySelector('[data-service-form]');
const fields = form.querySelectorAll('input, select, textarea');

Use stable, meaningful selectors. Check whether a query can return null, and avoid assuming that a selector matched the element you intended.

querySelector() returns the first match. querySelectorAll() returns a static NodeList of all matches.


Traverse relationships deliberately

const row = event.target.closest('tr[data-request-id]');
const requestId = row?.dataset.requestId;

Useful relationships include:

  • parentElement;
  • children;
  • nextElementSibling;
  • closest();
  • matches().

Traversal should express the component boundary, not depend on fragile visual nesting.


Create and insert elements safely

const message = document.createElement('p');
message.textContent = userMessage;
message.className = 'status-message';
container.append(message);

Prefer textContent for untrusted text. Do not use innerHTML as a convenient string formatter when the value may contain user or remote data.


Attributes and properties are connected, not identical

checkbox.setAttribute('aria-checked', 'true');
checkbox.checked = true;

Attributes are markup-facing values. Properties are live JavaScript-facing state. Boolean controls make the distinction especially visible:

checkbox.defaultChecked // initial value
checkbox.checked        // current live value

Events connect users to the DOM

button.addEventListener('click', (event) => {
  submitRequest();
});

Events carry an interaction through a tree. Understanding that path is more reliable than attaching random handlers until the interface appears to work.


Event propagation has phases

flowchart LR
    subgraph Capture["1. Capture Phase"]
        W1[window] --> D1[document] --> A1[ancestors]
    end
    subgraph Target["2. Target Phase"]
        A1 --> T[target element]
    end
    subgraph Bubble["3. Bubble Phase"]
        T --> A2[ancestors] --> D2[document] --> W2[window]
    end

The event can be observed at different points in the path. This is why a parent can respond to an interaction that occurred on a child.


target and currentTarget differ

list.addEventListener('click', (event) => {
  console.log(event.target);        // where it started
  console.log(event.currentTarget); // this listener's element
});

Confusing these values is a common source of bugs in delegated interfaces.


Default actions and propagation are different

link.addEventListener('click', (event) => {
  event.preventDefault(); // cancel navigation
  event.stopPropagation(); // stop this event moving farther
});

stopPropagation() does not cancel the browser’s default action. Use each operation only for the problem it solves.


Event delegation scales repeated UI

table.addEventListener('click', (event) => {
  const button = event.target.closest('[data-action="cancel"]');
  if (!button) return;
  cancelRequest(button.dataset.requestId);
});

One listener can handle current and future matching descendants. The handler must still validate the target and preserve keyboard-accessible controls.


Put the public-service interface together

<html lang="en">
  <body>
    <header>...</header>
    <nav aria-label="Primary">...</nav>
    <main>
      <h1>Citizen Services</h1>
      <section aria-labelledby="form-title">...</section>
      <section aria-labelledby="table-title">...</section>
    </main>
    <footer>...</footer>
  </body>
</html>

Now add language metadata, labels, native controls, meaningful table headers, and safe DOM updates before adding custom interaction.


Web Components extend HTML

Web Components provide browser-native mechanisms for reusable boundaries:

  • Custom elements define a new element name;
  • Shadow DOM can isolate internal markup and styles;
  • Slots allow controlled content insertion.

They are tools for encapsulation, not a reason to discard document meaning.


A custom element still needs a contract

class RequestStatus extends HTMLElement {
  connectedCallback() {
    this.textContent = this.getAttribute('status') ?? 'Unknown';
  }
}

customElements.define('request-status', RequestStatus);

The element needs clear inputs, output meaning, lifecycle behaviour, and an accessibility story. Encapsulation does not make an interface accessible by itself.


Shadow DOM is not a security boundary

Shadow DOM can hide implementation details and scope styles.

It does not automatically provide:

  • security isolation;
  • an accessible name;
  • keyboard behaviour;
  • correct focus management;
  • a good component API.

Treat it as an encapsulation mechanism, not a trust boundary.


Slots preserve controlled composition

<request-card>
  <span slot="status">In review</span>
</request-card>

Slots allow consumers to provide content to named insertion points. The component still owns the surrounding semantics and must define how slotted content participates in the accessible interface.


The component should not erase the document

Good component design preserves:

  • meaningful HTML where native elements already work;
  • accessible names and labels;
  • predictable keyboard behaviour;
  • language and direction metadata;
  • a clear DOM and event contract.

Web Components are an extension point, not a replacement for semantic HTML.


Misconceptions to leave behind (Part 1)

MisconceptionBetter model
“If it looks like a heading, it is a heading.”Meaning comes from the element and structure.
“A <div> can replace any element.”Native elements bring behaviour and semantics.
“ARIA makes custom controls accessible.”ARIA describes; code must implement interaction.
“Placeholder text is a label.”Use a real label and associate it with the control.

Misconceptions to leave behind (Part 2)

MisconceptionBetter model
“RTL means align everything right.”Direction, layout, icons, and mixed text need separate decisions.
“DOM means HTML.”HTML is source; the DOM is the live runtime tree.
“stopPropagation() prevents navigation.”Default actions and propagation are different.
“Shadow DOM is security.”Shadow DOM is encapsulation, not isolation.

The practical lab

Build a semantic multilingual service interface.

The runnable code will be in the separate Playground repository. This deck provides the learning sequence and the current repository provides the instructions.


Practical stages 1–3: structure and meaning

  1. Create the document hierarchy with header, nav, main, and footer.
  2. Add meaningful headings, sections, articles, lists, a figure, and a table.
  3. Build a request form with labels, native controls, a fieldset, and a legend.

Check that the page still communicates its structure without CSS.


Practical stages 4–6: interaction and language

  1. Test the interface with keyboard-only navigation.
  2. Add English and Arabic content with correct lang and dir values.
  3. Inspect the accessibility representation and repair missing names or relationships.

Do not treat a passing visual inspection as an accessibility test.


Practical stages 7–10: DOM and components

  1. Query and update the DOM with safe text insertion.
  2. Add event delegation for repeated request rows.
  3. Create a small Web Component with an explicit contract.
  4. Draw the final architecture and identify what remains native HTML.

The exercise is complete when the component boundary is explainable, not merely when the screen looks correct.


Try it yourself

Choose one interface feature and review it in four views:

  1. What does the raw HTML say?
  2. What does the keyboard user experience?
  3. What name and role does assistive technology receive?
  4. What does the DOM event path do when the user interacts?

Write one improvement that helps at least two of these views at once.


Troubleshooting questions

SymptomFirst question
The control cannot be reachedIs it a native interactive element?
The label is not announcedIs the label associated with the control?
The Arabic text changes layout unexpectedlyAre language and direction declared at the right boundary?
The delegated handler does nothingWhat are target and currentTarget?
Text displays unexpectedly as markupDid code use innerHTML where textContent was needed?
The custom element is confusingWhat is its semantic and accessibility contract?

Change one thing at a time and inspect the actual DOM after each change.


Completion check

  • I can choose semantic elements without relying on their default styles.
  • I can describe the page using landmarks and a heading hierarchy.
  • Every form control has a meaningful accessible name.
  • I can use the keyboard through the interface and see focus.
  • I can distinguish lang from dir and handle mixed-direction text.
  • I can query and update the DOM without unsafe string insertion.
  • I can explain event propagation, default actions, and delegation.
  • I can describe what Web Components add and what they do not guarantee.
  • I completed the Chapter 2 practical in the companion Playground.

Next: Modern CSS Architecture & Layout Systems

Chapter 3: the cascade, layout systems, responsive design, container queries, layers, design tokens, and the difference between visual flexibility and structural chaos.

Modern Front-End Engineering

3 - Modern CSS Architecture & Layout Systems

Chapter 3: make the cascade predictable, compose responsive layouts, support internationalized interfaces, and choose a maintainable styling architecture.

Modern CSS Architecture & Layout Systems

From cascade decisions to responsive, internationalized component systems

Chapter 3

Polla Fattah


Today’s goal

Stop treating CSS as a collection of visual fixes.

Today we will use CSS as:

  • a precedence system;
  • a value and token system;
  • a layout system;
  • a responsive system;
  • an internationalization system;
  • and an architecture for change.

By the end of today you can

  • explain why a CSS declaration wins or loses;
  • use cascade layers to control ownership;
  • distinguish raw tokens from semantic tokens;
  • choose normal flow, Flexbox, Grid, or positioning intentionally;
  • use intrinsic sizing, minmax(), clamp(), and subgrid;
  • distinguish viewport responsiveness from container responsiveness;
  • build direction-independent layouts with logical properties;
  • use modern selectors without creating specificity traps;
  • compare plain CSS, CSS Modules, utility CSS, and CSS-in-JS by problem;
  • design a responsive multilingual catalogue without JavaScript measurement.

CSS is a system, not decoration

A dashboard must answer architectural questions:

  • What happens when the page narrows?
  • What happens when a card moves into a narrow sidebar?
  • How do cards align when descriptions have different lengths?
  • How do long translations affect the layout?
  • Which rule wins when a library and the application disagree?
  • How can one design decision update the whole interface?

These are system-design questions expressed through CSS.


The chapter’s progression

flowchart TD
    A[Cascade & Precedence] --> B[Values & Tokens]
    B --> C[Layout Systems]
    C --> D[Responsive Design]
    D --> E[Internationalized Layout]
    E --> F[Modern Selectors]
    F --> G[Styling Architecture]

The goal is not to memorise properties. It is to make CSS behaviour predictable enough to change safely.


The cascade is the foundation

<button class="action primary">Submit request</button>
button  { background: gray; }
.action { background: blue; }
.primary { background: green; }

The cascade decides which declaration supplies the final value.

The answer is not always “the last rule”.


Why declarations compete

Candidate declarations may differ by:

  1. origin and importance;
  2. cascade layer;
  3. specificity;
  4. scope or proximity where relevant;
  5. source order.

The useful question is:

Why did this declaration win?

That question is more valuable than memorising a score.


Browser, user, and author styles

The page starts with more than our stylesheet:

  • the browser supplies user-agent styles;
  • users may supply preferences or styles;
  • the author supplies application styles.

An unstyled <h1> is already bold, large, and separated from nearby text.

Author CSS is not the only styling authority. This matters especially for accessibility preferences and !important declarations.


Inheritance is an architectural tool

body {
  color: #222;
  font-family: system-ui, sans-serif;
}

Descendant text normally receives these values without repeating them on every element.

Properties such as color and font-family commonly inherit. Properties such as margin, padding, border, and width normally do not.

Use inheritance for broad design decisions, then override at meaningful boundaries.


Specificity is not a quality ranking

.card h2 { color: navy; }

#dashboard .card h2.title {
  color: purple;
}

The second selector is harder to override, but that does not make it better.

High specificity increases the cost of future changes.

Prefer a selector that is easy to reason about and easy to replace.


Do not reduce specificity to arithmetic alone

Specificity is one stage in the cascade, not the entire cascade.

Before comparing selectors, ask:

  • Are both declarations in the same origin?
  • Are they in the same layer?
  • Is one declaration important?
  • Is a scoped rule involved?
  • Only then: which selector is more specific?

The browser is resolving a precedence system, not grading CSS style.


Source order is the final tie-breaker

.notice { color: blue; }
.notice { color: green; }

When the earlier cascade stages tie, the later declaration wins.

Source order is useful for intentional overrides. It is dangerous when a stylesheet becomes a sequence of unexplained patches.


Cascade layers make ownership explicit

@layer reset, base, components, utilities;

@layer reset { /* normalize browser differences */ }
@layer base { /* element defaults and typography */ }
@layer components { /* cards, forms, navigation */ }
@layer utilities { /* small explicit overrides */ }

The layer order is declared once. A later layer can win without requiring an increasingly specific selector.


Layer precedence comes before specificity

@layer base {
  #app .button { color: blue; }
}

@layer components {
  button { color: green; }
}

The simple selector in the later components layer can win over the highly specific selector in the earlier base layer.

This lets architecture control precedence before selector complexity grows.


A practical layer architecture

@layer reset, tokens, base, components, utilities, overrides;

Possible ownership:

LayerResponsibility
resetpredictable browser baseline
tokenscustom properties and theme values
baseelements, typography, document defaults
componentsreusable UI boundaries
utilitiessmall intentional helpers
overridesexplicit application exceptions

The names matter less than the stable ownership contract.


Custom properties participate in the cascade

:root {
  --color-accent: #1464a0;
}

.button {
  background: var(--color-accent);
}

Custom properties are live CSS values. They inherit, cascade, and can be overridden by component or theme boundaries.

They are more powerful than text replacement.


Custom properties can represent decisions

.card {
  --card-gap: 1rem;
  gap: var(--card-gap);
}

.card--compact {
  --card-gap: 0.5rem;
}

The component keeps one layout rule. A meaningful boundary changes the value.

This is different from searching and replacing every 1rem in a project.


Raw tokens and semantic tokens

Raw token:

--blue-600: #1464a0;

Semantic token:

--color-action-primary: var(--blue-600);

Components should usually consume semantic meaning. If the brand palette changes, the component should not need to know the replacement colour.


Token ownership has direction

flowchart TD
    A[Raw palette] --> B[Semantic meaning]
    B --> C[Component usage]
    C --> D[Page composition]

Avoid letting a component reach backward into arbitrary raw palette values.

Stable semantic names reduce the blast radius of visual change.


Theming with custom properties

:root {
  --surface-page: #ffffff;
  --text-primary: #172a42;
}

[data-theme="dark"] {
  --surface-page: #0c2238;
  --text-primary: #f5f7fa;
}

body {
  background: var(--surface-page);
  color: var(--text-primary);
}

The component stays attached to meaning. The theme changes the values.


Start with normal flow

Normal flow is the layout system already provided by the browser.

<main>
  <h1>Catalogue</h1>
  <p>Results appear below the heading.</p>
  <ul>...</ul>
</main>

Let content determine height and let blocks participate in the document before reaching for positioning.


Positioning changes the relationship

.badge {
  position: absolute;
  inset-block-start: 0.5rem;
  inset-inline-end: 0.5rem;
}

Absolute positioning can be correct for a badge anchored to a card. It is a poor replacement for a page layout system when content can grow or translate.

Use fixed and sticky positioning with equally explicit viewport and scrolling assumptions.


Flexbox is one-dimensional

Use Flexbox when items are primarily arranged along one main axis:

.toolbar {
  display: flex;
  align-items: center;
  justify-content: space-between;
  gap: 1rem;
}

The main axis follows flex-direction. The cross axis is perpendicular to it.


Main axis and cross axis

.toolbar {
  display: flex;
  flex-direction: row;
}

For a row:

  • the main axis is inline/horizontal in the usual LTR case;
  • the cross axis is block/vertical.

Use justify-content for main-axis distribution and align-items for cross-axis alignment. Always consider writing direction before describing an axis as “left” or “right”.


Flex sizing is not just alignment

.toolbar__search {
  flex: 1 1 20rem;
  min-inline-size: 0;
}

Flex sizing considers:

  • the flex basis;
  • grow and shrink factors;
  • available free space;
  • minimum sizes;
  • intrinsic content size.

min-inline-size: 0 can be necessary when a flexible child must be allowed to shrink rather than force overflow.


Grid is two-dimensional

.catalogue {
  display: grid;
  grid-template-columns: repeat(3, minmax(0, 1fr));
  gap: 1.25rem;
}

Grid describes rows and columns together. It is a strong choice when the relationship between both axes matters.


Grid tracks express constraints

.catalogue {
  grid-template-columns:
    repeat(auto-fit, minmax(min(100%, 16rem), 1fr));
}

This says:

  • create as many tracks as fit;
  • do not make a card narrower than its useful minimum;
  • share remaining space between tracks.

The layout responds to available space without a device list.


Intrinsic sizing asks the content

.label {
  inline-size: fit-content;
}

Useful intrinsic concepts include:

  • min-content: the smallest size without avoidable breaking;
  • max-content: the size needed by the unwrapped content;
  • fit-content(): a bounded content-aware size.

Intrinsic sizing is especially important when translations and user content change the amount of text.


Alignment is a separate decision

.card-grid {
  display: grid;
  place-items: start stretch;
}

Do not mix up:

  • track sizing;
  • item alignment inside a track;
  • content distribution across tracks.

The right property depends on which relationship you are trying to control.


Grid or Flexbox?

QuestionPrefer
Is the layout mainly one axis?Flexbox
Are rows and columns both meaningful?Grid
Does content define a small toolbar?Flexbox
Does the page define a shared card matrix?Grid
Does one element need to distribute remaining space?Flexbox

They are complementary systems, not competing religions.


Nested Grid has a boundary

A nested grid creates a new grid context:

.card-list { display: grid; }
.card { display: grid; }

The inner card does not automatically inherit the outer grid’s tracks. This is often correct, but it means related content may fail to align across siblings.


subgrid shares tracks intentionally

.cards {
  display: grid;
  grid-template-columns: repeat(3, 1fr);
}

.card {
  display: grid;
  grid-template-rows: subgrid;
  grid-row: span 3;
}

subgrid allows descendants to participate in the parent grid’s track sizing. It is useful when card headings, descriptions, and actions should align across cards with different content lengths.


Do not use subgrid automatically

Use it when shared track alignment is a real requirement.

Avoid it when:

  • cards have independent internal structure;
  • the extra coupling makes the component harder to reuse;
  • a simpler normal flow is sufficient;
  • the alignment is only decorative.

Layout power should follow a demonstrated relationship.


Responsive design is not a device list

Avoid designing only for:

flowchart LR
    A[Phone] --> B[Tablet] --> C[Desktop]

Real interfaces encounter:

  • split-screen windows;
  • zoom;
  • long translations;
  • embedded cards;
  • sidebars;
  • accessibility text size changes;
  • unusual aspect ratios.

Responsive design responds to constraints, not product labels for devices.


Begin with a fluid layout

.page {
  inline-size: min(100% - 2rem, 72rem);
  margin-inline: auto;
}

.title {
  font-size: clamp(1.75rem, 4vw, 3.5rem);
}

Let the interface adapt continuously before adding a breakpoint.


Media queries answer environment questions

@media (width < 48rem) {
  .site-nav {
    display: none;
  }
}

Media queries respond to the viewport or other environment features:

  • viewport width;
  • colour scheme;
  • pointer precision;
  • reduced motion;
  • contrast preferences;
  • print.

They are not the only responsive tool.


Container queries answer component-space questions

.catalogue-shell {
  container: catalogue / inline-size;
}

@container catalogue (width < 36rem) {
  .catalogue-toolbar {
    flex-direction: column;
  }
}

The component responds to the space it actually receives, whether that space comes from the page, a sidebar, or a nested layout.


Media queries and container queries differ

QuestionTool
Is the viewport narrow?Media query
Is this component’s parent narrow?Container query
Does the user prefer reduced motion?Media query
Should a card switch from horizontal to vertical?Container query

Container queries do not replace media queries. They solve a different scope of responsive reasoning.


Name containers when the relationship matters

.results-panel {
  container: results / inline-size;
}

@container results (width > 42rem) {
  .result-card { grid-template-columns: 1fr auto; }
}

Names make the component’s dependency explicit and avoid accidental reliance on an unrelated ancestor container.


Container query units

.card-title {
  font-size: clamp(1rem, 3cqi, 1.6rem);
}

Container-relative units include:

  • cqw and cqh;
  • cqi and cqb for logical axes;
  • cqmin and cqmax.

Use them when the component’s typography or spacing should track its container.


Responsive typography needs limits

h1 {
  font-size: clamp(2rem, 5vw, 4.5rem);
  line-height: 1.05;
}

clamp(min, preferred, max) prevents a fluid value from becoming unusably small or excessively large.

Readable text is a layout constraint, not a decorative afterthought.


Modern viewport units

Some mobile viewport units distinguish the small, dynamic, and large viewport:

.app-shell {
  min-block-size: 100dvb;
}

Choose viewport units according to whether browser UI changes should affect the layout. Do not assume every mobile viewport is a stable 100vh rectangle.


Responsive images belong to layout

<img
  src="catalogue-small.webp"
  srcset="catalogue-small.webp 640w,
          catalogue-large.webp 1280w"
  sizes="(width < 48rem) 100vw, 50vw"
  alt="A service catalogue dashboard">

Image dimensions, aspect ratios, loading behaviour, and object fitting affect layout stability and performance together.


CSS has logical axes

Prefer logical properties when the interface can change direction:

.card {
  margin-inline: auto;
  padding-block: 1rem;
  padding-inline: 1.25rem;
  border-inline-start: 0.25rem solid var(--color-accent);
}

This expresses the relationship instead of hard-coding physical left and right values.


Inline and block are relationships

flowchart TD
    subgraph Inline["Inline Axis"]
        I[Direction text progresses: horizontal LTR/RTL or vertical]
    end
    subgraph Block["Block Axis"]
        B[Direction blocks stack: perpendicular to inline axis]
    end

In horizontal English text these often resemble horizontal and vertical. In RTL or vertical writing modes, the physical interpretation changes.

Logical CSS keeps the component contract stable.


Direction-independent components

.toolbar {
  display: flex;
  gap: 1rem;
  margin-inline-start: auto;
}

.status {
  border-inline-start: 0.3rem solid var(--status-color);
  padding-inline-start: 0.75rem;
}

Test the component with different dir values rather than copying the whole stylesheet and swapping every physical property.


Some things should not mirror

Direction-aware layout does not mean every visual symbol flips.

Consider separately:

  • text and reading order;
  • navigation arrows;
  • media-play controls;
  • brand marks;
  • charts and geographic maps;
  • numbers and code.

The correct decision follows meaning, not a blanket mirror operation.


Modern selectors express relationships

CSS can now describe relationships more directly:

/* Group alternatives without repeating a block */
:is(.primary, .secondary) { border-radius: 0.5rem; }

/* Match without adding specificity */
:where(.card h2) { margin-block-end: 0.5rem; }

Use these features to make intent clearer, not to hide an unmaintainable selector strategy.


:has() selects based on a relationship

.field:has(input:invalid) {
  border-color: var(--color-error);
}

The parent can respond to the state of a descendant without JavaScript merely to add a class.

Still check whether the selector expresses a stable relationship and whether a class would make the ownership clearer.


CSS nesting needs readable boundaries

.card {
  padding: 1rem;

  & > h2 {
    margin-block: 0;
  }

  &:hover {
    box-shadow: 0 0.5rem 1rem rgb(0 0 0 / 12%);
  }
}

Nesting can keep a component’s local rules together. Deep nesting still creates coupling and should not replace a clear component contract.


User preferences are part of responsiveness

@media (prefers-reduced-motion: reduce) {
  *, *::before, *::after {
    animation-duration: 0.01ms !important;
    animation-iteration-count: 1 !important;
    transition-duration: 0.01ms !important;
    scroll-behavior: auto !important;
  }
}

Other preferences can affect colour, contrast, and interaction assumptions. Responsive design includes the user’s environment, not only its dimensions.


Motion should communicate something

Use motion to communicate:

  • a change of state;
  • a relationship between locations;
  • progress or feedback;
  • continuity during an interaction.

Avoid motion that delays reading, causes discomfort, or exists only because a component has an animation library available.


Styling architecture is a dependency system

Every approach answers questions about:

  • where styles live;
  • how names are scoped;
  • how styles are composed;
  • how themes are represented;
  • how overrides work;
  • how unused styles are removed;
  • how runtime and build-time costs are distributed.

The right choice depends on the project’s constraints and ownership boundaries.


Plain CSS

Strengths:

  • native browser model;
  • no framework requirement;
  • easy to inspect in DevTools;
  • supports layers, custom properties, and modern selectors.

Risks:

  • unclear naming can create collisions;
  • global ownership can become ambiguous;
  • unused styles need a deliberate strategy.

Architecture is still required even when no tool generates the CSS.


Component-oriented CSS and CSS Modules

Component-oriented CSS groups styles around an interface boundary.

CSS Modules add build-time scoping:

flowchart LR
    A[Card.module.css] --> B[Build Step] --> C[Generated Local Scoped Class Names]

They reduce accidental collisions, but they do not decide:

  • token ownership;
  • layout responsibility;
  • accessibility;
  • responsive strategy;
  • whether a component boundary is well designed.

Utility-first CSS

Utility classes make small decisions explicit in markup:

<div class="grid gap-4 md:grid-cols-3"></div>

Benefits can include consistency and fast composition. Costs can include noisy markup, difficult component-level meaning, or a false belief that utilities remove architecture.

Utilities are a vocabulary, not a complete design system.


CSS-in-JS is not one technology

CSS-in-JS can mean different things:

  • runtime style generation;
  • build-time extraction;
  • component-local syntax;
  • theme-aware value functions;
  • atomic style generation.

Compare the actual tool’s runtime cost, debugging model, SSR behaviour, accessibility support, and ownership boundaries - not the category label.


Compare problems, not fashion

Ask:

  • How many teams own the styles?
  • Do styles need to work without a framework runtime?
  • How important is runtime theming?
  • What is the build and delivery environment?
  • How will developers debug the final CSS?
  • How will shared components evolve?

Choose the smallest architecture that meets the real constraints.


The running catalogue interface

We will build a responsive multilingual catalogue with:

  • semantic page structure;
  • a toolbar;
  • summary cards;
  • a product grid;
  • theme tokens;
  • container-aware cards;
  • RTL-safe spacing;
  • reduced-motion support.

The interface is a test of relationships, not a gallery of CSS tricks.


Establish the cascade and tokens

@layer reset, tokens, base, components, utilities;

@layer tokens {
  :root {
    --surface-page: #f5f7fa;
    --surface-card: #ffffff;
    --text-primary: #172a42;
    --color-accent: #1d588c;
    --space-3: 0.75rem;
  }
}

Start by deciding who owns values and who owns component rules.


Build the page layout

.page {
  display: grid;
  grid-template-columns: minmax(0, 1fr);
  gap: var(--space-6);
  inline-size: min(100% - 2rem, 80rem);
  margin-inline: auto;
}

@media (width > 60rem) {
  .page { grid-template-columns: 16rem minmax(0, 1fr); }
}

The page breakpoint responds to the viewport. The components inside it should still respond to the space each one receives.


Build the toolbar with Flexbox

.toolbar {
  display: flex;
  flex-wrap: wrap;
  align-items: center;
  gap: 0.75rem;
}

.toolbar__search {
  flex: 1 1 18rem;
  min-inline-size: min(100%, 12rem);
}

Wrapping is part of the design. It is better than forcing every control into a single row that cannot contain translated labels.


Build the product grid with Grid

.product-grid {
  display: grid;
  grid-template-columns: repeat(
    auto-fit,
    minmax(min(100%, 16rem), 1fr)
  );
  gap: 1rem;
}

The grid expresses a useful minimum and lets the available container determine how many columns fit.


Adapt cards to their container

.product-grid { container: products / inline-size; }

.product-card {
  display: grid;
  gap: 0.75rem;
}

@container products (width > 42rem) {
  .product-card {
    grid-template-columns: 1fr auto;
  }
}

The card does not need to know whether the product grid is on the full page or inside a sidebar.


Use subgrid only for a shared alignment

.product-card {
  grid-template-rows: subgrid;
  grid-row: span 3;
}

If product names, descriptions, and actions should line up across a row, subgrid can share the parent tracks.

If cards do not need shared rows, keep their internal layout independent.


Make the catalogue RTL-safe

.catalogue {
  padding-inline: 1rem;
  margin-inline: auto;
}

.status {
  border-inline-start: 0.25rem solid var(--status-color);
  padding-inline-start: 0.75rem;
}

Test English, Arabic, and Sorani Kurdish content with long labels. Do not assume that replacing left with right is localization.


Respect reduced motion

@media (prefers-reduced-motion: reduce) {
  .catalogue * {
    transition: none;
    animation: none;
  }
}

The user preference is a real requirement. It belongs in the component’s architecture and verification, not only in a final accessibility audit.


CSS architecture has dependencies

flowchart TD
    Tokens[Tokens] --> Components[Components] --> Composition[Layout Composition]
    Tokens -.-> Themes[Themes]
    Components -.-> Responsive[Responsive Contexts]
    Composition -.-> Pages[Pages]

When a component reaches into page-specific selectors, or a page owns a component’s internal spacing, the dependency direction becomes unclear.

Good CSS makes change flow through deliberate boundaries.


Misconceptions to leave behind (Part 1)

MisconceptionBetter model
Specificity decides every conflict.Origin, layers, specificity, and order all matter.
The most specific selector is best.The most maintainable selector is usually better.
Flexbox replaced Grid.Flexbox and Grid solve different dimensional problems.
Grid replaced Flexbox.Tool choice follows the relationship being laid out.
Responsive means phone/tablet/desktop.Respond to constraints and user environment.

Misconceptions to leave behind (Part 2)

MisconceptionBetter model
Container queries replace media queries.They answer different scope questions.
RTL means swap every left and right.Use logical properties and inspect meaning.
CSS variables are text substitution.They cascade, inherit, and change at runtime.
CSS Modules solve architecture.Scoping helps; ownership and layout still need design.
Utility CSS removes architecture.A utility vocabulary still needs rules and boundaries.

The practical lab

Build a responsive multilingual catalogue.

The runnable implementation belongs in the separate Playground repository. The practical instructions define the evidence that the layout works.


Practical stages 1–3

  1. Establish semantic HTML and a small token layer.
  2. Add reset, base, components, and utilities cascade layers.
  3. Convert raw values into semantic tokens and add light/dark theme values.

At each stage, inspect which layer and token supplied the final style.


Practical stages 4–6

  1. Build the page layout with normal flow and Grid.
  2. Build the toolbar with wrapping Flexbox.
  3. Build the product grid with minmax() and intrinsic sizing.

Test narrow widths and long translated labels before adding more breakpoints.


Practical stages 7–9

  1. Demonstrate subgrid where card rows must align.
  2. Add container queries for cards and the toolbar.
  3. Make spacing, borders, and alignment direction-independent.

Use DevTools to test the same component in more than one container.


Practical stages 10–12

  1. Respect reduced motion.
  2. Use :where(), :is(), :has(), or nesting where they make relationships clearer.
  3. Draw the CSS architecture and document ownership boundaries.

The final deliverable is an explainable system, not merely a screenshot.


Try it yourself

Take one product card and test it in four environments:

  • wide page content;
  • narrow sidebar content;
  • English text;
  • Arabic or Sorani Kurdish text.

Record one layout decision that remains valid in all four environments and one decision that must adapt.


Troubleshooting questions (Part 1)

SymptomFirst question
A rule will not winWhich layer, specificity, origin, and order are involved?
A flexible item overflowsIs its minimum size preventing shrinkage?
Cards do not alignIs shared track alignment actually required?
A component breaks in a sidebarIs it using a viewport query instead of a container query?

Troubleshooting questions (Part 2)

SymptomFirst question
Arabic text collides with controlsAre logical properties and direction boundaries correct?
A theme change needs many editsAre components consuming semantic tokens?
Motion is uncomfortableIs reduced motion handled at the component boundary?

Completion check

  • I can explain why a declaration wins.
  • I can use cascade layers to represent ownership.
  • I can distinguish raw tokens from semantic tokens.
  • I can choose Flexbox, Grid, or normal flow for a reason.
  • I can use intrinsic sizing and minmax() without a device list.
  • I can explain when subgrid is useful.
  • I can distinguish media queries from container queries.
  • I can use logical properties for direction-independent layout.
  • I can respect reduced motion and long translated content.
  • I can compare styling architectures by project constraints.
  • I completed the responsive multilingual catalogue practical.

Next: Modern JavaScript & Asynchronous Programming

Chapter 4: lexical scope, closures, asynchronous work, promises, cancellation, and the event loop as an application runtime.

Modern Front-End Engineering

4 - Modern JavaScript & Asynchronous Programming

Chapter 4: understand scope, closures, modules, promises, concurrency, cancellation, and locale-aware output.

Modern JavaScript & Asynchronous Programming

Control state, boundaries, timing, cancellation, and data flow

Chapter 4

Polla Fattah


Today’s goal

Move beyond syntax toward the language model needed for modern front-end work.

We will connect:

  • scope and closures;
  • objects and composition;
  • data transformation and immutability;
  • modules and dynamic loading;
  • promises and asynchronous execution;
  • concurrency, races, and cancellation;
  • locale-aware output.

By the end of today you can

  • explain where a JavaScript binding lives;
  • use closures deliberately without confusing them with copied values;
  • describe prototype delegation and why classes do not remove it;
  • choose data transformation methods for clarity;
  • update nested state without accidental mutation;
  • design module boundaries and dynamic imports;
  • distinguish sequential work from concurrent work;
  • handle rejected promises and cancellation explicitly;
  • prevent stale results from winning a race;
  • format numbers, dates, and plurals for a locale.

The chapter’s progression

flowchart TD
    A[Scope & Closures] --> B[Objects & Composition]
    B --> C[Data Operations & Immutability]
    C --> D[Modules & Iteration]
    D --> E[Promises & async/await]
    E --> F[Concurrency & Cancellation]
    F --> G[Locale-Aware Output]

The core idea is:

Modern JavaScript controls state, boundaries, timing, and data flow.


Lexical scope: where a name can be used

const applicationName = "Citizen Portal";

function showName() {
  console.log(applicationName);
}

showName();

The function can look outward to the scope that surrounds it.

An outer scope cannot look inward into a function’s private bindings.

Scope follows the source-code structure. It is lexical, not based on which function happened to call another function at runtime.


Block scope is a lifetime boundary

if (loggedIn) {
  const message = "Welcome back";
  console.log(message);
}

console.log(message); // ReferenceError

Practical rule:

  • prefer const by default;
  • use let when reassignment is required;
  • understand that both respect block scope;
  • treat a block as a useful lifetime boundary.

Closures retain bindings

function createCounter() {
  let count = 0;

  return function increment() {
    count += 1;
    return count;
  };
}

const counter = createCounter();
counter(); // 1
counter(); // 2

createCounter() finished, but increment retains access to its lexical environment.

The closure does not copy count. It retains access to the binding.


Closures appear everywhere in front-end code

function attachFilter(input, applyFilter) {
  input.addEventListener("input", () => {
    applyFilter(input.value);
  });
}

The callback remembers input and applyFilter from the surrounding scope.

Common uses include:

  • event handlers;
  • callback configuration;
  • private state factories;
  • debounced functions;
  • memoized computations;
  • request controllers.

Closures can keep data alive

function createLogger(records) {
  return (message) => {
    records.push({ message, at: Date.now() });
  };
}

As long as the returned logger is reachable, records remains reachable too.

Closures provide privacy and state, but they can also retain large objects, DOM nodes, or subscriptions longer than intended.

Release listeners, timers, and references when the feature is destroyed.


Objects have a prototype model

An object can delegate property lookup to another object through its prototype.

const service = {
  describe() {
    return "public service";
  }
};

const request = Object.create(service);
request.id = 1042;
request.describe(); // delegated lookup

The object does not need to own every method directly.


Prototype delegation is not copying

Property lookup conceptually follows:

flowchart TD
    A[Request object itself] -->|if absent| B[Service prototype]
    B -->|if absent| C[Object.prototype]
    C -->|if absent| D[undefined]

This affects identity, mutation, method lookup, and debugging. Understand the delegation even when your code uses class syntax.


Classes are syntax over prototypes

class RequestRecord {
  constructor(id) {
    this.id = id;
  }

  describe() {
    return `Request ${this.id}`;
  }
}

The method is associated with the class prototype. Class syntax improves readability for some designs; it does not replace JavaScript’s prototype model.


Composition is often simpler than inheritance

Instead of a deep hierarchy:

flowchart LR
    A[BaseRecord] --> B[ServiceRecord] --> C[UrgentServiceRecord] --> D[LocalisedUrgentRecord]

compose focused capabilities:

const record = {
  ...withIdentity(1042),
  ...withStatus("in-review"),
  ...withLocale("ar")
};

Composition can reduce coupling and make a capability easier to reuse or replace.


Destructuring names the data you need

const request = {
  id: 1042,
  status: "in-review",
  applicant: { name: "Sara" }
};

const {
  id,
  status,
  applicant: { name }
} = request;

Destructuring can make data flow obvious, but avoid extracting so much that the source of a value becomes difficult to follow.


Rest and spread have different jobs

const { id, ...details } = request;
const updated = { ...request, status: "approved" };
  • rest collects the remaining values;
  • spread expands values into a new object or array.

Both are shallow operations.


Spread is not a deep copy

const original = { preferences: { theme: "dark" } };
const copy = { ...original };

copy.preferences.theme = "light";
console.log(original.preferences.theme); // light

The outer object is new. The nested object is still shared.

Know which references remain shared before calling something an immutable update.


Optional chaining is a guarded lookup

const phone = applicant?.contacts?.primaryPhone;

It is useful when a value is legitimately absent.

It does not:

  • validate the data shape;
  • supply a meaningful fallback;
  • prove that a missing value is acceptable;
  • repair a broken API contract.

Use it where absence is part of the model, not everywhere as a blanket guard.


Nullish coalescing preserves meaningful zeroes

const pageSize = settings.pageSize ?? 20;

?? falls back only for null or undefined.

That differs from ||:

0 || 20   // 20
0 ?? 20   // 0

Choose the operator according to the domain meaning of empty strings, zero, false, null, and undefined.


Transform collections by intent

const labels = requests.map(request => request.statusLabel);
const open = requests.filter(request => request.status === "open");
const selected = requests.find(request => request.id === id);
const hasUrgent = requests.some(request => request.priority === "urgent");
const allValid = fields.every(field => field.valid);

Each method communicates a different question. Prefer the method whose name matches the operation.


Use reduce() when it clarifies

const totals = requests.reduce((summary, request) => {
  summary[request.status] = (summary[request.status] ?? 0) + 1;
  return summary;
}, {});

reduce() can express aggregation, but it can also hide a complicated algorithm inside one callback.

Use a loop or named helper when the steps matter more than compactness.


Immutable update patterns

const nextState = {
  ...state,
  filters: {
    ...state.filters,
    status: "open"
  }
};

The update creates new references along the changed path and preserves unrelated references.

This helps state systems compare identity and makes change boundaries visible.


Identity and change detection

const same = state.filters === nextState.filters; // false
const sameList = state.requests === nextState.requests; // true

Identity can communicate which part of a state tree changed.

That does not mean every object must be recreated on every update. Preserve references when the value did not change.


Do not turn immutability into dogma

Mutation can be reasonable when:

  • the object is local to one operation;
  • no other consumer observes it;
  • the mutation improves clarity or performance;
  • ownership is explicit.

The architectural question is not “mutation or no mutation?”

It is:

Who can observe this value, and what change signal do they rely on?


ES Modules create boundaries

// format-status.js
export function formatStatus(status) {
  return status.replaceAll("-", " ");
}

// screen.js
import { formatStatus } from "./format-status.js";

Modules provide:

  • explicit dependencies;
  • module scope;
  • reusable exports;
  • a graph the build tool can inspect.

Modules are architecture, not only file organization.


Named and default exports

Named exports make the public vocabulary explicit:

export function parseRequest() {}
export function validateRequest() {}

Default exports can represent one primary value:

export default class RequestClient {}

Choose a consistent convention. Avoid making consumers guess whether a module has one canonical thing or a set of named capabilities.


Module scope is private by default

const cache = new Map();

export function getCached(key) {
  return cache.get(key);
}

cache is not a global just because the module can access it. Consumers see only what the module exports.

This is a useful boundary for implementation details and state ownership.


The module graph is a dependency graph

flowchart TD
    Screen[screen.js] --> ReqClient[request-client.js]
    Screen --> FormatStatus[format-status.js]
    ReqClient --> HTTP[http.js]

The graph affects:

  • evaluation order;
  • bundling;
  • caching;
  • code splitting;
  • circular-dependency risk;
  • ownership and testing boundaries.

Design imports as deliberately as public API calls.


Dynamic imports load on demand

button.addEventListener("click", async () => {
  const { openReport } = await import("./report.js");
  openReport();
});

Dynamic imports can reduce initial work and defer rarely used features.

They add asynchronous failure, loading state, and chunk-boundary concerns. A dynamic import is not free merely because it is written in one line.


Iteration is more than for loops

An iterable provides a protocol for producing values over time.

for (const request of requests) {
  renderRequest(request);
}

Arrays, strings, Maps, Sets, DOM collections, and custom objects can participate in iteration.

The protocol separates “how values are produced” from “how they are consumed.”


Iterators produce one value at a time

const iterator = ["queued", "open"].values();
iterator.next(); // { value: "queued", done: false }
iterator.next(); // { value: "open", done: false }
iterator.next(); // { value: undefined, done: true }

The done flag is part of the contract. Iterators can represent sequences that are calculated lazily rather than stored as a complete array.


Generators pause and resume

function* statuses() {
  yield "queued";
  yield "in-review";
  yield "approved";
}

Generators provide a convenient way to define an iterator. They can model progressive work, but they do not automatically make expensive work asynchronous or cancellable.


Async iteration represents arriving values

for await (const event of stream) {
  renderEvent(event);
}

Async iteration is useful when values arrive over time. The consumer still needs to understand completion, errors, backpressure, and cancellation.


From callbacks to promises

Callbacks can express completion and failure, but nested workflows become hard to compose:

loadUser(id, user => {
  loadRequests(user, requests => {
    render(requests);
  }, handleError);
}, handleError);

Promises represent a future result and provide composable success and failure paths.


A promise represents a future result

const request = fetch("/api/requests/1042");

request.then(response => response.json())
  .then(data => render(data))
  .catch(error => showError(error));

A promise is not the result itself. It represents a pending, fulfilled, or rejected computation.


Creating promises is an ownership decision

function wait(ms) {
  return new Promise(resolve => {
    setTimeout(resolve, ms);
  });
}

Create a promise when adapting a callback or exposing an asynchronous contract.

Do not wrap an API in a new promise without a reason; unnecessary wrappers can hide errors and complicate cancellation.


Promise chaining carries values forward

fetch("/api/requests")
  .then(response => response.json())
  .then(requests => requests.filter(isVisible))
  .then(renderRequests);

Each handler returns the value or promise for the next step.

Returning nothing intentionally passes undefined; forgetting to return is a common source of a broken chain.


Promise errors propagate through the chain

loadRequests()
  .then(validateResponse)
  .then(renderRequests)
  .catch(error => showError(error));

An exception or rejection skips forward to the next compatible rejection handler. Place recovery at the boundary that can actually decide what to do.


async and await improve expression

async function loadScreen() {
  const response = await fetch("/api/requests");
  const requests = await response.json();
  renderRequests(requests);
}

await makes a promise result available inside the function. It does not turn the browser into a blocking environment.


Async functions always return promises

async function answer() {
  return 42;
}

answer().then(console.log); // 42

Even a plain return value is wrapped in a fulfilled promise. Callers must still decide how to handle rejection.


Handle errors at the right boundary

async function loadScreen() {
  try {
    const response = await fetch("/api/requests");
    if (!response.ok) throw new Error(`HTTP ${response.status}`);
    return await response.json();
  } catch (error) {
    showRecoverableError(error);
    throw error;
  }
}

Do not catch an error merely to log it and then pretend the operation succeeded.


Not every failure is a network error

An async operation can fail because:

  • the request was blocked or offline;
  • the server returned a non-2xx response;
  • the body was invalid JSON;
  • validation rejected the data;
  • rendering code threw;
  • the user cancelled the work;
  • a stale result was deliberately ignored.

Your error model should preserve the difference when the UI needs it.


Sequential work is sometimes correct

const user = await loadUser();
const permissions = await loadPermissions(user.id);

The second operation depends on the first. Sequential execution expresses that dependency.

The cost is that total time includes both waits.


Accidental sequential work is a performance bug

const categories = await loadCategories();
const featured = await loadFeatured();

If the requests are independent, this waits unnecessarily.

Start independent work together, then await the combined result.


Promise.all() expresses all-or-nothing coordination

const [categories, featured] = await Promise.all([
  loadCategories(),
  loadFeatured()
]);

The operations start without waiting for one another. The combined promise fulfils when all fulfil and rejects when one rejects.

Use it when one failure should prevent the combined result from being used.


Promise.allSettled() preserves every outcome

const results = await Promise.allSettled([
  loadRecommendations(),
  loadAnnouncements(),
  loadOptionalMetrics()
]);

Use it when partial success is meaningful and each result needs its own status.

The choice communicates product behaviour, not only JavaScript preference.


Concurrency is not parallelism

Several asynchronous operations can be in flight at once while JavaScript continues on its event loop.

That does not mean JavaScript is running those callbacks on several CPU cores.

Concurrency is about overlapping waiting and coordinating completion.

Parallelism is about simultaneous execution.


Race conditions can happen in one thread

sequenceDiagram
    participant UI as User Interface
    participant Net as Network
    Note over UI: Types "ca" (Req A)
    UI->>Net: Request A ("ca")
    Note over UI: Types "car" (Req B)
    UI->>Net: Request B ("car")
    Net-->>UI: Response B arrives (150ms)
    Note over UI: UI shows "car" results
    Net-->>UI: Response A arrives late (800ms)
    Note over UI: Race! Stale "ca" overwrites "car"

If every response renders immediately, request A can overwrite the newer result for care.

Single-threaded JavaScript does not remove ordering races between asynchronous operations.


Race conditions are correctness problems

The question is not only:

Which request finished last?

It is:

Which result is still valid for the current user intent?

Freshness is an application rule. The network does not know which response the user currently cares about.


Strategy 1: ignore stale results

let latestQuery = "";

async function search(query) {
  latestQuery = query;
  const result = await fetchResults(query);
  if (query !== latestQuery) return;
  renderResults(result);
}

This is simple and safe when the old work is harmless. It still spends the resources needed to finish the old request.


Strategy 2: cancel outdated work

let controller;

async function search(query) {
  controller?.abort();
  controller = new AbortController();
  const response = await fetch(`/api/search?q=${query}`, {
    signal: controller.signal
  });
  return response.json();
}

Cancellation can save work and communicate that the previous intent is no longer relevant.

The UI must distinguish cancellation from an unexpected failure.


AbortController is a general signal

const controller = new AbortController();

fetch(url, { signal: controller.signal });
timer(signal, controller.signal);

controller.abort();

The signal is a shared cancellation contract. The operation decides how to listen and clean up when the signal aborts.


Cancellation is an application design problem

Ask:

  • What work belongs to the cancelled intent?
  • Who owns the controller?
  • What resources need cleanup?
  • Should cancellation be silent or visible?
  • Can a new operation replace the old one safely?
  • What happens if cancellation occurs after the result arrives?

Adding AbortController without defining ownership only moves the ambiguity.


Debouncing delays the start of work

function debounce(callback, delay) {
  let timer;
  return (...args) => {
    clearTimeout(timer);
    timer = setTimeout(() => callback(...args), delay);
  };
}

Debouncing waits for a pause in input before starting work.

It reduces the number of operations. It does not cancel a request that already started.


Debouncing and cancellation solve different problems

flowchart TD
    A["Debounce: delay unnecessary starts"] --> B["Cancellation: stop outdated work in flight"]
    B --> C["Freshness: reject obsolete results"]

Search interfaces often need all three.


Error propagation needs an owner

flowchart TD
    T["Transport layer: reports request failure"] --> D["Data layer: validates and normalizes data"]
    D --> F["Feature layer: decides recoverable UI state"]
    F --> S["Screen layer: presents feedback and retry"]

Do not let every layer catch and replace the same error with a vague message.

Each boundary should add context or make a decision.


finally is for cleanup

setLoading(true);

try {
  await loadResults();
} catch (error) {
  showError(error);
} finally {
  setLoading(false);
}

Use finally for work that should occur after success, failure, or cancellation: loading flags, controller references, locks, and temporary UI state.


Avoid unhandled promise rejections

Every promise chain needs an intentional owner:

void saveDraft().catch(error => {
  reportSaveFailure(error);
});

If a caller intentionally starts background work, make that intent visible and attach a failure path. A rejected promise should not disappear silently.


Internationalization is more than translation

Locale-aware interfaces consider:

  • language and script;
  • number formatting;
  • currencies;
  • dates and time zones;
  • relative time;
  • plural rules;
  • segmentation;
  • direction and layout.

Do not concatenate English assumptions into strings and call the result localized.


Format numbers with Intl.NumberFormat

const formatter = new Intl.NumberFormat("ar-IQ", {
  maximumFractionDigits: 2
});

formatter.format(12345.67);

The locale controls grouping, decimal conventions, digits, and other display rules. Formatting belongs near presentation, not in the domain value itself.


Currency is a meaning, not a symbol

new Intl.NumberFormat("en-US", {
  style: "currency",
  currency: "USD"
}).format(1234.5);

Do not hard-code a dollar sign or assume that the currency code determines the entire visual format.

Keep the numeric amount and currency identity separate until presentation.


Dates need a locale and a time-zone decision

new Intl.DateTimeFormat("en-GB", {
  dateStyle: "medium",
  timeZone: "Asia/Baghdad"
}).format(date);

Formatting is not the same as deciding what instant or calendar meaning the application intends. Store and transmit a clear temporal representation.


Relative time communicates change

const relative = new Intl.RelativeTimeFormat("en", { numeric: "auto" });
relative.format(-1, "day"); // yesterday

The correct unit and rounding policy are application decisions. A relative label should not hide the exact timestamp when the exact time matters.


Plural rules are grammatical rules

const plural = new Intl.PluralRules("en");
plural.select(1); // one
plural.select(2); // other

Different languages have different plural categories. Avoid building messages with count === 1 ? "item" : "items" as a universal model.


Intl.Segmenter respects language boundaries

const segmenter = new Intl.Segmenter("ar", {
  granularity: "word"
});

for (const part of segmenter.segment(text)) {
  console.log(part.segment);
}

String length and substring boundaries are not always equivalent to visible words or user-perceived characters.


A cancelable search: step 1

Debounce the user’s input so each keystroke does not immediately start work:

const scheduleSearch = debounce((query) => {
  runSearch(query);
}, 250);

This controls the rate of starts. It does not yet control in-flight requests.


A cancelable search: step 2

Cancel the previous request when a newer query becomes authoritative:

let activeController;

async function runSearch(query) {
  activeController?.abort();
  activeController = new AbortController();

  const response = await fetch(`/api/search?q=${encodeURIComponent(query)}`, {
    signal: activeController.signal
  });

  return response.json();
}

The feature owns the controller because it owns the current search intent.


A cancelable search: step 3

Handle cancellation separately from failure:

try {
  const results = await runSearch(query);
  renderResults(results);
} catch (error) {
  if (error.name === "AbortError") return;
  showSearchError(error);
}

Cancellation is expected control flow. It should not flash an error message for the user.


A cancelable search: the complete contract

flowchart TD
    A[Input event] --> B[Debounce delay]
    B --> C[Create current AbortController]
    C --> D[Cancel previous active request]
    D --> E[Fetch network data with signal]
    E --> F[Validate response & token freshness]
    F --> G[Format locale-aware values]
    G --> H[Render current results only]
    H --> I[Cleanup controller in finally]

Each arrow is a boundary where an error, race, or ownership decision can occur.


The practical lab

Build a cancelable locale-aware search.

The runnable Vite and TypeScript implementation will live in the separate Playground repository. This deck describes the JavaScript behaviour and the evidence to collect.


Practical stages 1–3

  1. Use closure state for a small search controller or cache.
  2. Apply immutable updates to the search state.
  3. Split the feature into modules with clear imports and exports.

Write down which module owns the current query, results, and cancellation controller.


Practical stages 4–6

  1. Dynamically import an optional result formatter or panel.
  2. Compare sequential and concurrent independent requests.
  3. Create a deliberate race condition by delaying responses differently.

Observe how a correct build can still show incorrect stale data.


Practical stages 7–9

  1. Prevent stale results from replacing newer results.
  2. Add AbortController cancellation.
  3. Add input debouncing.

Test rapid typing, clearing the input, slow responses, cancellation, and network failure separately.


Practical stages 10–11

  1. Add locale-aware numbers, dates, plurals, and direction.
  2. Draw the asynchronous flow and label every ownership boundary.

The final diagram should show what starts work, what can cancel it, what can reject, and what is allowed to update the UI.


Try it yourself

Choose one asynchronous feature in an application you know.

Answer:

  • Can two operations be in flight together?
  • What makes an old result stale?
  • Who owns cancellation?
  • Which failures are expected user flow?
  • Where should locale formatting happen?

Write one explicit contract before changing the code.


Troubleshooting questions (Part 1)

SymptomFirst question
A variable is unavailableWhich lexical scope owns it?
A callback uses old stateWhich binding did the closure retain?
A promise chain returns undefinedDid each handler return its next value?
Independent requests are slowAre they accidentally awaited sequentially?

Troubleshooting questions (Part 2)

SymptomFirst question
Old search results appearWhat defines freshness and cancellation?
Loading never endsIs cleanup in finally?
Cancellation shows as an errorIs AbortError handled separately?
Dates look different by machineIs locale and time zone explicit?

Completion check

  • I can explain lexical scope and closures.
  • I understand prototype delegation beneath class syntax.
  • I can choose clear collection transformations.
  • I can update nested state while preserving useful identity boundaries.
  • I can design a module graph with explicit dependencies.
  • I can distinguish sequential work from concurrent work.
  • I can explain promise rejection and error ownership.
  • I can identify and prevent stale asynchronous results.
  • I can use cancellation and debouncing for different purposes.
  • I can format numbers, dates, plurals, and text for a locale.
  • I completed the cancelable locale-aware search practical.

Next: TypeScript, Runtime Contracts & Safe Data Boundaries

Chapter 5: static types, inference, runtime validation, trusted boundaries, and the difference between what a compiler knows and what a browser receives.

Modern Front-End Engineering

5 - TypeScript, Runtime Contracts & Safe Data Boundaries

Chapter 5: model data precisely, narrow unknown values, use generics, and validate every external boundary.

TypeScript, Runtime Contracts & Safe Data Boundaries

Make the boundary explicit

Chapter 5

Polla Fattah


Today’s goal

Understand what TypeScript can prove, what it cannot know, and how runtime validation turns untrusted input into trusted application data.

We will connect:

  • inference and explicit types;
  • unions, narrowing, and exhaustive state design;
  • generics and reusable contracts;
  • strictness, DOM typing, and assertions;
  • APIs, URLs, storage, forms, and configuration;
  • parsers, schemas, branded values, and trusted domain data.

By the end of today you can

  • model domain states instead of decorating arbitrary objects;
  • use unknown at an external boundary;
  • narrow values with evidence and discriminated unions;
  • write generic result and collection APIs;
  • keep strict TypeScript useful rather than ceremonial;
  • distinguish a type assertion from a runtime conversion;
  • validate API, URL, storage, form, and configuration data;
  • separate transport types from domain types;
  • create branded identifiers only after validation;
  • explain why static types do not replace tests or runtime contracts.

The central lesson

flowchart LR
    A["TypeScript Static Types\n(Accepted by Compiler)"] -.->|Type Erasure| B["Runtime JavaScript"]
    C["Runtime Data\n(Delivered from Outside)"] --> D["Runtime Validation Boundary"]

An annotation is useful for the compiler and editor.

It is not a force field around a value arriving from a network, URL, browser, or user.


The chapter’s progression

flowchart TD
    A[Type Inference] --> B[Data Modeling]
    B --> C[Unions & Narrowing]
    C --> D[Generics]
    D --> E[Strictness & DOM Typing]
    E --> F[Trust Boundaries]
    F --> G[Runtime Validation]
    G --> H[Trusted Domain Data]

The order matters: prove small facts first, then compose them into safe boundaries.


JavaScript has values; TypeScript adds a model

const response = await fetch("/api/products");
const data = await response.json();

JavaScript executes the values that arrive.

TypeScript can describe the values we expect while developing, but its annotations are removed before the browser runs the code.


A type annotation is not validation

type User = { id: string; name: string };

const user = JSON.parse(input) as User;
console.log(user.name.toUpperCase());

The assertion changes the compiler’s story.

It does not inspect input, check id, or create a missing name.


Start with inference

const pageSize = 20;
const status = "loading";
const visible = true;

// pageSize: number
// status: string (in a mutable binding)
// visible: boolean

Inference keeps local code readable and lets the implementation remain the source of truth.

Add an annotation when it documents intent, constrains a boundary, or catches an important mistake.


Widening and literal values

const fixedStatus = "loading"; // "loading"
let changingStatus = "loading"; // string

const config = { mode: "dark" }; // { mode: string }

Literal information can widen when a value must remain mutable.

Use literal types deliberately when a small vocabulary is part of the contract.


Model the domain, do not label a guess

type Product = {
  id: string;
  title: string;
  priceCents: number;
};

This is a useful domain model for trusted application data.

It is not evidence that a random JSON object already satisfies the model.


Interfaces describe object shape

interface Product {
  id: string;
  title: string;
  priceCents: number;
}

function formatProduct(product: Product) {
  return `${product.title}: ${product.priceCents / 100}`;
}

Interfaces work well for object contracts that may be extended or implemented across a codebase.


Type aliases compose precisely

type ProductId = string;
type Currency = "IQD" | "USD";
type ProductSummary = Pick<Product, "id" | "title">;

Aliases are especially expressive for unions, tuples, mapped types, conditional types, and domain vocabulary.

Choose the form that communicates the design; both participate in structural typing.


Literal types create a controlled vocabulary

type Theme = "light" | "dark";
type SortDirection = "ascending" | "descending";

function setTheme(theme: Theme) {}
setTheme("dark");
// setTheme("blue"); // compile-time error

Small finite sets should be visible in the type rather than repeated as undocumented strings.


Unions represent alternatives

type Result = Product | { message: string };

function readTitle(value: Result) {
  if ("title" in value) return value.title;
  return value.message;
}

The safe question is not “what do I hope this is?”

It is “what evidence distinguishes the possible members?”


Narrow with typeof when the evidence fits

function toLabel(value: string | number) {
  if (typeof value === "number") {
    return value.toFixed(2);
  }
  return value.toUpperCase();
}

Narrowing is a proof step. After the branch, TypeScript allows operations supported by the proven type.


Narrow with property checks

type Failure = { message: string; code: number };
type Success = { data: Product[] };

function describe(value: Failure | Success) {
  if ("data" in value) return `${value.data.length} products`;
  return `${value.code}: ${value.message}`;
}

Property checks are useful, but a stable discriminant is often clearer for important state machines.


Discriminated unions make state explicit

type LoadState<T> =
  | { status: "idle" }
  | { status: "loading" }
  | { status: "success"; data: T }
  | { status: "error"; message: string };

Each state carries exactly the data that makes sense in that state.

This prevents impossible combinations such as status: "loading" with stale error and data fields.


Render state by its discriminant

function renderProducts(state: LoadState<Product[]>) {
  switch (state.status) {
    case "idle": return "Choose a search";
    case "loading": return "Loading…";
    case "success": return `${state.data.length} results`;
    case "error": return state.message;
  }
}

The branch gives the renderer the exact fields valid for that state.


Intersections combine capabilities

type Identified = { id: string };
type Timestamped = { updatedAt: string };
type StoredProduct = Product & Identified & Timestamped;

Use intersections when one value genuinely satisfies multiple independent contracts.

Do not use them to hide a domain model that is becoming difficult to understand.


unknown is the honest boundary type

function receiveExternalValue(value: unknown) {
  // value.toString(); // not allowed without proof
  if (typeof value === "string") return value.trim();
  return "not a string";
}

unknown says: a value exists, but this function has not earned knowledge about its shape yet.

It is safer than any because every operation requires evidence.


A type guard is a proof function

function isProduct(value: unknown): value is Product {
  if (typeof value !== "object" || value === null) return false;
  const record = value as Record<string, unknown>;
  return typeof record.id === "string"
    && typeof record.title === "string"
    && typeof record.priceCents === "number";
}

The return type tells TypeScript what follows when the function returns true.

The implementation must justify that claim at runtime.


Type guards can still be wrong

function isProduct(value: unknown): value is Product {
  return typeof value === "object";
}

This compiles, but it lies: arrays, dates, and empty objects pass the test.

A type predicate is not automatically verified by the compiler. Treat it as safety-critical code.


never protects exhaustive decisions

function assertNever(value: never): never {
  throw new Error(`Unhandled state: ${String(value)}`);
}

function label(state: LoadState<unknown>) {
  switch (state.status) {
    case "idle": return "Idle";
    case "loading": return "Loading";
    case "success": return "Ready";
    case "error": return "Failed";
    default: return assertNever(state);
  }
}

Adding a new state now creates a compile-time reminder at every incomplete decision.


Generics preserve relationships

function first<T>(items: T[]): T | undefined {
  return items[0];
}

const firstProduct = first(products); // Product | undefined
const firstNumber = first([1, 2, 3]); // number | undefined

Generics are not “any with extra syntax”. They carry a relationship between inputs and outputs.


A generic result keeps success data precise

type ApiResult<T> =
  | { ok: true; data: T }
  | { ok: false; error: string };

function show(result: ApiResult<Product[]>) {
  if (!result.ok) return result.error;
  return result.data.map(product => product.title);
}

The common result contract is reusable while T preserves the domain-specific payload.


Constrain a generic when it needs a capability

function getById<T extends { id: string }>(items: T[], id: string) {
  return items.find(item => item.id === id);
}

The constraint does not say every T is exactly an object with only id.

It says the function may safely rely on id while preserving additional fields.


Generic components should describe relationships

type TableProps<T> = {
  rows: T[];
  getKey: (row: T) => string;
  renderRow: (row: T) => string;
};

The table does not need to know whether a row is a product, lecture, or user.

The caller supplies the domain-specific relationship once.


Utility types transform an existing contract

type ProductDraft = Omit<Product, "id">;
type ProductPreview = Pick<Product, "id" | "title">;
type EditableProduct = Partial<Product>;

Utility types are useful for views and transitions, but they should not replace clear domain states.


Strictness catches boundary mistakes early

Prefer a strict baseline:

{
  "strict": true,
  "noUncheckedIndexedAccess": true,
  "exactOptionalPropertyTypes": true
}

The exact set depends on the project, but the principle is stable: make uncertainty visible where it matters.


Avoid implicit any

function format(value) {
  return value.name;
}

An implicit any turns off the very feedback TypeScript is meant to provide.

If a value is unknown, write unknown and decide how to prove it.


Null and undefined are part of the model

function firstTitle(items: Product[]): string | undefined {
  return items[0]?.title;
}

const title = firstTitle(products) ?? "No products";

An empty result is not a failed type. It is a state the caller must handle.


Indexed access can be unsafe

const first = products[0];

if (first) {
  console.log(first.title);
}

Even when an array is typed as Product[], an arbitrary index may not contain a product.

noUncheckedIndexedAccess makes that possibility explicit.


DOM APIs are runtime boundaries too

const form = document.querySelector("#search-form");

if (!(form instanceof HTMLFormElement)) {
  throw new Error("Search form is missing");
}

The selector returns a nullable, broad value. Narrow it before using form-specific behavior.


Type the element and the event

const input = document.querySelector<HTMLInputElement>("#query");

input?.addEventListener("input", (event) => {
  const target = event.currentTarget as HTMLInputElement;
  console.log(target.value);
});

Use the most specific safe DOM type available, and remember that the element can still be absent.


The dangerous API shortcut

const data = (await response.json()) as Product[];

This is convenient at the exact point where convenience is most dangerous.

The network response is outside the compiler’s control. Treat it as unknown until it is parsed.


Assertions, casts, and conversions are different

const value = input as number;       // assertion: no runtime work
const number = Number(input);        // conversion: runtime work
const parsed = JSON.parse(input);    // parsing: produces a runtime value

An assertion changes static knowledge.

A conversion or parser changes or inspects a runtime value.


What counts as a trust boundary?

flowchart TD
    subgraph Boundaries["Untrusted External Trust Boundaries"]
        API[API responses]
        URL[URL & query string]
        Storage[Browser storage]
        Forms[Form input]
        Post[postMessage]
        SDK[Third-party SDKs]
    end

Every source outside the current trusted function may be malformed, stale, partial, or surprising.


URL values are strings, even when they look numeric

const rawPage = new URLSearchParams(location.search).get("page");
const page = rawPage === null ? 1 : Number(rawPage);

if (!Number.isInteger(page) || page < 1) {
  throw new Error("Invalid page");
}

Parsing and domain checks are separate decisions. Number("") becoming 0 may be a valid JavaScript conversion but an invalid page.


Browser storage is untrusted text

const raw = localStorage.getItem("preferences");
const value: unknown = raw === null ? null : JSON.parse(raw);

Storage may contain old versions, manually edited values, or invalid JSON.

Read it as an external payload, then validate a versioned shape.


Forms produce strings and absence

const formData = new FormData(form);
const rawEmail: unknown = formData.get("email");

if (typeof rawEmail !== "string" || !rawEmail.includes("@")) {
  return showError("Enter a valid email");
}

The HTML control type is not enough to guarantee the value your domain logic expects.


Manual validation is a parser

function parseProduct(value: unknown): Product {
  if (typeof value !== "object" || value === null) {
    throw new Error("Product must be an object");
  }

  const record = value as Record<string, unknown>;
  if (typeof record.id !== "string") throw new Error("Product id is invalid");
  if (typeof record.title !== "string") throw new Error("Product title is invalid");
  if (typeof record.priceCents !== "number") throw new Error("Product price is invalid");

  return { id: record.id, title: record.title, priceCents: record.priceCents };
}

The returned Product is earned by checks and reconstruction, not by a blind assertion.


Parse, then trust the result

const payload: unknown = await response.json();
const product = parseProduct(payload);

// Product-specific application code starts here.
renderProduct(product);

Keep parsing close to the boundary. Keep the rest of the application free from repeated shape checks.


Validation errors should help the caller

Useful errors identify:

  • which boundary failed;
  • which field was unexpected;
  • what kind of value was expected;
  • whether the response was malformed or the request failed.

Do not include tokens, passwords, or complete sensitive payloads in error messages.


Schema libraries make repetitive checks composable

const ProductSchema = z.object({
  id: z.string(),
  title: z.string(),
  priceCents: z.number().int().nonnegative(),
});

A schema can centralize parsing, reuse nested rules, report paths, and derive static types.

The library does not remove the need to design the domain contract.


Safe parsing returns a controlled result

const result = ProductSchema.safeParse(payload);

if (!result.success) {
  return { ok: false, error: result.error.issues };
}

return { ok: true, data: result.data };

This makes invalid input an explicit branch instead of an unexpected exception deep in rendering code.


Transport data is not always domain data

type ProductResponse = {
  product_id: string;
  display_name: string;
  price_cents: number;
};

type Product = {
  id: string;
  title: string;
  priceCents: number;
};

The boundary can validate the transport shape and map it into the vocabulary used by the application.


A complete API boundary

async function loadProducts(): Promise<ApiResult<Product[]>> {
  let response: Response;
  try {
    response = await fetch("/api/products");
  } catch {
    return { ok: false, error: "Network request failed" };
  }

  if (!response.ok) return { ok: false, error: `HTTP ${response.status}` };

  const payload: unknown = await response.json();
  return parseProducts(payload);
}

Transport failure and schema failure are separate facts and should remain distinguishable.


Keep raw data unknown at the edge

flowchart TD
    A[fetch / storage / URL / form] --> B[unknown]
    B --> C[parse + validate]
    C --> D[trusted transport data]
    D --> E[map to domain]
    E --> F[trusted domain model]
    F --> G[components and features]

The farther a value travels from the edge, the less useful it is to keep asking whether it is valid.


The boundary architecture

flowchart TD
    Adapter["Adapter: knows external format"] --> Parser["Parser: checks runtime shape"]
    Parser --> Mapper["Mapper: creates domain vocabulary"]
    Mapper --> App["Application: consumes trusted values"]
    App --> View["View: renders explicit state"]

Each layer has a small responsibility and a clear reason to change.


The parse-then-trust principle

function useProducts(products: Product[]) {
  // No API-shape checks here.
  return products.map(product => product.title);
}

Trust should be established once at the boundary, then preserved by types and module boundaries.

Repeated checks in every component usually signal that the boundary is in the wrong place.


satisfies checks without widening useful literals

const routes = {
  home: "/",
  playground: "/playground",
} satisfies Record<string, `/${string}`>;

The object is checked against the contract while retaining precise keys and values for later inference.

This is often better than annotating the whole object as Record<string, string>.


When assertions are legitimate

An assertion can be reasonable when:

  • a platform API is typed too broadly;
  • a checked invariant is understood by the compiler but not expressible locally;
  • a test fixture deliberately models a controlled case;
  • the assertion is close to the proof and documented.

It is risky when it is used to skip parsing, null checks, or domain decisions.


Avoid double assertions

const impossible = value as unknown as Product;

This pattern can force unrelated types together and hides the missing proof.

If the boundary is uncertain, keep the value as unknown and write the parser that makes the transition explicit.


Non-null assertions hide a missing state

const element = document.querySelector("#app")!;

The ! silences null without changing runtime behavior.

Prefer a checked lookup, a required-element helper, or an initialization path that makes absence impossible.


Structural typing is useful and subtle

type User = { id: string };
type Product = { id: string };

const product: Product = { id: "p-1" };
const user: User = product;

Both shapes satisfy the same structure, even if the domain meanings differ.

Shape compatibility does not automatically communicate semantic identity.


Branded types protect semantic identifiers

type ProductId = string & { readonly __brand: "ProductId" };
type UserId = string & { readonly __brand: "UserId" };

function loadProduct(id: ProductId) {}

The compiler can now distinguish two strings that represent different kinds of identifier.

The brand has no runtime representation.


Create a brand only after validation

function parseProductId(value: unknown): ProductId {
  if (typeof value !== "string" || !/^p-[a-z0-9-]+$/.test(value)) {
    throw new Error("Invalid product id");
  }
  return value as ProductId;
}

The assertion is local and justified by the parser’s runtime rule.

Never export a public “brand anything” helper that bypasses the boundary.


Brands are useful when the cost is justified

Consider them for:

  • identifiers that are frequently confused;
  • validated URLs or paths;
  • normalized currency codes;
  • security-sensitive tokens with distinct lifecycles.

Do not brand every primitive. Extra vocabulary should reduce real mistakes, not create ceremony.


Static contracts are not shared runtime validation

flowchart TD
    A["Shared TypeScript type: agreement during development"]
    B["Runtime schema: verification of delivered data"]
    C["Generated API types: synchronized description"]
    A -.->|requires| B
    C -.->|requires| B

Even generated types can become stale or be bypassed by a broken server.

The receiving application still owns its trust boundary.


Choose validation proportionately

flowchart TD
    A["Small local constant: direct type / simple check"]
    B["Stable internal module: typed constructor or parser"]
    C["External API: schema and useful errors"]
    D["Security-sensitive data: strict validation & logging policy"]

The goal is reliable boundaries, not maximum ceremony everywhere.


Typed errors preserve failure meaning

type BoundaryError =
  | { kind: "network"; message: string }
  | { kind: "http"; status: number }
  | { kind: "schema"; issues: string[] };

type Result<T> =
  | { ok: true; data: T }
  | { ok: false; error: BoundaryError };

Callers can render, retry, report, or ignore failures based on their actual cause.


TypeScript does not replace tests

Types catch many inconsistencies before execution.

Tests still verify behavior, boundary cases, compatibility, and user-visible outcomes.

Test malformed payloads, missing fields, old storage versions, invalid query strings, and empty results.


TypeScript does not replace security

Compile-time types do not sanitize HTML, authorize a user, protect a secret, or prevent a malicious server response.

Security checks must happen at runtime and at the correct trust boundary.


A malformed response should fail clearly

const response = {
  products: [
    { id: "p-1", title: "Valid", priceCents: 1000 },
    { id: "p-2", title: 42, priceCents: "free" },
  ],
};

Do not render the first item and silently accept the second.

Decide whether the contract is all-or-nothing, item-level recovery, or partial data with an explicit warning.


Replace the assertion with unknown

const payload: unknown = response;

const parsed = parseProducts(payload);
if (!parsed.ok) {
  showBoundaryError(parsed.error);
}

The compiler now forces the design to answer what happens when the payload is not valid.


Validate, then construct the trusted model

function parseProducts(value: unknown): ApiResult<Product[]> {
  if (!Array.isArray(value)) {
    return { ok: false, error: "products must be an array" };
  }

  const products = value.map(parseProduct);
  return { ok: true, data: products };
}

In production, make the item failure shape explicit and avoid allowing an exception to erase useful context.


The trusted core should stay boring

function total(products: Product[]) {
  return products.reduce((sum, product) => sum + product.priceCents, 0);
}

Once data is trusted, domain functions should focus on domain behavior.

If every function still checks whether priceCents is a number, the boundary has not done enough work.


Common misconceptions

MisconceptionBetter mental model
as Product validates JSONIt only changes static interpretation
unknown is inconvenientIt marks work the boundary must do
any is fasterIt removes useful feedback
shared types guarantee the APIThey describe an agreement, not delivery
schemas make domain design unnecessaryThey validate a contract you still choose
strictness means more code everywhereIt makes important uncertainty visible

Practical lab: Build a Safe Data Boundary

Create a small API boundary that parses unknown, validates the response, and exposes trusted product data to application code.

The practical moves from static modeling to malformed payloads, manual parsing, schema validation, URL and storage boundaries, and branded IDs.


Practical stages 1–3: model uncertainty

  1. Model the domain product and its load states.
  2. Test inference, literal types, and generic result types.
  3. Introduce unknown where the API response enters.

Do not begin by asserting that the response is already a Product[].


Practical stages 4–6: make state and failure explicit

  1. Render every discriminated load state exhaustively.
  2. Add a generic ApiResult<T>.
  3. Create an intentionally unsafe API response with valid and malformed records.

Observe how the unsafe version spreads assumptions into the UI.


Practical stages 7–9: replace trust with parsing

  1. Replace the assertion with unknown.
  2. Write a manual parser with useful field errors.
  3. Compare the parser with a schema library.

Record the maintenance trade-off: readability, reuse, error paths, dependency cost, and change frequency.


Practical stages 10–13: finish the boundary map

  1. Validate URL state.
  2. Validate browser storage.
  3. Create branded IDs only after validation.
  4. Draw the trust architecture from raw input to trusted domain data.

Verification: invalid data cannot enter the trusted model silently.


Try this yourself

Add a new archived product state and update every renderer.

Then add a server field that is optional in the transport response but required by the domain model.

Where should the default be applied?

At the parser, mapper, or component? Explain the boundary decision.


Troubleshooting guide

SymptomLikely cause
Everything became anyAn untyped boundary leaked inward
Assertions appear everywhereParsing is missing or too far away
Components repeat shape checksTrusted data is not established once
Union branches feel awkwardAdd a discriminant or redesign the state
A guard compiles but fails in productionThe predicate claims more than it checks
Old storage breaks after a releaseAdd versions and migration/validation rules

Completion checklist

  • local values rely on inference where appropriate;
  • domain states use explicit, meaningful unions;
  • external values enter as unknown;
  • parsers return useful failure information;
  • transport and domain models are separated where needed;
  • branded values are created only after validation;
  • tests cover malformed and stale inputs;
  • trusted core functions do not repeat boundary checks.

The chapter in one sentence

Use TypeScript to describe and connect trusted program states; use runtime validation to earn trust at every external boundary.


Next: Chapter 6

The next chapter turns these contracts into component architecture:

  • component boundaries;
  • props and composition;
  • controlled and uncontrolled state;
  • reusable UI contracts;
  • accessibility as a design constraint;
  • testing behavior at the boundary.

Questions

Which value in your current application is trusted only because an assertion says it is?

That is the first boundary worth making explicit.

6 - Component-Driven Architecture & Design Patterns

Chapter 6: find component boundaries, design stable APIs, compose behavior, and organize front-end systems by responsibility.

Component-Driven Architecture & Design Patterns

Make responsibility visible

Chapter 6

Polla Fattah


Today’s goal

Learn to decide where a component should begin and end.

We will connect:

  • responsibility and reasoning boundaries;
  • inputs, outputs, props, events, and slots;
  • controlled and uncontrolled state;
  • compound and headless components;
  • context and dependency injection;
  • reuse, abstraction, and duplication;
  • feature, layer, and domain-oriented organization;
  • accessibility, performance, internationalization, and security.

By the end of today you can

  • identify a coherent component responsibility;
  • distinguish a component boundary from a state boundary;
  • design intent-oriented public APIs;
  • choose controlled or uncontrolled ownership deliberately;
  • compose behavior without forcing styling decisions;
  • use context or injection without hiding dependencies;
  • recognize oversized and over-fragmented components;
  • extract components in a safe sequence;
  • organize a front-end by feature, layer, or domain;
  • refactor a catalogue into a maintainable component architecture.

The central principle

A component should exist because it owns a coherent responsibility - not merely because some markup can be extracted into another file.

The file boundary is an implementation detail.

The responsibility boundary is the architecture.


The chapter’s progression

flowchart TD
    A[Complex Interface] --> B[Identify Responsibilities]
    B --> C[Find Boundaries]
    C --> D[Define Inputs & Outputs]
    D --> E[Compose Components]
    E --> F[Decide State Ownership]
    F --> G[Share Dependencies Carefully]
    G --> H[Organize by Feature, Domain, or System]

Every step reduces accidental coupling while preserving understandable flow.


Why components exist

Consider a catalogue page containing:

  • navigation;
  • search and filters;
  • product cards;
  • pagination;
  • basket actions;
  • loading and error states;
  • keyboard behavior and analytics.

A single function can render it.

That does not mean one function is a good unit of reasoning.


Components reduce the amount we understand at once

flowchart TD
    Page["CataloguePage"]
    Page --> Search["SearchControls"]
    Page --> Filter["FilterPanel"]
    Page --> Grid["ProductGrid"]
    Page --> Paging["Pagination"]
    Grid --> Card["ProductCard"]
    Card --> Price["Price"]
    Card --> Stock["StockStatus"]
    Card --> AddBtn["AddToCartButton"]

If price formatting changes, we should not need to understand pagination, filters, and cart state at the same time.


Reuse is useful, but not required

AnnualBudgetApprovalPanel

It may appear once and still deserve a component because it owns:

  • a coherent domain responsibility;
  • substantial behavior;
  • a meaningful state model;
  • a natural testing boundary.

Reuse is one reason to create a component. Responsibility is stronger.


Isolation is encapsulation

A date picker can expose:

value
onChange
disabled

while hiding:

  • calendar-grid generation;
  • month navigation;
  • keyboard behavior;
  • focus management;
  • date-cell implementation.

The caller gets what it needs without controlling every internal detail.


Testing boundaries follow behavior

A date picker can be tested for:

  • selected-date behavior;
  • keyboard navigation;
  • disabled dates;
  • accessible labels.

The surrounding page should not have to reproduce every internal interaction to test the date picker.

Test where behavior and risk are concentrated - not merely where files happen to exist.


Ownership makes team work clearer

A useful component boundary can make it clear:

  • which team owns the behavior;
  • which API is stable;
  • which implementation is private;
  • which changes require coordination.

Architecture is also communication between people, not only organization for the compiler.


Bad extreme: one giant component

CataloguePage
  search state
  filter state
  fetch logic
  URL synchronization
  product rendering
  cart events
  modal behavior
  analytics
  keyboard handling
  responsive layout

Typical symptoms:

  • everything can reach everything;
  • a small change creates a large regression surface;
  • tests require excessive setup;
  • state ownership is ambiguous.

Bad extreme: component explosion

flowchart TD
    Page["Page"] --> Section["Section"]
    Section --> Wrapper["Wrapper"]
    Wrapper --> Row["Row"]
    Row --> Text["Text"]
    Text --> Label["Label"]

Small is not automatically good.

Too many trivial components increase navigation cost, obscure control flow, and make a simple change require a tour of the repository.


Find boundaries by asking five questions

  1. What changes together?
  2. What owns the behavior?
  3. What represents one domain concept?
  4. What is genuinely reusable?
  5. What should remain private?

The answers should describe responsibility, not just the shape of the markup.


What changes together?

If a visual structure, its interaction rules, and its tests always change together, they may belong in one component.

If a value, layout, and data-fetching policy change independently, forcing them into one unit creates coupling.

Change patterns are evidence for boundaries.


What owns the behavior?

SearchInput       owns input interaction
SearchController  owns query coordination
ProductGrid       owns product layout
ProductCard       owns product presentation

Do not move state merely because a child renders the element.

The owner is the unit that makes the decision and coordinates its consequences.


What represents one domain concept?

ProductCard
ApprovalPanel
ShippingAddress
LectureNavigation

Domain concepts are often stronger boundaries than generic visual fragments.

They give the API meaningful vocabulary and give tests a clear subject.


What should remain private?

Not every internal helper needs to become a public component.

Keep implementation details local when they:

  • have no independent responsibility;
  • are used in one place;
  • would expose unstable structure;
  • make the public API harder to understand.

Local components are valid architecture.


A deliberately monolithic shape

function CataloguePage() {
  const [query, setQuery] = useState("");
  const [filters, setFilters] = useState(defaultFilters);
  const { data, loading, error } = useProducts(query, filters);

  return (
    <main>
      {/* search, filters, cards, pagination, errors, and cart */}
    </main>
  );
}

The problem is not that it renders JSX.

The problem is that unrelated responsibilities have no visible boundaries.


First refactoring: stable visual responsibilities

function CataloguePage() {
  return (
    <main>
      <SearchControls />
      <FilterPanel />
      <ProductGrid />
      <Pagination />
    </main>
  );
}

Extract the stable responsibilities first.

Keep coordination in the parent until the data and ownership decisions are understood.


A component boundary is not necessarily a state boundary

ProductCard       renders one product
CataloguePage     owns selected products and filters
CartStore         owns basket state

A child can be a useful rendering boundary while state remains higher in the tree.

Do not force every component to own every value it displays.


Inputs and outputs define the contract

flowchart LR
    Inputs["Inputs / Props\n(Outside Values)"] --> Comp["Component\n(Encapsulated Logic)"]
    Comp --> UI["Rendered Output\n(Visible Result)"]
    Comp --> Outputs["Outputs / Callbacks\n(Events & Intent)"]

The public contract should be smaller and more stable than the implementation.


Props describe data and capabilities

type ProductCardProps = {
  product: Product;
  onAddToCart: (productId: ProductId) => void;
  isAdded: boolean;
};

Good props answer:

  • what does this component need?
  • what decisions can it request?
  • what state is intentionally owned elsewhere?

Vue props and events express the same boundary

<ProductCard
  :product="product"
  :is-added="isAdded"
  @add-to-cart="addToCart"
  />

The syntax differs from React.

The architectural question is the same: what enters, what leaves, and who owns the decision?


Callbacks and events should communicate intent

Prefer:

onAddToCart(productId)
onSearch(query)
onApprove(requestId)

Over exposing implementation details:

onButtonClick(event)
setInternalCartState(nextState)

Intent-oriented APIs let the component change its internal markup without breaking callers.


Children and slots enable composition

<Panel>
  <Panel.Header>Filters</Panel.Header>
  <Panel.Body>
    <FilterForm />
  </Panel.Body>
</Panel>

The panel owns structure and semantics.

The caller supplies content that varies.


Named slots make variation explicit

<ProductCard>
  <template #badge>
    <StockStatus />
  </template>
  <template #actions>
    <AddToCartButton />
  </template>
</ProductCard>

Named composition points communicate where variation belongs without adding a boolean prop for every possibility.


Composition over inheritance

flowchart TD
    Card["Card Container"]
    Card --> Header["Header (Slot / Child)"]
    Card --> Body["Body (Slot / Child)"]
    Card --> Actions["Actions (Slot / Child)"]

Compose small capabilities and content rather than building a deep inheritance hierarchy.

Composition keeps variation near the caller and avoids base classes accumulating unrelated assumptions.


Wrapper components should add a reason

A wrapper is valuable when it adds:

  • semantics;
  • layout responsibility;
  • accessibility behavior;
  • a state boundary;
  • a stable composition API.

If it only forwards every prop and renders one child unchanged, it may be navigation cost without architectural value.


Controlled components: the parent owns the value

<SearchInput
  value={query}
  onChange={setQuery}
  placeholder="Search products"
/>

The component renders and reports interaction.

The parent owns the source of truth and can synchronize it with URL state, data fetching, or another control.


Uncontrolled components: the component owns the value

<SearchInput
  defaultValue=""
  onSubmit={submitSearch}
/>

The component manages its internal editing state.

This can simplify local interactions when the parent does not need every intermediate value.


Controlled versus uncontrolled is an ownership decision

flowchart TD
    A{"State Ownership Need"}
    A -- Synchronize with external state / URL --> B["Controlled Component\n(Parent owns state via props & callbacks)"]
    A -- Only final submission value required --> C["Uncontrolled Component\n(Component manages internal state)"]
    A -- Both local reactivity & parent control --> D["Deliberate Bridge Pattern\n(Explicit value + onChange synchronization)"]

Neither mode is universally better.

Choose based on who must make decisions about the state.


Avoid half-controlled APIs

<Tabs value={value} defaultValue="overview" onChange={setValue} />

What wins if value is absent? What happens when it appears later?

Ambiguous ownership creates warnings, stale state, and surprising transitions.

Use separate APIs or document the controlled/uncontrolled contract precisely.


Good component APIs minimize invalid combinations

Instead of:

disabled + loading + error + success + compact + outline + danger + iconOnly

Model meaningful states and relationships.

type SubmitButtonProps =
  | { status: "ready"; onSubmit: () => void }
  | { status: "loading" }
  | { status: "error"; message: string; onRetry: () => void };

Keep the public surface small and stable

flowchart TD
    subgraph PublicContract["Public Interface (Contract)"]
        Props["Props / Attributes"]
        Events["Events / Callbacks"]
        Slots["Slots / Children"]
    end
    subgraph PrivateImpl["Private Implementation (Encapsulated)"]
        State["Internal State"]
        Helpers["Private Helpers & Handlers"]
        DOM["Internal DOM Nodes"]
    end
    PublicContract --> PrivateImpl

Every public prop is a promise.

Every public event is a dependency for callers.

Expose the smallest contract that supports the component’s responsibility.


Compound components share a local vocabulary

<Tabs>
  <Tabs.List>
    <Tabs.Tab value="overview">Overview</Tabs.Tab>
    <Tabs.Tab value="reviews">Reviews</Tabs.Tab>
  </Tabs.List>
  <Tabs.Panel value="overview">...</Tabs.Panel>
  <Tabs.Panel value="reviews">...</Tabs.Panel>
</Tabs>

The pieces are separate in markup but belong to one conceptual system.


Why compound components can help

They can provide:

  • readable structure;
  • explicit composition points;
  • shared state without prop repetition;
  • a constrained vocabulary;
  • flexible placement of tabs and panels.

The API communicates the relationship between the parts.


The cost of compound components

They also add:

  • hidden coordination rules;
  • context or injection dependencies;
  • more concepts to document;
  • more invalid combinations to prevent;
  • debugging work when pieces are used incorrectly.

Use the pattern when the relationship is real and repeated - not because the API looks advanced.


Headless components separate behavior from styling

flowchart TD
    subgraph HeadlessLogic["Headless Behavior Layer"]
        KB["Keyboard Navigation Rules"]
        Sel["Selection State Machine"]
        ARIA["ARIA Attributes & Relationships"]
        Focus["Focus Management"]
    end
    subgraph ConsumerUI["Consumer UI Layer"]
        Markup["Consumer-Owned Markup"]
        Styles["Tailwind / Custom CSS Styles"]
        Layout["Flexible Component Layout"]
    end
    HeadlessLogic --> ConsumerUI

A headless component supplies interaction logic without imposing a visual design.


Why headless UI exists

It is useful when teams need:

  • shared accessibility behavior;
  • different visual systems;
  • consistent keyboard interaction;
  • application-specific layout;
  • reusable state machines.

The abstraction is behavioral, not merely visual.


Dependency sharing: explicit props first

<ProductCard
  product={product}
  currency={currency}
  locale={locale}
  onAddToCart={onAddToCart}
/>

Explicit inputs make dependencies visible and make the component easy to render in isolation.

Prop passing becomes a problem only when it is genuinely burdensome or crosses unrelated layers repeatedly.


React context shares a local dependency

const TabsContext = createContext<TabsContextValue | null>(null);

function useTabsContext() {
  const context = useContext(TabsContext);
  if (!context) throw new Error("Tabs parts must be inside Tabs");
  return context;
}

Context can keep compound parts coordinated without repeating the same props at every level.


Vue provide/inject expresses the same idea

provide(TabsKey, {
  selected,
  select,
});

const tabs = inject(TabsKey);

The framework syntax changes; the architectural trade-off remains: the dependency is less visible at the call site.


Context is not automatically better than props

Context can:

  • hide where a value comes from;
  • make isolated rendering harder;
  • widen a component’s implicit dependency surface;
  • cause broad updates when the value changes.

Use it for a real shared relationship, not just to avoid writing one more prop.


Dependency injection makes variation explicit at setup

type ProductRepository = {
  list(): Promise<Product[]>;
};

function createCatalogue(repository: ProductRepository) {
  return { load: () => repository.list() };
}

The feature depends on a capability, not on one concrete network implementation.


Why dependency injection helps

It can improve:

  • testing with fakes;
  • environment-specific adapters;
  • separation of domain behavior from transport;
  • migration between implementations.

The dependency remains a design decision rather than a hidden import.


Dependency injection can also be overused

If every helper receives a container with dozens of services, the dependency graph becomes harder to understand.

Inject the smallest capability needed.

Prefer a direct import for a stable, genuinely global constant when injection adds no meaningful variation.


Reusable UI components versus application components

Reusable: Button, Dialog, Tabs, Field
Application: ProductCard, CatalogueFilters, ApprovalPanel

Reusable UI components should avoid application-specific assumptions.

Application components should speak the domain language and may coordinate several reusable primitives.


Reuse has levels

flowchart TD
    A["One Feature Boundary\n(Lowest cost, highest velocity)"] --> B["One Application Boundary\n(Shared across feature teams)"]
    B --> C["Multiple Products Boundary\n(Multi-app shared packages)"]
    C --> D["Global Design System Package\n(Highest contract cost, strict versioning)"]

The wider the reuse boundary, the more expensive the public contract becomes.

Do not design a global package API for a problem that only exists in one feature.


Premature abstraction freezes assumptions

An abstraction created before the second use often encodes:

  • accidental naming;
  • the first layout’s constraints;
  • one feature’s state model;
  • props that do not generalize.

Wait until the shared behavior and variation are understood.


Duplication can be cheaper than the wrong abstraction

Two similar components may differ in:

  • ownership;
  • accessibility requirements;
  • lifecycle;
  • domain vocabulary;
  • future change direction.

Temporary duplication preserves independent evolution.

Remove duplication when the shared concept - not only the current markup - is real.


Design methodologies are lenses, not laws

Atomic design, feature folders, layers, and domain modules can all be useful.

None can decide a boundary without understanding:

  • change patterns;
  • ownership;
  • dependencies;
  • product vocabulary;
  • team constraints.

Use a methodology to ask better questions, not to avoid judgment.


Atomic design: strength and limitation

Atomic design encourages a vocabulary from primitives to composed interfaces.

It can help teams discover reusable visual patterns.

But visual size does not always match responsibility.

An “organism” may be a domain feature, while a tiny “atom” may still contain complex behavior.


Feature-oriented decomposition

features/catalogue/
  components/
  state/
  api/
  tests/
features/cart/
  components/
  state/
  api/
  tests/

Feature organization keeps the code that changes together close together.

It is often a strong default for application-scale work.


Layer-oriented structure

components/
hooks/
services/
state/
utils/
pages/

Layer organization makes technical roles easy to scan.

Its risk is scattering one feature across many directories and weakening domain ownership.


Domain-oriented decomposition

catalogue/
  Product.ts
  ProductCard.tsx
  catalogue-api.ts
  catalogue-state.ts
  catalogue.test.ts
cart/
  Cart.ts
  CartSummary.tsx

Domain organization keeps vocabulary, behavior, and tests near the concept they serve.


Technical reuse still crosses domains

shared/ui/Button
shared/ui/Dialog
shared/forms/Field

A reusable technical primitive can be shared across domains without forcing domain components into one generic model.

The shared layer should remain intentionally small.


Domain components should speak domain language

Prefer:

<AddToCartButton productId={product.id} />

Over:

<Button onClick={() => dispatch({ type: "CART_ADD", payload: product })} />

The domain component hides coordination details and exposes the intent relevant to its caller.


Avoid boolean prop explosion

<Button primary compact rounded loading danger iconOnly />

Many booleans create a combinatorial API and states nobody designed.

Prefer meaningful variants or modeled states:

type ButtonVariant = "primary" | "danger" | "quiet";
type ButtonState = "ready" | "loading" | "disabled";

Stable components should not depend on the whole application

A reusable component should not know:

  • the entire route tree;
  • the global store shape;
  • the current user object;
  • every feature’s analytics policy.

Pass the smallest data and capabilities needed.

Application coordination belongs above the reusable boundary.


Smart and presentational is a useful distinction

flowchart TD
    Container["CatalogueContainer\n(Fetches data, owns state, coordinates features)"]
    Grid["ProductGrid\n(Lays out collection, handles viewport)"]
    Card["ProductCard\n(Presents one item, emits user intent)"]
    Container --> Grid
    Grid --> Card

The names are less important than the separation of coordination from presentation.


React catalogue architecture

function CataloguePage() {
  const state = useCatalogueState();

  return (
    <CatalogueLayout>
      <SearchControls value={state.query} onChange={state.setQuery} />
      <ProductGrid products={state.products} onAdd={state.addToCart} />
    </CatalogueLayout>
  );
}

The page coordinates. The children expose focused responsibilities.


React product grid

function ProductGrid({ products, onAdd }: ProductGridProps) {
  return (
    <ul className="product-grid">
      {products.map(product => (
        <li key={product.id}>
          <ProductCard product={product} onAddToCart={onAdd} />
        </li>
      ))}
    </ul>
  );
}

The grid owns collection layout.

The card owns one product’s presentation and interaction surface.


Vue can express the same architecture

<template>
  <CatalogueLayout>
    <SearchControls v-model="query" />
    <ProductGrid :products="products" @add-to-cart="addToCart" />
  </CatalogueLayout>
</template>

Compare the responsibilities and data flow, not the framework punctuation.


Compare architecture, not syntax

flowchart LR
    subgraph React["React Paradigm"]
        RProps["Props + Callbacks"]
        RChild["Children"]
        RCtx["React Context"]
        RHooks["Custom Hooks"]
    end
    subgraph Vue["Vue Paradigm"]
        VProps["Props + Emits"]
        VSlots["Slots"]
        VPI["provide / inject"]
        VComp["Composables"]
    end
    RProps <-->|Architectural Equivalent| VProps
    RChild <-->|Architectural Equivalent| VSlots
    RCtx <-->|Architectural Equivalent| VPI
    RHooks <-->|Architectural Equivalent| VComp

Frameworks provide mechanisms.

They do not decide ownership, coupling, or the right domain boundary for you.


Recognize an oversized component

Warning signs:

  • many unrelated state variables;
  • long conditional render branches;
  • repeated markup with slightly different behavior;
  • effects that coordinate unrelated systems;
  • tests that require the entire application setup;
  • a name that no longer describes one responsibility.

Extract by responsibility, not by line count alone.


Recognize over-fragmentation

Warning signs:

  • a component has no meaningful API;
  • understanding one behavior requires many file jumps;
  • props are forwarded unchanged through several layers;
  • every markup element has its own file;
  • local details become global vocabulary.

Navigation cost is part of the architecture.


Use the extraction test

Ask:

  • Does it have its own responsibility?
  • Does it have a meaningful public API?
  • Does it change for a different reason?
  • Is it reused or likely to be reused for a real reason?
  • Will extraction reduce or increase navigation cost?

If most answers are no, keep it local for now.


Colocation keeps private behavior near its owner

features/catalogue/
  ProductCard.tsx
  ProductCard.test.tsx
  ProductCard.module.css
  product-card-format.ts

Colocation shortens the path from behavior to styles to tests.

Move files outward only when they become shared or when the repository’s structure makes ownership clearer elsewhere.


Component boundaries and performance

Boundaries can help with:

  • reducing rerender scope;
  • memoizing stable subtrees;
  • lazy-loading feature code;
  • isolating expensive calculations.

But splitting every element is not a performance strategy.

Measure the actual bottleneck and preserve a clear data flow.


Component boundaries and accessibility

An accessible pattern has a behavioral contract:

  • roles and relationships;
  • keyboard interaction;
  • focus movement;
  • names and descriptions;
  • disabled and busy states.

Keep these rules with the component that owns the interaction, especially for dialogs, tabs, menus, and composite widgets.


Component boundaries and internationalization

Do not bury locale assumptions in generic components.

<Price amount={product.priceCents} currency={currency} locale={locale} />

The component can format correctly while the feature decides which domain value and locale it is displaying.

Text, plural rules, direction, and date conventions are part of the public behavior.


Component boundaries and security

Keep trust-sensitive behavior near the boundary that understands it.

flowchart LR
    External["External / User Content"] --> Sanitize["Sanitize & Validate Boundary"]
    Sanitize --> Trusted["Trusted Display Component"]

Do not make a generic renderer responsible for deciding whether arbitrary HTML is safe.

Make the safe path the easiest public API.


The public API is a contract

Changing a public prop can affect:

  • callers;
  • tests;
  • stories and examples;
  • accessibility behavior;
  • analytics;
  • downstream packages.

Treat component APIs with the same care as a service boundary: explicit, intentional, and proportionate.


A component is not automatically a design-system component

flowchart LR
    App["Application Component"] --- AppDesc["Speaks one product's domain language"]
    DS["Design-System Component"] --- DSDesc["Cross-product stable behavior & appearance"]

Promoting a local component too early creates a public API before variation is understood.

Keep application components local until real reuse proves the broader boundary.


One possible layering model

flowchart TD
    Shell["Application Shell"] --> Feature["Feature Components"]
    Feature --> Domain["Domain Components"]
    Domain --> Adapter["Shared Behavior & Data Adapters"]
    Adapter --> UI["Reusable UI Primitives"]

The dependency direction should be intentional.

Lower-level primitives should not import feature-specific decisions.


Practical refactoring strategy

  1. Identify responsibilities.
  2. Identify shared state.
  3. Extract stable visual responsibilities.
  4. Keep coordination in the parent initially.
  5. Refine APIs.
  6. Move truly reusable primitives downward.
  7. Add context or injection only when explicit passing is genuinely burdensome.

Refactoring is a sequence of smaller decisions, not a single rewrite.


Refactor the monolith in observable steps

flowchart TD
    Monolith["Monolithic Interface"] --> Layout["1. Stable Layout Shell"]
    Layout --> Collection["2. Product Collection Boundary"]
    Collection --> Item["3. Product Item Boundary (ProductCard)"]
    Item --> Search["4. Controlled Search & Filter Boundary"]
    Search --> Composition["5. Composition Slot / Children Boundary"]
    Composition --> Shared["6. Shared Primitives (Only where justified)"]

After each step, preserve behavior and re-check ownership.


Practical lab: Compound Headless Tabs

Build a composable tabs system that separates keyboard behavior and state ownership from styling and content layout.

The practical applies the chapter’s component API, controlled state, compound composition, accessibility, and headless behavior principles.


Practical stages 1–3: name responsibilities

  1. Define the responsibility of the root, tab list, tab, and panel.
  2. Implement controlled and uncontrolled selection.
  3. Add role, aria-selected, aria-controls, and tabindex behavior.

Start with the contract before choosing context, slots, or styling.


Practical stages 4–6: test the interaction

  1. Test arrow-key navigation, disabled tabs, and dynamic panels.
  2. Compare explicit props, context/provide-inject, and a headless API.
  3. Verify that selected tab and visible panel remain synchronized.

The keyboard behavior should be independently testable from the visual treatment.


Practical extension: lazy panels

Add lazy panel loading.

Document the states explicitly:

flowchart LR
    Unrequested["Not Requested"] --> Loading["Loading"]
    Loading --> Ready["Ready"]
    Loading --> ErrorState["Error"]
    ErrorState -->|User Retry| Loading

The tabs contract should make loading and failure visible rather than hiding them inside a boolean prop combination.


Try this yourself

Add a disabled tab that:

  • cannot receive selection;
  • remains represented in the tab order according to your chosen accessibility policy;
  • exposes its disabled state;
  • does not load its panel;
  • does not break arrow navigation.

Write the behavioral contract before the implementation.


Troubleshooting guide (Part 1)

SymptomLikely cause
Props are forwarded through many layersOwnership or a real shared dependency is unclear
Every component has many booleansInvalid combinations are not modeled
A child and parent fight over stateThe controlled contract is ambiguous
Context appears everywhereDependencies are hidden instead of designed

Troubleshooting guide (Part 2)

SymptomLikely cause
Reusable component needs feature knowledgeThe boundary is too low or too broad
Refactor creates dozens of filesExtraction followed markup, not responsibility
Keyboard behavior is duplicatedInteraction logic lacks one owner

Completion checklist

  • each extracted component has a coherent responsibility;
  • public props and events express intent;
  • state ownership is explicit;
  • controlled and uncontrolled modes are not mixed accidentally;
  • compound APIs have a real relationship to model;
  • headless behavior is independent from styling;
  • context or injection is used only where it clarifies composition;
  • accessibility behavior belongs with its interaction owner;
  • the final structure reflects feature or domain change patterns.

Misconceptions to leave behind (Part 1)

MisconceptionBetter mental model
Every repeated markup needs a componentResponsibility and change patterns matter
Components exist mainly for reuseReasoning, isolation, and ownership matter too
Smaller components are always betterNavigation cost is part of quality
Each component owns all its stateThe owner is the decision-making unit

Misconceptions to leave behind (Part 2)

MisconceptionBetter mental model
More props mean more flexibilityMore public combinations mean more obligations
Context is better than prop drillingContext trades repetition for visibility
Reusable means genericReuse should preserve meaningful vocabulary
A framework decides architectureFrameworks provide mechanisms, not boundaries

The chapter in one sentence

Design components around coherent responsibility, explicit ownership, and small stable contracts; let composition provide variation.


Next: Chapter 7

The next chapter will build on component architecture with:

  • application state and data flow;
  • URL and server state;
  • local versus shared state;
  • synchronization and derived state;
  • predictable updates and debugging.

Questions

Which component in your current application owns too many decisions?

Which one has an API so generic that its domain responsibility is no longer visible?

7 - Reactivity & Rendering Mechanics

Chapter 7: follow state through render trees, dependency graphs, scheduling, identity, and side-effect boundaries.

Reactivity & Rendering Mechanics

Trace state to the screen

Chapter 7

Polla Fattah


Today’s goal

Understand what happens between a state change and the UI the user sees.

We will connect:

  • state-driven rendering;
  • snapshots, batching, and identity;
  • reconciliation and keys;
  • derived values and memoization;
  • effects and cleanup;
  • React and Vue mental models;
  • fine-grained reactivity and signals;
  • scheduling, measurement, and unnecessary work.

By the end of today you can

  • describe UI as a function of state;
  • distinguish render calculation from DOM commitment;
  • explain reconciliation and component identity;
  • use keys to preserve or reset state intentionally;
  • update from previous state safely;
  • separate source state from derived values;
  • reserve effects and watchers for external synchronization;
  • explain React snapshots, Vue proxies, computed values, and batching;
  • trace a reactive dependency graph;
  • optimize only after measuring the actual work.

The central lesson

flowchart TD
    A["Source State"] -->|dependencies| B["Render / Derivation"]
    B -->|scheduling| C["Commit / Synchronization"]
    C --> D["Screen & External Systems"]

Reactivity is not magic.

It is a system that tracks relationships, schedules work, preserves identity, and crosses explicit side-effect boundaries.


The chapter’s progression

flowchart TD
    A[Manual DOM Updates] --> B[State-Driven UI]
    B --> C[Render & Reconciliation]
    C --> D[Identity & Snapshots]
    D --> E[Derived State & Effects]
    E --> F[React Mechanics]
    F --> G[Vue Mechanics]
    G --> H[Fine-Grained Systems & Signals]
    H --> I[Scheduling & Measurement]

The same UI goal can be implemented with different reactive mechanisms.


Manual DOM updates are imperative

let count = 0;

button.addEventListener("click", () => {
  count += 1;
  countElement.textContent = String(count);
});

The event handler must remember every DOM node affected by the change.

As the interface grows, manual synchronization becomes a coordination problem.


State-driven UI describes a result

function renderCount(count) {
  countElement.textContent = String(count);
}

function increment() {
  count += 1;
  renderCount(count);
}

The UI is derived from state rather than updated through a growing list of unrelated DOM instructions.


UI as a function of state

UI = render(state)

The function is conceptual, not necessarily a literal full-page rewrite.

It gives the system a useful question:

Given this state, what should the interface represent?


Reactivity does not mean everything runs automatically

A reactive system must decide:

  • what depends on what;
  • what work is invalidated;
  • when work is scheduled;
  • what can be reused;
  • what must be synchronized;
  • when cleanup runs.

“Reactive” describes a mechanism, not a guarantee that every operation is free or immediate.


State is not the same as every variable

const label = "Products";       // stable configuration
let localCounter = 0;            // ordinary mutable variable
const [query, setQuery] = useState(""); // render-relevant state

Use state when a change must participate in the UI’s update model.

Do not put every value in state just because it changes somewhere.


A small running example

query: "phone"
products: 120 records
filter: "in-stock"

The visible list depends on all three values.

The selected count may be derived.

The network request is an external operation.

The reactive architecture should make each relationship visible.


The React mental model

React calls component functions to calculate a description of UI.

function Counter({ count }: { count: number }) {
  return <output>{count}</output>;
}

Calling the function does not mean the browser DOM is immediately rewritten.


A React render is a calculation

flowchart TD
    Update["State Update Triggered"] --> Call["React calls component function"]
    Call --> Desc["New Element Descriptions (VNodes)"]
    Desc --> Reconcile["Reconciliation / Diffing Engine"]
    Reconcile --> Commit["Commit Necessary Host DOM Changes"]

The render phase should be free of observable side effects.


Render should behave like a pure calculation

function ProductCount({ products }: Props) {
  // Good: calculate a value
  const count = products.length;
  return <span>{count}</span>;
}

Avoid in render:

  • network requests;
  • subscriptions;
  • DOM mutation;
  • timers;
  • analytics calls.

Reconciliation compares descriptions

flowchart LR
    subgraph Prev["Previous VNode Tree"]
        Old["&lt;h1&gt;Old&lt;/h1&gt;"]
    end
    subgraph Next["Next VNode Tree"]
        New["&lt;h1&gt;New&lt;/h1&gt;"]
    end
    Prev -->|Diff / Reconciliation| Patch["DOM Mutation: textContent = 'New'"]
    Next -.-> Patch

The framework compares what was previously described with what is now described, then determines the smallest host update required.

The component function may run even when the DOM change is tiny or absent.


Virtual DOM without the mythology

“Virtual DOM” is a representation and comparison strategy.

It does not mean:

  • the entire real DOM is rebuilt every time;
  • virtual work is automatically faster than all alternatives;
  • every render creates a visible browser update;
  • architecture no longer matters.

Measure the work that actually matters.


The commit phase changes external reality

After reconciliation, the framework commits required changes to the host environment.

flowchart LR
    RenderPhase["Render Phase: Pure Calculation\n(Calculate VNodes; zero DOM side-effects)"] --> CommitPhase["Commit Phase: Host Mutation\n(Apply DOM diffs, attach refs, run effects)"]

Keeping these phases conceptually separate explains why render code should remain predictable.


Parent rendering and child rendering

When a parent renders, its child functions may be called again.

That does not automatically mean:

  • every child DOM node changed;
  • every child state reset;
  • every expensive calculation must rerun forever.

Rendering, reconciliation, commitment, and state preservation are related but distinct questions.


Component identity is part of behavior

flowchart TD
    Check{"Component Check"}
    Check -- Same component type + same key + same position --> Preserve["Preserve state & instance"]
    Check -- Different component type OR different key --> Destroy["Destroy old instance & initialize new state"]

Identity determines whether a component is treated as the same logical instance.

This affects focus, input values, animations, and user experience - not only performance.


State is associated with a position in the render tree

{showDetails && <DetailsForm />}

If the same component remains at the same logical position, its state may persist across parent renders.

Changing the structure or key can intentionally create a new identity.


Keys are about identity, not silence

{items.map(item => (
  <ProductCard key={item.id} product={item} />
))}

The key tells the renderer which item is which across changes.

It is not merely a way to remove a warning.


Why array index keys can be dangerous

items.map((item, index) => <Row key={index} item={item} />)

If items are inserted, removed, or reordered, the index can now identify a different item.

Local state may move to the wrong row.

Use a stable identity from the data whenever the list can change.


Keys can intentionally reset state

<Editor key={documentId} document={document} />

When documentId changes, the editor receives a new identity.

This can be useful when switching between records should discard local draft state.

Resetting is a deliberate product behavior, not a rendering trick.


State is a snapshot

function handleClick() {
  setCount(count + 1);
  console.log(count); // current render's snapshot
}

The handler closes over the state value from the render that created it.

Calling a setter schedules a future render; it does not mutate the current snapshot in place.


Batching combines updates

setCount(count + 1);
setCount(count + 1);

Both updates may read the same snapshot and request the same next value.

Batching reduces unnecessary intermediate work, but it means update intent must be expressed correctly.


Use the previous state for dependent updates

setCount(previous => previous + 1);
setCount(previous => previous + 1);

Each updater receives the latest queued value.

Use this form whenever the next state depends on the previous state.


Derived values should usually be calculated

const visibleProducts = products
  .filter(product => product.title.includes(query))
  .filter(product => product.inStock);

If a value can be directly calculated from current inputs, storing it separately creates another synchronization obligation.


Source state versus derived state

source: products, query, filter
derived: visibleProducts, resultCount

Store the source values.

Calculate the derived values during rendering or in a memoized calculation when measurement justifies it.


Duplicated state creates inconsistency

const [products, setProducts] = useState<Product[]>([]);
const [visibleProducts, setVisibleProducts] = useState<Product[]>([]);

Now every product update and every filter update must keep both arrays synchronized.

One source of truth is usually simpler and safer.


Expensive derived values are a measurement question

const visibleProducts = expensiveFilter(products, query);

First ask:

  • is the calculation actually expensive?
  • how often does it run?
  • how large is the input?
  • which dependencies change?

Do not add memoization because a calculation exists.


Memoization preserves a calculation result

const visibleProducts = useMemo(
  () => expensiveFilter(products, query),
  [products, query],
);

The cache is valid only while its dependencies represent the same inputs.

Memoization is an optimization with cost, not a correctness requirement.


Referential identity affects memoization

const options = { sort: "price" };

This object is new on each render.

Passing it to a memoized child or using it as a dependency may invalidate the optimization even when its contents look unchanged.

Understand identity before optimizing around it.


Effects synchronize with something outside rendering

useEffect(() => {
  document.title = `${count} products`;
}, [count]);

The document title is external to React’s render calculation.

The effect synchronizes it after the committed UI reflects the new state.


Effects are not for ordinary derivation

Avoid:

useEffect(() => {
  setVisibleProducts(filter(products, query));
}, [products, query]);

This creates an extra state update and an intermediate render for a value that can be calculated directly.

Use an expression or memoized calculation instead.


Events and effects answer different questions

event: what should happen because the user did this?
effect: what external system must be synchronized with committed state?

Submit an order because the user activated submit.

Update a subscription because committed state says the subscription should exist.

Do not turn every event response into an effect chain.


Effect cleanup prevents stale work

useEffect(() => {
  const controller = new AbortController();
  loadProducts(query, controller.signal);

  return () => controller.abort();
}, [query]);

Cleanup runs when dependencies change or the component leaves the tree.

It prevents old subscriptions, timers, and requests from outliving the state that created them.


React rendering summary

flowchart TD
    A["State Update"] --> B["Snapshot-based Render Calculation"]
    B --> C["Reconciliation & Identity Matching"]
    C --> D["Commit Host DOM Changes"]
    D --> E["Effects Synchronize External Systems"]

This is a model for reasoning, not a promise that every implementation detail is synchronous or simple.


The Vue mental model

Vue tracks reactive dependencies more directly through refs, reactive proxies, computed values, and watchers.

The goal remains the same:

state → dependencies → render / synchronization

The tracking mechanism and timing model differ from React’s component recalculation model.


Vue ref() wraps a reactive value

const count = ref(0);

count.value += 1;

The ref object gives Vue a stable reactive container.

In templates, Vue can unwrap refs for convenient access.


Vue reactive() proxies an object

const filters = reactive({
  query: "",
  inStock: false,
});

filters.query = "phone";

The proxy intercepts reads and writes so Vue can track which reactive effects depend on which properties.


The proxy and original object differ

const original = { query: "" };
const state = reactive(original);

state !== original;

Use the reactive proxy consistently.

Identity assumptions become important when comparing, storing, or passing reactive objects.


Destructuring can break a reactive connection

const state = reactive({ query: "" });
const { query } = state;

The local query is no longer a reactive property reference in the same way.

Use toRefs, a computed value, or access through the reactive object when the connection must be preserved.


Vue rendering is a reactive effect

When a component renders, Vue records which reactive values it reads.

When one of those values changes, Vue knows that the component’s rendered output may need updating.

This is dependency tracking at the property level rather than a blanket statement that every component always reruns.


Vue DOM updates are scheduled

state.query = "phone";
state.inStock = true;

await nextTick();

The state assignments are observable to Vue immediately, but DOM work is typically queued and batched.

Do not assume the DOM reflects a mutation on the next line.


Vue batching reduces intermediate work

Several synchronous mutations can be grouped into one update cycle.

This is why code that needs the updated DOM may need nextTick() or an equivalent lifecycle boundary.

The scheduling model should be part of the component’s reasoning, especially around focus and measurement.


Computed values represent derivation

const visibleProducts = computed(() =>
  products.value
    .filter(product => product.title.includes(query.value))
    .filter(product => product.inStock),
);

Computed values declare a dependency relationship and can cache until their dependencies invalidate.


Computed values are not ordinary state

flowchart TD
    P["products (source)"] --> C["computed: visibleProducts"]
    Q["query (source)"] --> C
    F["filter (source)"] --> C

The computed result should not be manually synchronized with every source update.

Keep derivation represented as derivation.


Vue watchers synchronize external systems

watch(query, value => {
  router.replace({ query: { q: value } });
});

This is appropriate when a change in reactive state must update a router, storage layer, network operation, or other external system.


Watchers should not replace computed values

watch([products, query], () => {
  visibleProducts.value = filter(products.value, query.value);
});

This stores a derived value and introduces a synchronization path.

Prefer computed when the result is a direct calculation.


watch() and watchEffect() differ

flowchart TD
    W1["watch(source, callback)"] --- W1D["Explicit dependency source & controlled comparison"]
    W2["watchEffect(callback)"] --- W2D["Automatic dependency discovery during execution"]

Use the most explicit form that communicates the intended relationship.


Watcher cleanup prevents stale effects

watch(query, async (value, _oldValue, onCleanup) => {
  const controller = new AbortController();
  onCleanup(() => controller.abort());
  await loadProducts(value, controller.signal);
});

If a newer query arrives, the old operation should not win after it becomes irrelevant.


React and Vue share the goal, not the mechanism

ConcernReactVue
primary modelcomponent calculationtracked reactive dependencies
derivationexpression / memocomputed
external synceffectwatch / watchEffect
state updatesetter schedules rendermutation invalidates dependencies
DOM timingcommit and effectsqueued update and nextTick

Learn the mechanism well enough to predict behavior; do not flatten the differences into slogans.


Reactivity has granularity

flowchart TD
    Coarse["Coarse Reactivity\n(Rerun broad component calculation)"]
    Fine["Fine-Grained Reactivity\n(Invalidate only computations reading changed property)"]

Coarser systems can be simple to reason about.

Finer systems can reduce work by tracking smaller dependencies.

Neither choice removes the need for good state design.


Fine-grained reactivity

flowchart LR
    SigA["Signal A"] --> Deriv["Derived C"]
    SigB["Signal B"] --> Deriv
    Deriv --> Eff["Effect (DOM / Output)"]

Only computations that depend on invalidated sources need to be reconsidered.

This graph is explicit in the runtime rather than reconstructed from a broad component render.


Signals are a family of ideas

const count = signal(0);
const doubled = computed(() => count() * 2);
effect(() => console.log(doubled()));

Different libraries use different APIs and scheduling rules.

“Signals” names a reactive primitive pattern, not one universal technology.


Signals are not automatically faster

Performance depends on:

  • graph shape;
  • update frequency;
  • computation cost;
  • scheduling;
  • DOM work;
  • memory and bookkeeping;
  • developer usage.

A fine-grained mechanism can still perform unnecessary work if the graph or state model is poorly designed.


Scheduling is part of the model

flowchart TD
    Mut["State Mutation"] --> Inval["Dependency Invalidation"]
    Inval --> Queue["Job Scheduler Queue"]
    Queue --> Flush["Microtask Flush"]
    Flush --> Exec["Batch Render / Commit / Effects"]

Scheduling determines what “immediately” means.

It affects batching, race conditions, DOM measurement, focus, and perceived responsiveness.


Why scheduling exists

Scheduling can:

  • combine several changes;
  • avoid repeated layout work;
  • prioritize urgent interaction;
  • defer expensive computation;
  • coordinate asynchronous results;
  • prevent recursive update storms.

The trade-off is that a state mutation and a visible result may be separated in time.


Unnecessary rendering work is not always a bug

Ask:

  • did the calculation actually cost enough to matter?
  • did the DOM change?
  • did the user notice?
  • did the work block input or layout?
  • does optimization add more complexity than it removes?

Not every rerender is a problem.


Example: expensive filtering

const visible = useMemo(
  () => products.filter(matchesQuery),
  [products, query],
);

This may help when products is large and the calculation is repeated.

It may do nothing useful when the array is small, dependencies change every time, or rendering dominates the cost.


Memoization should follow measurement

flowchart LR
    Measure["Measure Performance"] --> Identify["Identify Repeated Work"]
    Identify --> Optimize["Narrowest Targeted Optimization"]
    Optimize --> Verify["Verify Behavior & Real Cost"]

Memoization has costs:

  • dependency maintenance;
  • memory;
  • identity management;
  • cognitive overhead.

Component memoization is not a correctness fix

const ProductCard = memo(function ProductCard(props: Props) {
  return <article>{props.product.title}</article>;
});

Memoization can skip a render when props are considered equal.

It cannot repair incorrect keys, duplicated state, stale closures, or a bad ownership boundary.


Vue’s selective tracking has its own costs

Property-level tracking can avoid broad updates.

But deep reactive objects, unstable identities, and unnecessary watchers can still make an application difficult to reason about.

Selectivity is a mechanism - not a substitute for a clear dependency graph.


Identity affects list rendering in every framework

flowchart TD
    Identity["Stable Item Key / Identity"]
    Identity --> State["Preserve correct row/sub-tree state"]
    Identity --> Focus["Preserve active focus & form inputs"]
    Identity --> Order["Deterministic list reordering & animations"]

Stable keys are a user-experience concern as much as a rendering optimization.


State preservation is architectural

When state disappears unexpectedly, ask:

  • did the component type change?
  • did its position change?
  • did its key change?
  • did a conditional branch replace its identity?
  • did the data identity change while the key stayed index-based?

The answer is often in the render tree, not in the state setter.


Conditional rendering can change identity

{mode === "edit" ? <Editor /> : <Preview />}

Switching between different component types replaces the identity at that position.

If two modes should preserve one shared draft, model them accordingly.

If switching should reset state, make that reset intentional.


Effects and watchers can create feedback loops

flowchart TD
    SChange["State Change"] --> Eff["Effect Runs & Calls setState"]
    Eff --> Loop["Triggers Re-render"]
    Loop --> SChange

Before adding synchronization, define:

  • the external source of truth;
  • the direction of the update;
  • the stopping condition;
  • cleanup and error behavior.

Think in reactive graphs

flowchart TD
    Query["query (source)"] --> Visible["computed: visibleProducts"]
    Products["products (source)"] --> Visible
    Stock["stockFilter (source)"] --> Visible
    Visible --> Grid["ProductGrid (View)"]
    Query --> URLEffect["URL Sync Effect"]

The graph reveals which values are sources, derivations, render consumers, and external effects.

It also reveals cycles and broad dependencies.


Compiler-assisted optimization

Compilers can sometimes infer:

  • stable expressions;
  • memoization opportunities;
  • dependency relationships;
  • component boundaries;
  • update paths.

They cannot infer product intent, correct identity, or whether a value belongs in state.

Optimization tools reduce some manual work; they do not remove architectural decisions.


React Compiler is an example, not a new mental model

Compiler assistance may reduce the need for some manual memoization.

The fundamentals remain:

  • pure render calculations;
  • stable identity;
  • correct dependencies;
  • explicit effects;
  • measured performance work.

Understand the behavior even when tooling automates an optimization.


Searchable product list: the shared dependency graph

flowchart TD
    Q["query (source)"] --> F["filteredProducts (derived)"]
    P["products (source)"] --> F
    Flt["filter (source)"] --> F
    F --> List["Product List View"]
    Q --> URL["URL Sync Effect"]
    List --> Analytics["Selection Analytics Effect"]

The same product feature can be implemented in React, Vue, or a signal system while preserving this conceptual graph.


React version: calculate and synchronize

const filteredProducts = useMemo(
  () => filterProducts(products, query, filter),
  [products, query, filter],
);

useEffect(() => {
  router.replace({ query });
}, [query]);

Derivation stays in the render model.

The router is synchronized in an effect.


React anti-pattern: derived state effect

useEffect(() => {
  setFilteredProducts(filterProducts(products, query, filter));
}, [products, query, filter]);

This creates:

  • duplicated state;
  • an extra render path;
  • a possible stale intermediate value;
  • more code to test and debug.

Calculate the list directly unless there is a measured reason not to.


Vue version: computed and watch

const filteredProducts = computed(() =>
  filterProducts(products.value, query.value, filter.value),
);

watch(query, value => router.replace({ query: { q: value } }));

The same graph is expressed with Vue’s dependency tracking primitives.


Trace React updates

For a state change, record:

  1. Which setter was called?
  2. Which snapshot created the handler?
  3. Which components calculate again?
  4. Which identities and keys are reused?
  5. Which host nodes commit changes?
  6. Which effects run and which clean up?

This sequence is more useful than saying “React rerendered everything.”


Trace Vue updates

For a reactive mutation, record:

  1. Which ref or proxy property changed?
  2. Which computed values depend on it?
  3. Which component render effects read it?
  4. Which watchers are triggered?
  5. When is the DOM queue flushed?
  6. Which cleanup functions run?

The goal is to expose the dependency graph and schedule.


What should we measure?

Measure questions such as:

  • how long does the expensive calculation take?
  • how many items are processed?
  • how often does the calculation run?
  • how many components commit changes?
  • does input responsiveness degrade?
  • does memory grow because of caches or subscriptions?

Choose a measurement that can change the decision.


Necessary UI change versus unnecessary work

flowchart TD
    StateChange["State Changed"] --> Recalc["Recalculate Derived Descriptions"]
    Recalc --> DiffCheck{"Are Descriptions Different?"}
    DiffCheck -- Yes --> DOMMut["Necessary DOM Changes Committed"]
    DiffCheck -- No --> Skip["Skip Host DOM Updates"]

Do not optimize away work before identifying which part is unnecessary.

Correctness and a truthful dependency graph come first.


State location affects rendering scope

State placed high in the tree can coordinate many consumers but broaden update scope.

State placed close to one interaction can reduce unrelated work but may require a deliberate communication path.

Choose location based on ownership and synchronization - not on a universal rule to lift or localize state.


Derived state affects rendering scope

Keep derivation near the data and consumers that need it.

If a derived value is shared, expose one clear calculation rather than duplicating it in multiple components.

If it is cheap and local, a plain expression is often the best design.


Side effects should cross a boundary

flowchart LR
    subgraph Pure["Pure Computation Zone"]
        Render["Render: State → Desired UI Description"]
    end
    subgraph Boundary["Side Effect Boundary"]
        Sync["Network, Storage, DOM, Timers, Subscriptions"]
    end
    Pure -->|Committed State| Boundary

The boundary helps answer:

  • when does this run?
  • what invalidates it?
  • how is it cleaned up?
  • what happens if the value changes again?

Do not let the reactive system become a mystery network

A healthy graph has:

  • visible sources;
  • named derivations;
  • few synchronization edges;
  • bounded effects;
  • explicit cleanup;
  • tests for identity and timing.

If changing one field triggers an unexplained chain of watchers and effects, simplify the graph before optimizing it.


A practical decision model

Ask:

  • Is another value directly calculable from this one?
  • Does a user action cause an operation?
  • Must an external system remain synchronized?
  • Is a calculation expensive and repeated unnecessarily?
  • Is the update scope too broad?

These questions map naturally to derived values, events, effects, memoization, and state placement.


Four reactive strategies

StrategyMain strengthMain risk
React render modelexplicit component calculationconfusing render with DOM work
Vue dependency trackingselective property updateshidden reactive connections
signalsfine-grained graphgraph and lifecycle complexity
compiler assistanceless manual optimizationfalse confidence about architecture

Use the model your team can explain and debug.


Practical lab: Reactive Computed Graph

Implement a tiny educational reactive graph to make source state, derived values, dependency tracking, and effects visible.

The implementation is for learning. It is not a production reactive runtime.


Practical stages 1–3: sources and derivation

  1. Implement a signal with subscribers.
  2. Add lazy computed values and invalidation.
  3. Add an effect with cleanup.

Make the graph observable with logs or counters so the execution order can be inspected.


Practical stages 4–5: cycles and comparisons

  1. Create a cycle deliberately and explain why production systems must guard against it.
  2. Compare the educational graph with React render calculation and Vue computed().

Ask which mechanism discovers dependencies, when invalidation occurs, and how cleanup is represented.


Practical extension: visualize the graph

Add a graph visualizer showing:

flowchart LR
    Source["Signal Source"] --> Comp1["Computed A"]
    Comp1 --> Comp2["Computed B"]
    Comp2 --> EffectNode["Effect Observer"]

Highlight:

  • invalidated nodes;
  • execution order;
  • cached nodes;
  • unused computations;
  • cycles.

The visualizer should make the invisible dependency model inspectable.


Try this yourself

Build a searchable list with:

  • source products;
  • query state;
  • a derived filtered list;
  • a URL synchronization effect;
  • cancellation for stale searches.

Then explain which updates are necessary and which calculations can remain untouched.


Troubleshooting guide (Part 1)

SymptomLikely cause
State update seems one step behindSnapshot or batching misunderstood
Input state moves to another rowUnstable or index-based keys
UI flashes an old derived valueDerived state synchronized through an effect
Effect runs foreverEffect updates one of its own dependencies

Troubleshooting guide (Part 2)

SymptomLikely cause
Vue value stopped updatingDestructuring removed the reactive connection
DOM is old after a mutationUpdate is queued; await the framework flush
Memoization changes nothingDependencies or calculation cost do not justify it
Async result wins after a newer queryMissing cleanup or cancellation

Completion checklist

  • source state is distinct from derived values;
  • render calculations are free of external side effects;
  • list keys represent stable identity;
  • previous-state updates are used for dependent changes;
  • effects and watchers synchronize external systems only;
  • cleanup prevents stale subscriptions and requests;
  • React and Vue timing differences are understood;
  • optimization decisions are supported by measurement;
  • the reactive graph is explainable from source to screen.

Misconceptions to leave behind (Part 1)

MisconceptionBetter mental model
State changed, so the DOM changed immediatelyUpdates are calculated and scheduled
A React render rebuilt the DOM subtreeRender and commit are different phases
Virtual DOM is always fasterPerformance depends on the measured work
Every rerender is a bugSome recalculation is necessary
Keys only remove warningsKeys preserve logical identity

Misconceptions to leave behind (Part 2)

MisconceptionBetter mental model
Derived values belong in stateDirect calculations usually stay derived
Effects respond to any state changeEffects synchronize external systems
Vue watchers calculate normal derivationscomputed represents derivation
Signals are one standard technologySignals are a family of mechanisms
Compiler optimization makes architecture irrelevantTools cannot choose ownership or intent

The chapter in one sentence

Design a truthful reactive graph: keep source state minimal, derivation explicit, identity stable, scheduling understood, and side effects at the boundary.


Next: Chapter 8

The next chapter will apply these rendering and state principles to:

  • asynchronous data and server state;
  • loading, error, empty, and success states;
  • request cancellation and stale results;
  • caching and synchronization;
  • resilient data-fetching architecture.

Questions

For one interaction in your application, can you draw the path from source state to derived value to committed UI?

Where does the first external side effect enter the graph?

8 - State Management, Routing & Form Architecture

Chapter 8: classify state, assign ownership, design URL-driven views, and build resilient form workflows.

State Management, Routing & Form Architecture

Put every value where it belongs

Chapter 8

Polla Fattah


Today’s goal

Stop treating all state as the same kind of problem.

We will connect:

  • local, shared, domain, server, URL, form, persistent, and derived state;
  • ownership, reducers, stores, and state machines;
  • routing as application state architecture;
  • paths, parameters, history, loading, and navigation UX;
  • controlled, uncontrolled, and hybrid forms;
  • validation, dirty state, dynamic fields, and multistep workflows;
  • state placement in a routed administrative catalogue.

By the end of today you can

  • classify a value before choosing a state tool;
  • keep state close to its practical owner;
  • explain why server data is not ordinary global state;
  • model URL state as a serializable public view;
  • design history and back-button behavior intentionally;
  • separate form drafts from domain models;
  • distinguish touched, dirty, validation, and submission state;
  • use reducers and state machines for explicit transitions;
  • choose local state, context, or a store proportionately;
  • review a feature for duplicated or misplaced state.

The central principle

State architecture is the deliberate placement of values according to ownership, lifetime, sharing, persistence, and transition rules.

The right question is not “which state library should we use?”

The better question is “what kind of state is this, and who is responsible for it?”


The chapter’s progression

flowchart TD
    A[Classify State] --> B[Assign Ownership]
    B --> C[Choose Update Model]
    C --> D[Design URL & Route State]
    D --> E[Design Form State]
    E --> F[Model Workflow Transitions]
    F --> G[Test Navigation & Persistence]

Tools come after the state model, not before it.


“State” is not one thing

flowchart TD
    subgraph StateTaxonomy["Application State Taxonomy"]
        S1["dialogOpen: Local UI State"]
        S2["selectedTab: Local / URL State"]
        S3["currentUser: Domain / Session State"]
        S4["products: Server State"]
        S5["query: URL / Local Draft State"]
        S6["formDraft: Form State"]
        S7["themePreference: Persistent Client State"]
        S8["resultCount: Derived State"]
    end

The category predicts who owns the value and how it should change.


A practical state taxonomy (Part 1)

CategoryTypical lifetimeExample
local UIone component or featuremodal open
shared UIseveral nearby componentsselected tab
domainbusiness workflowapproval status
serverremote sourceproduct list

A practical state taxonomy (Part 2)

CategoryTypical lifetimeExample
URLshareable viewpage and filters
formunfinished inputdraft email
persistentacross sessionstheme preference
derivedcalculatedfiltered results

Local UI state

const [isOpen, setIsOpen] = useState(false);
const [focusedIndex, setFocusedIndex] = useState(0);

Local state is usually:

  • interaction-specific;
  • short-lived;
  • not useful to unrelated routes;
  • safe to discard when the component leaves the tree.

Keep it local unless another owner genuinely needs it.


Shared UI state

Tabs.List  ↔ selectedTab ↔ Tabs.Panel

Shared UI state coordinates a small group of related components.

Use a parent, compound component context, or a local feature store when the relationship is real.

Do not promote it globally merely because more than one component reads it.


Domain state

type ApprovalState =
  | { status: "draft" }
  | { status: "submitted"; submittedAt: string }
  | { status: "approved"; approvedBy: string }
  | { status: "rejected"; reason: string };

Domain state represents business meaning and rules.

It should not be confused with whether a modal is open or a request is currently loading.


Server state is a different category

Server state has:

  • a remote owner;
  • latency and failure;
  • freshness and staleness;
  • caching concerns;
  • invalidation rules;
  • multiple possible consumers.

Treating it as one ordinary global variable usually loses important behavior.


Server state is not “just another global variable”

request → pending → success / error
             ↓
       cache and freshness
             ↓
     refetch / invalidate / retry

The state includes the lifecycle of synchronization, not only the latest payload.


Cached data needs an ownership policy

Ask:

  • who populated the cache?
  • when is it stale?
  • who invalidates it after a mutation?
  • can two requests race?
  • can an older response overwrite a newer one?
  • what happens offline?

Caching is architecture, not merely a performance toggle.


URL state is public view state

/catalogue?query=phone&sort=price&page=2

URL state can survive:

  • reloads;
  • sharing;
  • bookmarks;
  • back and forward navigation;
  • opening the view in another tab.

That makes it valuable - but also part of the public contract.


Form state is unfinished work

draft value
touched fields
dirty fields
validation messages
submission status
server errors

A draft is not automatically a valid domain object.

Form architecture should represent the journey from incomplete input to accepted data.


Persistent client state has a migration problem

localStorage → parse → version check → migrate → trusted preferences

Persisted values can outlive the code that created them.

Treat storage as an external boundary and design for old or malformed versions.


Derived state should usually be calculated

const visibleProducts = products
  .filter(matchesQuery)
  .sort(compareProducts);

If visibleProducts is directly calculable from source state, storing it separately creates another value that can become stale.

Use memoization only when repeated computation is measurably expensive.


Ownership: who is responsible?

For every value, ask:

  • who decides it?
  • who needs to read it?
  • who changes it?
  • how long should it survive?
  • should it be shareable or persisted?
  • what external system owns the truth?

The answers identify the correct boundary better than a favorite library does.


Keep state as close as practical

smallest common owner
        ↓
components that genuinely need the value

Local ownership reduces coordination and update scope.

Move state outward only when a real consumer, lifecycle, or synchronization rule requires it.


State should move outward for a reason

Good reasons include:

  • two siblings must coordinate;
  • a route must own the view;
  • a server cache is shared;
  • a workflow spans multiple screens;
  • another system must synchronize with the value.

“It might be useful later” is not an ownership rule.


Global state has a cost

Global state can create:

  • invisible dependencies;
  • broad update scope;
  • unclear ownership;
  • difficult isolated tests;
  • stale values that survive too long;
  • accidental coupling between features.

Make a value global because its lifetime and sharing demand it, not because global access is convenient.


Reducers make transitions explicit

type Action =
  | { type: "searchChanged"; query: string }
  | { type: "submitted" }
  | { type: "reset" };

function reducer(state: State, action: Action): State {
  switch (action.type) {
    case "searchChanged": return { ...state, query: action.query };
    case "submitted": return { ...state, status: "submitted" };
    case "reset": return initialState;
  }
}

The transition vocabulary makes state changes inspectable and testable.


Unidirectional data flow

flowchart LR
    State["Current State"] --> Render["Render UI"]
    Render --> Intent["User Action / Intent"]
    Intent --> Action["Dispatch Action"]
    Action --> Reducer["Transition Function"]
    Reducer --> Next["New State Snapshot"]
    Next -.-> State

One direction makes it easier to answer:

  • what caused this value?
  • which action changed it?
  • which consumers should update?

It does not mean every value belongs in one global store.


Stores are boundaries, not magic containers

A store should define:

  • the state it owns;
  • actions or methods that change it;
  • selectors or derived values;
  • effects and external dependencies;
  • initialization and cleanup.

If a store becomes the home for every value, it has stopped communicating ownership.


Reducers versus stores

ReducerStore
transition functionlonger-lived owner
often puremay coordinate effects
easy to test as input/outputmay expose selectors and subscriptions
useful inside local featuresuseful for shared domain workflows

They can be combined. Neither is automatically the correct scale.


State machines make workflows visible

flowchart LR
    Idle["Idle"] --> Editing["Editing"]
    Editing --> Submitting["Submitting"]
    Submitting --> Success["Success"]
    Submitting --> ErrorState["Error"]
    ErrorState -->|Retry| Submitting
    ErrorState -->|Edit again| Editing

State machines are useful when legal transitions matter more than storing a collection of booleans.


Replace impossible boolean combinations

Avoid:

isLoading = true;
isSuccess = true;
isError = true;

Prefer one explicit status:

type Status = "idle" | "loading" | "success" | "error";

The model should make impossible states difficult to represent.


Routing is state architecture

A route decides more than the address.

It can determine:

  • which feature is active;
  • which data loads;
  • which layout persists;
  • what can be shared;
  • what the back button restores;
  • which code is loaded.

Routing is a state-lifetime and ownership decision.


Paths represent resource identity

/products
/products/p-42
/orders/o-17

Use a path segment when the value identifies the resource or nested location being viewed.

The route should speak domain language, not reveal component filenames.


Route parameters identify a resource

const { productId } = route.params;

Validate the parameter before using it as a domain identifier.

The string from the URL is external input, not automatically a valid ProductId.


Query parameters describe a view

/products?query=phone&sort=price&page=2

Query parameters are suitable for filters, sorting, pagination, search, and other shareable view choices.

They should be serializable, parseable, and stable enough to form a public contract.


Path parameter or query parameter?

Path:  /products/p-42
       “which resource?”

Query: /products?sort=price
       “how should this collection be viewed?”

Use the semantic distinction, not personal preference.


Nested routes express nested ownership

/admin
  /catalogue
    /products
    /products/:id/edit

Parent routes can own layout, permissions, data context, and persistent navigation while child routes own the active view.


Layout routes preserve context

flowchart TD
    Layout["AdminLayout (Persistent Shell)"]
    Layout --> Sidebar["Sidebar Nav"]
    Layout --> Header["Top Header & User Info"]
    Layout --> Outlet["Child Route &lt;Outlet /&gt;"]

The layout can remain mounted while the child changes.

This preserves navigation context and avoids rebuilding shared structure unnecessarily.


Navigation changes state and history

flowchart TD
    Push["push: New history entry (back returns here)"]
    Replace["replace: Update current entry (no new history step)"]
    Back["back / forward: Restore historical view snapshot"]

Choose history semantics deliberately.

Typing a filter may replace the current entry.

A meaningful page transition may push a new entry.


Redirects are state transitions

Redirect when:

  • a user lacks access;
  • a resource moved;
  • a form completed a workflow;
  • a canonical URL should replace an invalid one.

Preserve useful context where appropriate, and avoid redirect loops that obscure the actual state.


File-based routing is a convention

File structure can make routes discoverable:

routes/
  products/index.tsx
  products/$productId.tsx
  products/$productId/edit.tsx

It helps organize the map, but it does not decide which state belongs in the route or how transitions should behave.


The URL is state

flowchart LR
    URL["URL Address Bar"] <-->|"Parse / Serialize"| State["Parsed Route State"]
    State <-->|"Render / Events"| View["Rendered View"]

Do not copy URL values into local state without a clear ownership reason.

Two sources of truth create synchronization work and surprising back-button behavior.


Shareable state belongs naturally in the URL

Good candidates:

  • search query;
  • filters;
  • sorting;
  • pagination;
  • selected public tab;
  • view density when sharing it matters.

If another person opens the URL, the meaningful view should be reconstructable.


What usually should not go in the URL?

Avoid placing:

  • passwords or secrets;
  • sensitive personal data;
  • large drafts;
  • ephemeral hover state;
  • internal implementation details;
  • values that cannot be serialized safely.

The URL is visible, copyable, logged, and often shared.


URL state should have one owner

Avoid:

URL query ↔ local query ↔ debounced query ↔ server query

Define the stages explicitly:

flowchart LR
    Draft["1. Local Draft Input\n(Immediate keystrokes)"] -->|"Commit (Enter/Blur/Debounce)"| URL["2. Committed URL Query\n(Shareable & Back-button aware)"]
    URL -->|"Data Fetch / Cache"| ServerReq["3. Server API Request\n(Cancellable async query)"]

Each stage has a different purpose and transition.


Parse URL state at the boundary

type CatalogueQuery = {
  query: string;
  page: number;
  sort: "relevance" | "price";
};

Parsing should:

  • provide defaults;
  • reject or normalize invalid values;
  • clamp unsafe ranges;
  • preserve only supported vocabulary;
  • return a trusted view model.

Example: URL-based filtering

const params = new URLSearchParams(location.search);

const state = {
  query: params.get("query") ?? "",
  sort: parseSort(params.get("sort")),
  page: parsePage(params.get("page")),
};

The view consumes state, not raw strings from location.search.


Routing UX is part of architecture

A route transition should define:

  • pending feedback;
  • loading behavior;
  • error boundaries;
  • scroll restoration;
  • focus placement;
  • unsaved-change handling;
  • code-loading behavior.

The route is a user interaction, not merely a URL replacement.


Pending navigation needs visible feedback

flowchart TD
    Req["Navigation Requested"] --> Indicator["Pending Feedback Indicator"]
    Indicator --> Resolve["Data & Code Bundle Resolve"]
    Resolve --> CommitRoute["New Route Commits to DOM"]

Keep the current context understandable while the next view is loading.

Avoid showing a blank screen for an operation that can preserve useful layout.


Loading states should belong to the right scope

flowchart TD
    S1["Route shell loading: Page-level fallback / spinner"]
    S2["Panel data loading: Panel skeleton loader"]
    S3["Button mutation: Local inline pending state"]

One global spinner often hides which part of the interface is actually unavailable.


Route errors are part of the route contract

Model distinct causes:

  • invalid route parameter;
  • missing resource;
  • permission failure;
  • network failure;
  • unexpected application error.

The user should receive the recovery action appropriate to the cause.


Scroll restoration is state restoration

Decide whether navigation should:

  • restore the previous scroll position;
  • start a new route at the top;
  • preserve a nested panel’s scroll;
  • maintain position while query filters update.

Surprising scroll behavior makes a correct route feel broken.


Focus after navigation is accessibility state

After a route changes, move focus to a meaningful landmark or heading when appropriate.

Do not leave keyboard users at an old location with no indication that the view changed.

Focus management belongs to the route transition boundary.


State preservation across routes is a choice

Ask:

  • should the parent layout remain mounted?
  • should the child form survive route changes?
  • should a tab selection be encoded in the URL?
  • should changing an ID reset local state?

Route nesting, component identity, and state ownership work together.


Route-level code splitting

flowchart TD
    Shell["Initial Application Shell"] --> CatRoute["Load Catalogue Route Chunk"]
    CatRoute --> EditRoute["Load Edit Route Chunk on Demand"]

Splitting at route boundaries can reduce initial work and align code loading with user navigation.

The route model and build model should reinforce one another.


Form architecture is state architecture

A form includes more than input values:

values + touched + dirty + validation + submission + server errors

Design these lifecycles explicitly instead of allowing them to emerge from scattered event handlers.


Controlled forms

<input
  value={email}
  onChange={event => setEmail(event.target.value)}
/>

Benefits:

  • immediate application visibility;
  • easy derived feedback;
  • explicit formatting and validation;
  • synchronization with other state.

Costs include rerenders, wiring, and more code for large forms.


Uncontrolled forms

<form onSubmit={handleSubmit}>
  <input name="email" defaultValue="" />
</form>

The browser owns current input values until submission or a deliberate read.

Native form behavior can be simple, efficient, and accessible when the application does not need every keystroke.


Hybrid form architectures

flowchart LR
    Native["Native Input Editing"] --> Draft["Local Form Draft State"]
    Draft --> Validation["Controlled Validation Engine"]
    Validation --> Command["Submitted Domain Command"]

Use control where coordination is needed and native behavior where it is enough.

Do not choose one mode for ideological reasons.


Form values are not automatically domain values

input value: "42"
domain value: 42

input value: ""
domain value: undefined or validation failure

Forms represent unfinished, string-heavy input.

Parse and validate before constructing a domain command.


Keep the form model separate when useful

type ProductDraft = {
  title: string;
  price: string;
  categoryId: string;
};

type ProductCommand = {
  title: string;
  priceCents: number;
  categoryId: CategoryId;
};

The draft supports editing.

The command represents validated intent.


Touched state records interaction

touched: false → user focused and left field → true

Use it to decide when a field-level message should appear.

Touched does not mean the value changed and does not mean the value is valid.


Dirty state records change from an initial value

dirty = currentDraft !== initialDraft

Dirty state answers whether there may be unsaved work.

It is different from touched state: a user can touch a field and return it to its original value.


Validation has several layers

flowchart TD
    F1["1. Field-Level: Format, length, regex (e.g. Email syntax)"]
    F2["2. Cross-Field: Relationship rules (e.g. End date after Start date)"]
    F3["3. Domain/Business: Entity rules (e.g. Permit status editable)"]
    F4["4. Server-Side: Authority validation (e.g. Unique registration number)"]
    F1 --> F2 --> F3 --> F4

Keep the layer visible so the UI can show the right message and recovery path.


Derive validation where practical

const emailError = email.length === 0
  ? "Email is required"
  : isEmail(email) ? undefined : "Email is invalid";

Do not store every validation message if it can be derived from the current draft and interaction state.

Store server results and asynchronous validation when they are not directly calculable.


Cross-field validation needs the whole draft

if (draft.password !== draft.confirmPassword) {
  return { confirmPassword: "Passwords do not match" };
}

The validation owner should receive the relevant model rather than forcing fields to synchronize through unrelated global state.


Dynamic fields need stable identity

type LineItemDraft = {
  key: string;
  productId: string;
  quantity: string;
};

Use a stable draft key for rows that can be inserted, removed, or reordered.

Index identity can move input values to a different line item.


Dynamic field operations are domain actions

add line item
remove line item
move line item
change quantity

Represent these operations explicitly in a reducer or form API.

Scattered array mutations make dirty state, validation, and focus behavior harder to preserve.


Multistep forms are workflows

flowchart LR
    Account["Account"] --> Profile["Profile"]
    Profile --> Review["Review"]
    Review --> Submit["Submit"]

Each step needs:

  • an entry condition;
  • validation scope;
  • persistence policy;
  • back behavior;
  • recovery from invalid or stale data.

Treat the wizard as a state machine when transitions are meaningful.


URL and multistep forms can work together

/onboarding/profile
/onboarding/review

Put the current step in the URL when it should survive reload, sharing, or navigation.

Keep sensitive drafts and unfinished values in an appropriate private owner.


A wizard state machine

flowchart LR
    Incomplete["profileIncomplete"] --> Complete["profileComplete"]
    Complete --> InReview["reviewStep"]
    InReview --> Submitting["submitting"]
    Submitting --> Done["success"]
    Submitting --> Fail["serverError"]
    Fail -->|Retry| Submitting

Explicit transitions prevent the UI from entering a review step without the required data.


When native HTML is enough

Prefer native forms when:

  • fields are simple;
  • browser validation is adequate;
  • every keystroke need not update application state;
  • submission can read FormData;
  • accessibility should follow platform behavior.

Native behavior is a feature, not an implementation failure.


When a form library is justified

A library can help with:

  • many fields and nested structures;
  • reusable validation patterns;
  • touched and dirty tracking;
  • field arrays;
  • controlled/uncontrolled integration;
  • performance and subscription granularity.

Adopt it for repeated complexity, not for a two-field form.


Administrative catalogue: classify before building

URL: query, filters, sort, page
server: products, categories
local UI: modal, focused row
form: draft product, touched, dirty, errors
derived: visible rows, result count
persistent: table density preference

The map explains where each value should live.


Avoid duplicating URL state

Bad:

flowchart LR
    subgraph Bad["Anti-Pattern: Duplicating URL State Across Stores"]
        B1["URL Page"] --> B2["Local Page State"] --> B3["Global Store Page"] --> B4["Request Page"]
    end
    subgraph Good["Architectural Best Practice: Direct Pipeline"]
        G1["URL Page"] --> G2["Parsed Route State"] --> G3["Request Query"]
    end

Create a local draft only when editing and committing are intentionally different states.


Avoid copying server data into the form too early

server product → form draft when edit begins

This is a deliberate transition, not a continuous mirror.

After editing starts, the draft can diverge safely until save or reset.

Keep server freshness and local unsaved work conceptually separate.


Modal state should stay local unless navigation owns it

click “delete” → local confirmation modal

Put the modal in the URL only when the modal itself should be deep-linkable, restorable, or part of browser history.

Not every visible state deserves a route.


Selected tabs: local or URL?

Use local state when:

  • the tab is ephemeral;
  • sharing the selection is not useful;
  • back navigation should not record every change.

Use URL state when:

  • the selected view is meaningful to share;
  • reload should preserve it;
  • browser history should restore it.

These often belong in the URL because they define the collection view.

But distinguish:

local suggestion input → committed search parameter

The user can type freely without creating a history entry for every keystroke, then commit the meaningful query deliberately.


Persistent preferences are not route state

table density → local storage preference
current page  → URL state

A preference follows the user across views.

A route value describes one navigable view.

Their lifetimes and ownership differ.


One value can change category over time

search draft       local form state
committed search   URL state
server results     server state
cached results     server cache

The value’s role changes through explicit transitions.

Do not force every stage into one state container.


A state placement decision tree

flowchart TD
    Q1{"Is it calculable from existing data?"} -- Yes --> A1["Derive it (Pure computed / inline)"]
    Q1 -- No --> Q2{"Is an external system the owner?"}
    Q2 -- Yes --> A2["Synchronize at boundary (API/Storage)"]
    Q2 -- No --> Q3{"Must it be bookmarkable / shareable?"}
    Q3 -- Yes --> A3["URL Query Parameter"]
    Q3 -- No --> Q4{"Must it survive reloads / sessions?"}
    Q4 -- Yes --> A4["Persistent Client Storage (localStorage/IndexedDB)"]
    Q4 -- No --> Q5{"Who needs it?"}
    Q5 -- Single component --> A5["Local Component State"]
    Q5 -- Component subtree --> A6["Context / Provide-Inject"]
    Q5 -- Cross-feature domain workflow --> A7["Dedicated Store / State Machine"]

This is a reasoning aid, not a mechanical law.


Common state smells

Watch for:

  • state copied from props without a reset rule;
  • derived values stored as state;
  • everything placed in a global store;
  • one value duplicated in URL and component state;
  • effects used to synchronize local copies;
  • server data copied into forms continuously.

Each smell suggests competing owners.


Routing smells

Warning signs:

  • route reflects implementation rather than domain;
  • important view state disappears on refresh;
  • back button behaves surprisingly;
  • child navigation destroys useful parent context;
  • a rarely used route inflates the initial bundle;
  • invalid parameters reach data-fetching code unparsed.

Form smells

Warning signs:

  • every keystroke updates a global store;
  • validation is duplicated across fields and submit handlers;
  • draft and domain models are forced to be identical;
  • dirty and touched flags are synchronized manually everywhere;
  • a multistep form uses unrelated booleans instead of transitions.

React URL-state catalogue

const query = useSearchParams();
const state = parseCatalogueQuery(query);

return (
  <CatalogueFilters
    value={state}
    onChange={next => navigate({ search: serialize(next) })}
  />
);

The URL owns committed view state.

The component receives a parsed model rather than raw strings.


Vue URL-state catalogue

const route = useRoute();
const router = useRouter();
const state = computed(() => parseCatalogueQuery(route.query));

function update(next: CatalogueQuery) {
  router.replace({ query: serialize(next) });
}

Again, compare ownership and transitions rather than framework syntax.


Forms in React and Vue

flowchart LR
    subgraph ReactWorld["React Paradigm"]
        R1["Controlled Input (value + onChange)"]
        R2["useReducer (state + dispatch)"]
        R3["useEffect (synchronization)"]
    end
    subgraph VueWorld["Vue Paradigm"]
        V1["v-model (two-way binding syntax)"]
        V2["reactive + action methods"]
        V3["watch / watchEffect"]
    end
    R1 <-->|Equivalent| V1
    R2 <-->|Equivalent| V2
    R3 <-->|Equivalent| V3

The important design questions remain:

  • what is the draft?
  • what is valid?
  • what is submitted?
  • who owns the transition?

A complex edit form model

type EditState = {
  draft: ProductDraft;
  initial: ProductDraft;
  touched: Set<string>;
  errors: Record<string, string>;
  status: "idle" | "saving" | "saved" | "error";
};

The model separates current values, comparison baseline, interaction history, validation, and submission lifecycle.


A form reducer makes operations visible

type FormAction =
  | { type: "fieldChanged"; name: string; value: string }
  | { type: "fieldBlurred"; name: string }
  | { type: "submitted" }
  | { type: "reset" };

Explicit actions make it possible to test dirty state, touched state, validation, and reset behavior as transitions.


Route plus form interaction

flowchart TD
    Route["1. Route: /products/p-42/edit"] --> Load["2. Load Server Product into Cache"]
    Load --> Init["3. Initialize Local Form Draft (Detached copy)"]
    Init --> Edit["4. User Edits Draft (Zero mutation of server cache)"]
    Edit --> Validate["5. Validate & Submit Command"]
    Validate --> Inval["6. Invalidate / Update Server State"]
    Inval --> Nav["7. Navigate to Canonical Resource View"]

Each arrow is a deliberate ownership transition.


Unsaved changes need a policy

When a dirty form meets navigation, decide:

  • block and confirm;
  • autosave;
  • preserve a draft;
  • discard explicitly;
  • allow navigation and make loss clear.

Do not let a route unmount silently destroy work the user believes is still present.


Persistence for long forms

Persist only what is appropriate:

  • version the draft;
  • exclude secrets;
  • expire stale drafts;
  • validate on restore;
  • show the user what was restored;
  • provide clear reset behavior.

Persistence is another external boundary with lifecycle and privacy decisions.


Local state versus context versus store

flowchart TD
    Local["Local State: One component / feature owns it"]
    Context["Context State: Related family shares it without prop drilling"]
    Store["Store State: Cross-feature domain workflow / server data"]
    Local --> Context --> Store

Start with the smallest scope that satisfies the real consumers.


Server state versus client state

flowchart TD
    S1["Server Data: Remote owner, asynchronous, stale, refetchable"]
    S2["Client UI State: Local user choices, immediate interaction, ephemeral"]
    S3["Domain State: Business transitions, validated models, domain rules"]

The same object may be represented in more than one layer, but each representation needs a clear owner and synchronization rule.


URL state versus persistent state

flowchart LR
    URL["URL State: Current shareable view & navigation bookmark"]
    Storage["Persistent Storage: Long-lived preference or offline draft across sessions"]

The URL is visible and navigable.

Persistence survives beyond one route and may need migration.

Do not use one as a substitute for the other.


State ownership and testing

Clear ownership makes focused tests possible:

  • parser tests for URL state;
  • reducer tests for transitions;
  • form tests for validation and dirty behavior;
  • route tests for history and redirects;
  • server-state tests for loading and stale responses.

If every test needs the whole application, ownership may be too global.


State ownership and team scale

As a team grows, implicit ownership becomes expensive.

Document:

  • which module owns a value;
  • which API changes it;
  • what is public URL state;
  • what is cached server data;
  • what can be persisted;
  • which transitions are legal.

Architecture reduces coordination cost when boundaries are explicit.


Practical lab: Routed Administrative Catalogue

Build a catalogue whose filters, sorting, pagination, and selected view are represented in the URL when they should survive reload, sharing, and history navigation.

Then add a routed edit form with explicit ownership for draft, validation, server state, and workflow transitions.


Practical stages 1–4: inventory and URL state

  1. Create the state inventory.
  2. Build URL-based filters.
  3. Parse and validate query parameters.
  4. Keep server state separate.

Verification: a copied URL reconstructs the same meaningful view without exposing sensitive data.


Practical stages 5–8: local UI and editing

  1. Add local modal state.
  2. Convert editing to a route.
  3. Create a complex edit form.
  4. Track touched and dirty state.

Do not copy server data into a continuously synchronized global form object.


Practical stages 9–12: validation and transitions

  1. Add cross-field validation.
  2. Add dynamic fields with stable keys.
  3. Add a reducer.
  4. Model the workflow as a state machine.

Make invalid combinations and illegal transitions visible in the model.


Practical stages 13–16: resilient navigation

  1. Add navigation UX.
  2. Add route-level code splitting.
  3. Add persistence with versioning and validation.
  4. Create the final state map.

Include loading, error, retry, unsaved-change, and back/forward behavior.


Add:

flowchart LR
    Input["Local Keystrokes"] --> Debounce["Debounce Timer"]
    Debounce --> Cancel["AbortController (Cancel in-flight)"]
    Cancel --> Cache["Server-State Cache (Query Result)"]

Keep the local suggestion draft separate from the committed URL query.

Ensure an old response cannot overwrite a newer query’s result.


Try this yourself

For a product editor, decide where these values belong:

modal open
product ID
draft title
server product
page number
selected tab
dirty flag
filtered rows

Write the owner and lifetime beside each one before writing components.


Troubleshooting guide (Part 1)

SymptomLikely cause
URL and input disagreeTwo owners exist for the same value
Back button feels noisyEvery transient edit pushed history
Refresh loses the viewShareable state stayed local
Old API data overwrites new dataServer-state race lacks cancellation or identity

Troubleshooting guide (Part 2)

SymptomLikely cause
Form loses work on navigationDirty policy is undefined
Validation flickersDraft, touched, and errors are conflated
Store contains everythingState categories were never classified
Child route feels like a full resetParent context or layout is not preserved

Completion checklist

  • every important value has a category and owner;
  • derived values are not duplicated unnecessarily;
  • server state has freshness and failure semantics;
  • URL state is serializable, validated, and shareable;
  • history semantics are intentional;
  • form drafts are separate from accepted domain commands;
  • dirty, touched, and validation state have distinct meanings;
  • reducers or state machines model meaningful transitions;
  • persistence is versioned and safe;
  • route and form behavior is tested from the user’s perspective.

Misconceptions to leave behind (Part 1)

MisconceptionBetter mental model
State management means choosing a libraryFirst classify ownership and lifetime
All shared state belongs in a global storeShare only across real consumers
Server data becomes ordinary client stateIt has freshness, cache, and synchronization rules
The URL is just routingIt is public, serializable view state
Every option belongs in the URLEphemeral and sensitive values have other owners
Forms are collections of controlled inputsForms are workflows with drafts and transitions

Misconceptions to leave behind (Part 2)

MisconceptionBetter mental model
Controlled forms are always betterChoose based on coordination needs
Draft and domain models must matchInput representation can differ from accepted data
Dirty and touched mean the same thingThey record different interaction facts
A reducer is only for global stateLocal complex transitions benefit too
Back/forward is only the router’s problemHistory behavior is product architecture

The chapter in one sentence

Classify state by ownership and lifetime, keep one clear source of truth, and make routing and forms explicit state machines for the user’s journey.


Next: Chapter 9

The next chapter will build on state architecture with:

  • resilient asynchronous data flows;
  • request lifecycle design;
  • caching, invalidation, and optimistic updates;
  • loading and error boundaries;
  • race-free server synchronization.

Questions

Which value in your current application has two owners?

What should happen to it on reload, sharing, back navigation, and an interrupted request?

9 - Client–Server Communication, APIs & Cache Management

Chapter 9: design reliable HTTP boundaries, represent remote UI states, and manage caches, mutations, retries, and recovery.

Client–Server Communication, APIs & Cache Management

Make remote data predictable

Chapter 9

Polla Fattah


Today’s goal

Build a client that treats the network as an unreliable, stateful boundary.

We will connect:

  • HTTP methods, headers, status codes, and request helpers;
  • runtime validation and transport/domain separation;
  • cancellation, timeouts, retries, and backoff;
  • REST, GraphQL, pagination, and API compatibility;
  • loading, empty, stale, error, and recovery states;
  • query keys, deduplication, freshness, and invalidation;
  • pessimistic and optimistic mutations;
  • server validation, authentication failures, and partial data.

By the end of today you can

  • implement a fetch boundary that handles HTTP errors explicitly;
  • distinguish network, transport, schema, domain, and authentication failures;
  • choose safe retry behavior;
  • model remote UI states without conflating loading and empty;
  • design stable cache keys and freshness policies;
  • deduplicate requests and cancel stale work;
  • invalidate or update caches after mutations;
  • use optimistic updates with rollback only when justified;
  • preserve useful partial data during dashboard failures;
  • draw the full client-server data architecture.

The central lesson

Remote data is not a value; it is a lifecycle of requests, cache entries, freshness, failures, mutations, and recovery.

The UI should make that lifecycle visible without forcing every component to understand HTTP details.


The chapter’s progression

flowchart TD
    A[HTTP Boundary] --> B[Transport & Domain Layers]
    B --> C[Request Lifecycle]
    C --> D[Remote UI States]
    D --> E[Query Cache]
    E --> F[Freshness & Invalidation]
    F --> G[Mutations & Rollback]
    G --> H[Resilient App Architecture]

Every layer answers a different question about remote data.


Browser and server have different responsibilities

browser: interaction, rendering, local drafts, navigation
server:  authority, persistence, authorization, business rules
network:  latency, loss, duplication, reordering, failure

The client cannot assume that the server is fast, available, current, or correct for every request.


HTTP is the communication foundation

request = method + URL + headers + body
response = status + headers + body

Each part carries contract information.

Do not treat an HTTP exchange as merely “fetch some JSON”.


HTTP methods express intent

MethodTypical intent
GETread a representation
POSTcreate or trigger an operation
PUTreplace a resource representation
PATCHpartially modify a resource
DELETEremove a resource

The exact API contract still matters; method names are not permission to guess semantics.


Safe and idempotent are different properties

safe      → intended not to change server state
idempotent → repeating the same request has the same intended result

GET is generally safe and idempotent.

PUT is commonly idempotent but not safe.

POST is not automatically idempotent, so retries require care.


Request headers are part of the contract

Accept: application/json
Content-Type: application/json
Authorization: Bearer …
If-None-Match: "version-42"

Headers communicate representation, credentials, caching, conditional requests, and client capabilities.


Accept and Content-Type answer different questions

Accept       what response representation can the client read?
Content-Type what representation is this request body?

Confusing them can produce content negotiation and parsing bugs that look like application errors.


Status codes are part of the API contract

2xx success
3xx redirection / cache interaction
4xx client or authorization problem
5xx server failure

The client should classify status codes into user-visible recovery paths instead of displaying one generic failure for everything.


fetch() does not reject for every HTTP error

const response = await fetch("/api/products");

if (!response.ok) {
  throw new HttpError(response.status);
}

A 404 or 500 response can resolve normally.

Network failures typically reject; HTTP failure status still needs explicit handling.


A small fetch helper

async function request<T>(input: RequestInfo, init?: RequestInit): Promise<T> {
  const response = await fetch(input, init);
  if (!response.ok) throw new HttpError(response.status);
  const payload: unknown = await response.json();
  return parseResponse<T>(payload);
}

The helper centralizes transport behavior, but it must not hide meaningful errors or bypass runtime validation.


Network and domain layers should not be confused

HTTP adapter → transport response → parser → domain model → feature

The adapter knows status codes and headers.

The domain feature knows what a valid product or workflow means.

Keeping those responsibilities separate makes change and testing cheaper.


JSON is a representation, not a type guarantee

const payload: unknown = await response.json();

JSON can represent an object that is:

  • missing fields;
  • using the wrong types;
  • from an older server version;
  • structurally valid but semantically invalid.

Parse at the boundary before trusting it.


Request bodies need an explicit representation

await fetch("/api/products", {
  method: "POST",
  headers: { "Content-Type": "application/json" },
  body: JSON.stringify(command),
});

The request body is a transport representation, not automatically the same as the form draft or domain entity.


Credentials and cookies change the boundary

fetch("/api/me", {
  credentials: "include",
});

Authentication, CSRF protection, CORS, expiration, and logout behavior are part of the client-server contract.

Never treat a missing credential as an ordinary empty response.


Cancellation is correctness, not only optimization

const controller = new AbortController();

fetch(url, { signal: controller.signal });
controller.abort();

Cancellation prevents obsolete work from consuming resources or updating a UI that now represents a different request.


Timeouts are application policy

const timeout = setTimeout(() => controller.abort(), 8_000);
try {
  return await fetch(url, { signal: controller.signal });
} finally {
  clearTimeout(timeout);
}

Choose timeout behavior based on user action, operation cost, connectivity, and retry policy.


Retry carefully

Retrying can help transient failures.

It can also:

  • duplicate a non-idempotent mutation;
  • overload a struggling server;
  • delay a useful error;
  • hide an authorization failure;
  • create a thundering herd.

Retry only when the operation and policy justify it.


Exponential backoff spreads retries

flowchart TD
    Att1["Attempt 1: Fails (Network / 5xx)"] -->|"Wait 1,000ms"| Att2["Attempt 2: Fails"]
    Att2 -->|"Wait 2,000ms + Jitter"| Att3["Attempt 3: Fails"]
    Att3 -->|"Wait 4,000ms + Jitter"| Att4["Attempt 4: Abort / Surface Error"]

Add jitter so many clients do not retry at the same instant.

Bound the number of attempts and expose a recovery action when automatic retry ends.


API design affects front-end architecture

The API determines:

  • what can be fetched independently;
  • how mutations are represented;
  • which fields are stable;
  • how pagination works;
  • which errors can be recovered;
  • which cache entries need invalidation.

Client architecture cannot fully compensate for an ambiguous server contract.


REST is resource-oriented

GET    /products
GET    /products/p-42
PATCH  /products/p-42
DELETE /products/p-42

Resource-oriented URLs can map naturally to cache keys and route identities.

They are a style, not a promise that every API has identical semantics.


Collections and resources have different concerns

collection: filters, sorting, pagination, totals
resource:   identity, fields, permissions, mutation

Do not assume that invalidating one product automatically answers what should happen to every filtered collection containing it.


Pagination is part of the cache model

products?page=1&limit=20
products?page=2&limit=20

Decide whether pages are:

  • independent cache entries;
  • concatenated into an infinite list;
  • invalidated together after a mutation;
  • prefetched based on navigation.

Filtering and sorting belong in the request identity

/products?query=phone&sort=price&stock=in

Changing any relevant input should produce a distinct query key or an explicit cache update.

Never reuse a cache entry for a request with different semantics.


API versioning protects compatibility

Versioning may be represented through:

  • URL paths;
  • headers;
  • content negotiation;
  • additive schema evolution.

The client should validate the response it actually receives, even when generated types say the version is known.


GraphQL asks for a shape of data

query Products($query: String) {
  products(query: $query) {
    id
    title
    priceCents
  }
}

The client requests fields and receives a schema-described response.

This changes transport shape and caching questions; it does not eliminate them.


GraphQL schemas and mutations

Schemas can provide:

  • typed fields;
  • discoverable operations;
  • nested data selection;
  • validation at the operation boundary.

Mutations still need authorization, error handling, cache updates, conflict policies, and runtime behavior.


REST and GraphQL can coexist

Use the boundary that fits the resource and team needs.

A system may use REST for uploads and resource mutations, GraphQL for composed reads, and a separate stream for live updates.

The architecture should make each boundary’s semantics explicit.


Remote data has UI states

type QueryState<T> =
  | { status: "idle" }
  | { status: "loading" }
  | { status: "success"; data: T; stale: boolean }
  | { status: "error"; error: RemoteError; previous?: T };

Represent the lifecycle instead of reducing it to data plus isLoading.


Loading is not the same as empty

loading → data not known yet
empty   → request succeeded; result set has zero items

The recovery and user message differ.

An empty catalogue may need onboarding.

A loading catalogue needs progress or preserved context.


Loading UI should match scope

whole route loading     → route fallback
table refreshing         → table status / preserved rows
save button submitting   → button pending state

A global spinner can erase useful context and make the application feel slower.


Spinner, skeleton, or existing content?

spinner    unknown short operation
skeleton   first load with known content shape
existing   background refresh where current data remains useful

Choose based on what the user can still do and how much layout is known.


Empty state is product state

An empty result may mean:

  • no records exist yet;
  • filters are too restrictive;
  • the user lacks access;
  • a search has no matches;
  • the resource was removed.

Show the cause and the next useful action when possible.


Error states need recovery

retry → same request again
edit filters → change request identity
sign in → resolve authentication
go back → leave invalid route

An error component should answer: what failed, what remains usable, and what can the user do next?


Client caching is a policy

A cache answers:

  • may this result be reused?
  • for how long?
  • when should it revalidate?
  • how is it invalidated?
  • can stale data remain visible?
  • who owns the cache?

Caching is not simply “store the last response”.


Freshness depends on the data

country list       long freshness window
product inventory  short freshness window
bank balance       very short / explicit refresh

Freshness is a product and domain decision, not one global number.


Stale does not mean wrong

Stale data can still be useful while a background request checks for newer data.

The UI should communicate that it is refreshing when the distinction matters.

Discarding useful data on every refresh failure can create a worse experience than showing stale data with a clear warning.


Stale-while-revalidate

flowchart LR
    CacheHit["Cache Hit"] --> RenderCached["Render Cached Data Immediately"]
    RenderCached --> BackgroundReq["Background Revalidation Request"]
    BackgroundReq --> Update["Update Cache & Silent UI Re-render"]

This pattern separates immediate usefulness from eventual freshness.

It requires a policy for errors, timestamps, and concurrent requests.


Cache keys identify query meaning

const key = ["products", { query, sort, page, stock }];

A cache key must include every input that changes the response.

It should be deterministic, serializable, and stable across callers.


Missing key inputs cause incorrect reuse

Bad:

["products"]

when the response varies by query, page, sort, account, or locale.

The cache then returns data that is valid for a different request but wrong for this view.


Deduplicate identical requests

flowchart LR
    CompA["Component A"] --> Dedup["In-Flight Request Deduplicator\n(Promise Cache)"]
    CompB["Component B"] --> Dedup
    CompC["Component C"] --> Dedup
    Dedup --> Single["Single HTTP Network Request"]
    Single --> Shared["Shared Server Result Broadcast"]

Deduplication reduces duplicate work and makes concurrent consumers observe one request lifecycle.

The cache must distinguish an in-flight promise from a completed fresh value.


Revalidation asks whether data changed

Revalidation can happen:

  • when data becomes stale;
  • when a route gains focus;
  • after a mutation;
  • on explicit refresh;
  • after reconnecting;
  • through a conditional HTTP request.

Choose triggers that match the user’s need for freshness.


Invalidation removes confidence, not necessarily data

flowchart TD
    Mut["Mutation Succeeds on Server"] --> Inval["Mark Related Cache Keys Stale / Invalid"]
    Inval --> Refetch["Trigger Active Query Revalidation"]
    Refetch --> Fresh["Render Canonical Server Truth"]

Invalidation is a statement that cached knowledge may no longer be current.

It does not always mean immediately deleting useful visible data.


Invalidation versus direct cache update

Use invalidation when:

  • many related queries may change;
  • server logic computes fields the client cannot reproduce;
  • correctness is more important than avoiding a request.

Directly update one entry when:

  • the mutation response is authoritative;
  • the affected cache shape is known;
  • the update is easy to verify.

Mutations have their own lifecycle

flowchart LR
    Idle["Idle"] --> Submitting["Submitting"]
    Submitting --> Success["Success"]
    Submitting --> ErrorState["Error"]
    ErrorState -->|Retry| Submitting

Add:

  • duplicate-submission protection;
  • server validation handling;
  • authentication behavior;
  • cache reconciliation;
  • draft preservation on failure.

Mutation UX communicates authority

flowchart TD
    Pess["Pessimistic: Wait for authoritative server 200 OK before updating UI"]
    Opt["Optimistic: Mutate UI immediately; rollback if server rejects request"]

Choose based on reversibility, conflict risk, user expectations, and the cost of being briefly wrong.


Prevent duplicate submission at several layers

Client disabling helps the interface.

It does not guarantee that duplicate requests cannot arrive.

Use server-side idempotency keys or operation identifiers where repeating an operation has real cost.


Pessimistic updates are conservative

flowchart TD
    Sub["Submit Command"] --> Pending["Display Inline Pending State"]
    Pending --> Ok["Server Success: Update Canonical UI"]
    Pending --> Err["Server Error: Preserve Draft & Show Recovery Action"]

Use this for high-risk, non-reversible, or conflict-sensitive operations where showing unconfirmed state would mislead users.


Optimistic updates trade certainty for responsiveness

flowchart LR
    Action["User Action"] --> Snapshot["1. Snapshot Previous Cache"]
    Snapshot --> OptimisticUI["2. Update Cache & UI Instantly"]
    OptimisticUI --> Network["3. Dispatch Server Request"]
    Network --> Confirm["4a. Server 200: Confirm & Revalidate"]
    Network --> Rollback["4b. Server Error: Restore Snapshot & Alert"]

The client must know how to undo the change and what to show if the server rejects it.


Optimistic rollback needs a snapshot

const previous = cache.get(key);
cache.set(key, nextValue);

try {
  await save(command);
} catch (error) {
  cache.set(key, previous);
}

In real systems, also consider concurrent edits, invalidation, and a final revalidation.


Not every mutation should be optimistic

Avoid optimism when:

  • the action is irreversible;
  • server rules are complex;
  • conflicts are likely;
  • authorization may fail often;
  • rollback is ambiguous;
  • the visible result depends on server-side computation.

Fast feedback is not worth misleading the user.


Conflict and concurrency need a policy

Two clients may update the same resource.

Possible strategies include:

  • last write wins;
  • version checks;
  • conditional requests;
  • conflict UI;
  • server merge rules;
  • explicit refresh before editing.

The cache cannot solve a domain conflict by itself.


Conditional requests use server versions

If-Match: "version-42"

The server can reject a mutation when the resource changed since the client read it.

This prevents a stale edit from silently overwriting newer data.


Browser HTTP cache versus application cache

Browser HTTP cacheApplication/query cache
controlled by HTTP semanticscontrolled by application policy
stores representationsstores parsed/query-aware data
keyed by request semanticskeyed by query and domain inputs
works below application codeexposes freshness and invalidation

They can cooperate, but they are not the same layer.


Cache-Control is a server instruction

Cache-Control: max-age=60, stale-while-revalidate=300

HTTP caching can reduce network work before application code runs.

The application still needs its own policy for query state, mutations, and visible stale/error behavior.


Forms and server mutations

flowchart TD
    Draft["1. Form Draft State"] --> Validate["2. Validate Form Locally"]
    Validate --> Transport["3. Submit Transport Command via Fetch"]
    Transport --> ServerVal["4. Server-Side Validation & Authorization"]
    ServerVal --> Commit["5. Commit or Preserve Draft on Error"]

A successful request does not automatically mean every form field, cache entry, and route is now synchronized.


Native form submission remains useful

Native submission provides:

  • keyboard behavior;
  • browser integration;
  • progressive enhancement;
  • a clear submit event;
  • FormData serialization.

Enhance it when application behavior requires it; do not discard platform behavior without a reason.


FormData is a transport boundary

const formData = new FormData(form);
const email = formData.get("email");

Values may be strings, files, or null.

Validate and convert them before building a domain command or JSON payload.


JSON form submission is an explicit conversion

const command = parseProductDraft(readDraft(formData));

await fetch("/api/products", {
  method: "POST",
  headers: { "Content-Type": "application/json" },
  body: JSON.stringify(command),
});

The parser is where unfinished input becomes a validated transport model.


Server validation errors are different from local errors

local: email format is invalid
server: email is already registered

Both may appear near a field, but they have different owners, timing, and recovery paths.

Do not overwrite a useful server response with a generic client message.


Map server errors to fields carefully

type ServerError =
  | { kind: "field"; field: string; message: string }
  | { kind: "form"; message: string };

Only map an error to a field when the server contract identifies that field reliably.

Some failures belong to the form or page, not one input.


Do not merge client and server validation carelessly

Track the sources distinctly:

clientErrors
serverErrors

Clear server errors when the relevant draft changes if the server response is no longer applicable.

Keep the message that explains the actual current failure.


Progressive enhancement for forms

flowchart TD
    Native["1. Native HTML form submit works without JavaScript"] --> Enh["2. Progressive Enhancement with JS: Pending states, client validation & SPA transitions"]

The enhanced path should preserve the core meaning of the native form rather than creating a completely different contract.


File uploads have different transport needs

const body = new FormData(form);

await fetch("/api/upload", {
  method: "POST",
  body,
});

Do not set a JSON content type for multipart data manually.

Model upload progress, cancellation, size limits, type validation, and partial failure explicitly.


Authentication failures are not ordinary validation errors

401 → establish or refresh identity
403 → identity exists but action is forbidden
422 → submitted data violates application rules

The correct response may be sign-in, permission explanation, or field correction - not a red error beneath an input.


Search needs query identity, cache, and cancellation

flowchart LR
    QA["Query A (/search?q=ca)"] --> CancelA["Abort in-flight Controller A"]
    CancelA --> QB["Query B (/search?q=cat)"]
    QB --> Network["Execute Request B"]

An old response must not replace results for a newer query.

The query key should include every input that changes the result.


Query keys and URL state fit naturally together

URL: /products?query=phone&page=2
key: ["products", "phone", 2]

The URL provides the committed view identity.

The query cache stores the server result for that identity.

Parse once and derive both from the same validated model.


Avoid copying query data into another store

query cache → components

Copying into a general client store creates:

  • duplicate ownership;
  • stale copies;
  • unclear invalidation;
  • extra synchronization effects.

Keep server data in a server-aware cache unless there is a deliberate transformation or offline model.


Dependent queries form a data graph

flowchart LR
    User["currentUser Query"] -->|"Resolves accountId"| Account["account Query"]
    Account -->|"Resolves credentials"| Tx["transactions Query"]

Start the dependent request only when its required input is known.

Represent the not-ready state instead of sending malformed requests with missing identifiers.


Request waterfalls cost time

A completes → B starts → C starts

If B and C do not depend on A, request them in parallel.

If they do depend, consider server composition, prefetching, or a route loader that understands the graph.


Parallel fetching

const [products, categories] = await Promise.all([
  loadProducts(),
  loadCategories(),
]);

Use parallel work for independent requests.

Choose Promise.allSettled() or explicit result handling when partial success is useful.


Partial failure can preserve useful data

flowchart TD
    Dash["Dashboard View"]
    Dash --> Q1["Sales Widget (Status: Success)"]
    Dash --> Q2["Inventory Widget (Status: Error + Retry Button)"]
    Dash --> Q3["Alerts Widget (Status: Success)"]

Do not blank the entire dashboard because one independent panel failed.

The architecture should allow panels to own their own request lifecycle.


Promise.all() versus partial failure

const results = await Promise.allSettled(requests);

Promise.all() is appropriate when all results are required for one operation.

allSettled() or independent query states are appropriate when each panel can recover separately.


Background refresh preserves context

fresh visible data + isRefreshing = true

Keep current rows visible while a refresh runs when stale data remains useful.

Show that the data is refreshing without turning every background check into a blank loading screen.


Stale data with an error is still a meaningful state

previous data + refresh failed

Tell the user that the displayed data may be out of date and offer retry.

This is often more useful than replacing known content with an empty error page.


Mutation followed by revalidation

flowchart LR
    Save["Mutation 200 OK"] --> Inval["Invalidate Related Query Keys"]
    Inval --> Refetch["Background Refetch of Canonical Server Truth"]

Revalidation is safest when the server computes fields, permissions, totals, or relationships the client cannot reproduce.


Mutation response as a cache update

flowchart LR
    Patch["PATCH /products/:id"] --> Resp["Response Payload Contains Canonical Product"]
    Resp --> Direct["Update ['product', id] Cache Directly (Zero extra GET)"]

This can avoid a request when the response is authoritative and the affected query shapes are known.

Still consider list ordering, filters, totals, and related cache entries.


Optimistic cache update

flowchart LR
    Snap["1. Cache Snapshot"] --> Prov["2. Provisional UI"]
    Prov --> Req["3. Async Request"]
    Req --> Resolution["4. Confirm / Rollback / Revalidate"]

The cache update should be isolated, reversible, and tested for failure and concurrency.


Server-state libraries provide mechanisms

They may offer:

  • query keys;
  • caching;
  • stale time;
  • retries;
  • deduplication;
  • invalidation;
  • mutation lifecycles;
  • optimistic updates.

They do not decide API semantics, domain ownership, or whether a mutation is safe to retry.


A useful separation of responsibilities

flowchart TD
    Adapter["Transport Adapter: HTTP methods, headers, status codes"]
    Query["Query Cache Layer: Freshness, deduplication, staleTime"]
    Mapper["Domain Mapper: Transforms untrusted JSON into domain entities"]
    Feature["Feature Coordinator: User intent, pagination, mutations"]
    View["View Layer: Renders data, skeletons, and error recovery"]
    Adapter --> Query --> Mapper --> Feature --> View

Each layer should expose the information its caller needs without leaking every lower-level detail.


Do not hide every HTTP detail behind one giant service

A universal api.requestEverything() wrapper often obscures:

  • status-specific recovery;
  • cancellation;
  • request identity;
  • cache behavior;
  • mutation semantics.

Centralize repeated mechanics, but keep meaningful operation contracts visible.


Transport errors versus domain errors

network unavailable     transport failure
HTTP 500                server/transport failure
malformed JSON          schema failure
valid but forbidden     authorization/domain failure
valid but unavailable   domain/business failure

Different causes require different UI and retry behavior.


Normalize errors without leaking secrets

type RemoteError = {
  kind: "network" | "http" | "schema" | "auth" | "domain";
  message: string;
  retryable: boolean;
};

Normalize enough for the UI to make a decision.

Do not expose stack traces, tokens, internal SQL, or sensitive payloads to users.


Administrative catalogue architecture

URL             query, filters, sort, page
server state    products, categories, cache
local UI        modal, focus, panel open
form state      draft, dirty, validation
transport       request, status, cancellation

The catalogue becomes easier to reason about when each concern has one owner.


Product query lifecycle

flowchart TD
    URL["1. Parse URL & Query State"] --> Key["2. Build Deterministic Query Key"]
    Key --> Cache["3. Read Fresh Cache or Trigger Fetch"]
    Cache --> Validate["4. Validate Schema at Trust Boundary"]
    Validate --> Expose["5. Expose Loading / Stale / Error / Data"]
    Expose --> UI["6. Render Data & Recovery Actions"]

The component should consume a query model, not reinvent this lifecycle.


Editing a product

flowchart TD
    Load["1. Load Canonical Product into Server Cache"] --> Draft["2. Create Detached Local Edit Draft"]
    Draft --> Validate["3. Validate Inputs Locally"]
    Validate --> Submit["4. Submit Transport Command via Fetch"]
    Submit --> ErrorBranch["5a. Failure: Preserve Draft & Show Alert"]
    Submit --> SuccessBranch["5b. Success: Update/Invalidate Cache & Navigate"]

Never let a failed mutation erase the user’s unfinished work by default.


Search and cancellation

let activeController: AbortController | undefined;

async function search(query: string) {
  activeController?.abort();
  activeController = new AbortController();
  return loadProducts(query, activeController.signal);
}

The cache and UI must also ensure that an older result cannot win after a newer query.


Pagination and cache retention

Decide:

  • whether previous pages remain visible;
  • whether the next page is prefetched;
  • how long old pages remain cached;
  • how a mutation changes page membership;
  • what happens when filters change.

Pagination is a user experience and cache policy together.


Prefetching remote data

Prefetch when:

  • the next destination is predictable;
  • the request is safe;
  • the cost is bounded;
  • the data is likely to be used.

Do not prefetch every possible route and overwhelm the network or cache.


A client cache is not a database

Cache entries can:

  • expire;
  • be evicted;
  • be incomplete;
  • fail to persist;
  • disagree with the server;
  • belong to one user or permission context.

Never use a cache as the sole authority for durable business data.


Server state and offline state differ

Offline support requires decisions about:

  • local persistence;
  • queued mutations;
  • conflict resolution;
  • retry scheduling;
  • user visibility;
  • security and data expiry.

A query cache alone does not create an offline architecture.


Practical lab: Cached Administrative API Client

Build a small API client that separates HTTP caching, application/query caching, runtime validation, and user-visible failure states.

The practical makes request identity, freshness, invalidation, mutation behavior, and partial failure observable.


Practical stages 1–4: establish the boundary

  1. Implement a fetch() wrapper that checks response.ok explicitly.
  2. Validate response data at the boundary.
  3. Model remote UI states.
  4. Add search and cancellation.

Simulate 401, 403, 404, 422, 500, timeout, and malformed-data responses.


Practical stages 5–8: query cache fundamentals

  1. Build a simple query cache.
  2. Add freshness.
  3. Deduplicate requests.
  4. Add pagination.

Verify that query keys include every relevant input and that stale data remains distinguishable from fresh data.


Practical stages 9–13: mutations and cache truth

  1. Add a mutation.
  2. Handle server validation.
  3. Invalidate after save.
  4. Directly update one cache entry.
  5. Add an optimistic toggle with rollback.

Document why each mutation uses pessimistic, direct-update, or optimistic behavior.


Practical stages 14–17: resilience and architecture

  1. Simulate partial dashboard failure.
  2. Compare browser cache and query cache.
  3. Add native FormData submission.
  4. Draw the full data architecture.

Verification: loading, empty, stale, error, recovery, and partial-success states are visible.


Practical extension: conditional requests

Add ETags and conditional requests:

If-None-Match: "version-42"

Document which layer owns:

  • the HTTP cache;
  • the query cache;
  • freshness;
  • invalidation;
  • the final domain model.

Try this yourself

Design the cache keys for:

products by query, sort, page
product by ID
current user
dashboard panel by account and date range

Then list which mutations invalidate or directly update each key.


Troubleshooting guide (Part 1)

SymptomLikely cause
500 response enters success coderesponse.ok was not checked
Old search result replaces new resultmissing cancellation or request identity
Different filters show the same datacache key omits an input
Every refresh blanks the screenstale data is discarded unnecessarily

Troubleshooting guide (Part 2)

SymptomLikely cause
Failed save loses the draftmutation lifecycle owns the form incorrectly
Retry duplicates an operationnon-idempotent request has no safety policy
One panel breaks the dashboardindependent queries were coupled with Promise.all()
Cache never updates after saveinvalidation or direct update is undefined

Completion checklist

  • HTTP errors and network errors are distinct;
  • responses are validated before entering domain code;
  • request cancellation prevents stale work;
  • retries are limited and operation-aware;
  • remote UI states distinguish loading, empty, stale, and error;
  • cache keys include all relevant inputs;
  • freshness and invalidation policies are explicit;
  • mutations preserve drafts on failure;
  • optimistic updates have rollback and revalidation rules;
  • partial failures preserve independent useful data.

Misconceptions to leave behind (Part 1)

MisconceptionBetter mental model
fetch() rejects for every 4xx or 5xxCheck response.ok explicitly
TypeScript proves the server responseRuntime validation proves delivered data
Every request should be retriedRetry only safe, useful operations
Loading and empty are the sameOne is pending; one is a successful zero result
Cache means correct foreverCache means reusable knowledge under a policy

Misconceptions to leave behind (Part 2)

MisconceptionBetter mental model
Stale means unusableStale data may be useful while revalidating
Every mutation should be optimisticChoose based on reversibility and conflict risk
Browser cache and query cache are identicalThey are different layers with different owners
A server-state library replaces API designIt provides mechanisms, not semantics
Successful save fixes every cacheRelated entries still need reconciliation

The chapter in one sentence

Treat remote data as a lifecycle with explicit request identity, validation, freshness, failure, mutation, and recovery policies.


Next: Chapter 10

The next chapter will build on client-server architecture with:

  • accessibility and inclusive interaction;
  • semantic structure and assistive technology;
  • keyboard, focus, and form behavior;
  • robust component contracts;
  • testing the experience rather than only the implementation.

Questions

For one request in your application, can you explain its key, freshness policy, cancellation rule, retry policy, failure states, and invalidation path?

10 - Real-Time Communication, Offline Systems & Client Persistence

Chapter 10: choose live communication deliberately, persist useful client data, and synchronize offline work safely.

Real-Time Communication, Offline Systems & Client Persistence

Design for disconnection

Chapter 10

Polla Fattah


Today’s goal

Build front-end systems that remain understandable when messages arrive, networks disappear, and local work must synchronize later.

We will connect:

  • polling, long polling, SSE, WebSockets, and WebRTC;
  • connection state, ordering, duplication, backpressure, and reconnection;
  • cookies, browser storage, IndexedDB, and Cache Storage;
  • service-worker lifecycle and caching strategies;
  • offline reads, offline writes, outboxes, and synchronization;
  • identity, conflicts, eventual consistency, and background sync;
  • a field-inspection architecture that works without a permanent connection.

By the end of today you can

  • choose the simplest communication model that satisfies the requirement;
  • prevent overlapping polls and stale live updates;
  • design SSE and WebSocket message envelopes;
  • model connection state as user-visible state;
  • distinguish transport delivery from application semantics;
  • choose browser storage by data meaning and lifetime;
  • explain service-worker install, activate, and fetch interception;
  • design cache-first, network-first, and stale-while-revalidate policies;
  • persist drafts and queued operations offline;
  • synchronize with idempotency, retry limits, and conflict detection.

The central lesson

Real-time and offline behavior are not switches; they are explicit policies for communication, persistence, synchronization, identity, and recovery.

The network may be delayed, duplicated, reordered, unavailable, or partially available.

The application must make those conditions meaningful rather than pretending they cannot happen.


The chapter’s progression

flowchart TD
    A["Communication Requirement"] --> B["Select Simplest Transport"]
    B --> C["Model Connection Lifecycle"]
    C --> D["Local Persistence Strategy"]
    D --> E["Service Worker & Cache Storage"]
    E --> F["Offline Reads & Writes"]
    F --> G["Outbox Synchronization"]
    G --> H["Conflict Resolution & Recovery"]

Complexity should be earned by a real product requirement.


“Real-time” is a requirement, not one technology

Ask:

  • how fresh must the data be?
  • who sends updates?
  • is communication one-way or two-way?
  • can a short delay be accepted?
  • must work continue without a network?
  • are messages durable or disposable?

The answers narrow the transport and persistence choices.


Start with the simplest communication model

flowchart LR
    A["Manual Refresh"] --> B["Short Polling"]
    B --> C["Long Polling"]
    C --> D["Server-Sent Events (SSE)"]
    D --> E["WebSockets / WebRTC"]

Use the least complex model that satisfies freshness and interaction needs.

Do not add a bidirectional socket when periodic reads are sufficient.


Polling is repeated HTTP

setInterval(async () => {
  const updates = await fetch("/api/updates");
  render(await updates.json());
}, 10_000);

Polling is easy to deploy, observe, authorize, and cache.

Its cost is repeated requests even when nothing changed and delayed delivery between intervals.


Polling interval is a trade-off

short interval → fresher data, more requests and battery use
long interval  → less work, more visible delay

Choose based on the business meaning of freshness, not an arbitrary “real-time” label.


Poll only when useful

Pause or reduce polling when:

  • the document is hidden;
  • the user leaves the relevant route;
  • the device is offline;
  • the data is not visible or actionable;
  • a long-lived connection already supplies updates.

Resume with an explicit refresh when the user returns.


Avoid overlapping poll requests

let active: AbortController | undefined;

async function poll() {
  active?.abort();
  active = new AbortController();
  return fetch("/api/updates", { signal: active.signal });
}

If a request takes longer than the interval, a naive timer can create concurrent requests and out-of-order responses.


Long polling is still repeated HTTP

sequenceDiagram
    autonumber
    participant Client
    participant Server
    Client->>Server: HTTP GET /events (hangs open)
    Note over Server: Server delays response until event occurs
    Server-->>Client: 200 OK (Event payload)
    Note over Client: Client processes event
    Client->>Server: HTTP GET /events (immediately reopens)

Long polling reduces empty responses while preserving an HTTP-shaped deployment model.

It still needs cancellation, timeout, reconnection, and duplicate handling.


Server-Sent Events are one-way streams

const events = new EventSource("/api/events");

events.onmessage = event => {
  const message: unknown = JSON.parse(event.data);
  handleEvent(message);
};

SSE is useful when the client sends commands through ordinary HTTP and the server streams updates back.


SSE architecture

flowchart LR
    subgraph Client["Browser Client"]
        Cmd["Mutation Action"]
        Listener["EventSource Listener"]
    end
    subgraph Server["Server API"]
        HTTP["POST /api/commands"]
        Stream["GET /api/events (text/event-stream)"]
    end
    Cmd -->|Standard HTTP POST| HTTP
    Stream -->|Unidirectional Stream| Listener

The one-way direction simplifies some authorization and infrastructure concerns compared with a fully bidirectional socket.


When SSE fits well

Use SSE for:

  • notifications;
  • progress updates;
  • monitoring dashboards;
  • assignment changes;
  • server-generated status events.

It is a poor fit when the client must exchange frequent messages in both directions over one connection.


Commands over HTTP, events over SSE

sequenceDiagram
    autonumber
    participant App as Browser Client
    participant API as Municipal REST API
    participant SSE as SSE Stream Server

    App->>API: POST /assignments/a-1/complete (HTTP)
    API-->>App: 200 OK { id: "a-1", status: "completed" }
    API->>SSE: Broadcast domain event
    SSE-->>App: event: assignment.updated
data: { id: "a-1", status: "completed" }

Keeping commands and events separate can make authorization, retries, and audit behavior clearer.

The event remains an announcement; the server remains authoritative.


Named SSE events improve intent

events.addEventListener("assignment.updated", event => {
  const payload: unknown = JSON.parse(event.data);
  handleAssignmentUpdate(payload);
});

Named events avoid one generic handler having to infer every message kind from an unstructured payload.


SSE reconnect is part of the contract

An interrupted stream should define:

  • reconnect delay;
  • maximum or bounded backoff;
  • authentication refresh;
  • missed-event recovery;
  • duplicate-event handling;
  • user-visible connection status.

Reopening the stream alone does not guarantee that no update was missed.


Events need identity when necessary

{
  "id": "event-1042",
  "type": "assignment.updated",
  "version": 8,
  "entityId": "assignment-7"
}

An event ID or entity version lets the client detect duplicates, gaps, and stale updates.


WebSockets are bidirectional transports

client ⇄ persistent WebSocket connection ⇄ server

They fit interactive collaboration, presence, chat, live controls, and high-frequency two-way communication.

They also create more lifecycle and protocol responsibility.


WebSocket architecture

stateDiagram-v2
    [*] --> Connecting: new WebSocket(url)
    Connecting --> Authenticating: onopen
    Authenticating --> Subscribed: auth token accepted
    Subscribed --> Active: bidirectional framing
    Active --> Active: message / ping / pong
    Active --> Reconnecting: onclose / onerror
    Reconnecting --> Connecting: exponential backoff
    Active --> Closed: user disconnect / logout
    Closed --> [*]

The socket is one part of the application architecture, not the architecture itself.


A WebSocket is a transport, not an application protocol

The transport does not decide:

  • message types;
  • authentication refresh;
  • ordering guarantees;
  • idempotency;
  • authorization;
  • persistence;
  • conflict resolution;
  • missed-event recovery.

Those rules belong in the message protocol and domain design.


Design message envelopes deliberately

type Message = {
  id: string;
  type: "assignment.updated" | "inspection.acknowledged";
  version: number;
  occurredAt: string;
  payload: unknown;
};

An envelope gives the client enough information to route, validate, order, and observe messages.


Runtime validation still applies

socket.addEventListener("message", event => {
  const value: unknown = JSON.parse(event.data);
  const message = parseMessage(value);
  if (!message.ok) return showProtocolError(message.error);
  applyMessage(message.data);
});

TypeScript cannot guarantee that a remote peer sent the expected data.


Connection state is UI state

type ConnectionState =
  | { status: "offline" }
  | { status: "connecting" }
  | { status: "connected" }
  | { status: "reconnecting"; attempt: number }
  | { status: "failed"; message: string };

Users need to know whether an action is live, queued, delayed, or unavailable.


Reconnection is not “just reopen the socket”

After reconnecting, the client may need to:

  • refresh credentials;
  • resubscribe;
  • request a snapshot;
  • replay safe commands;
  • detect missed versions;
  • discard obsolete local assumptions.

Reconnect is a synchronization sequence.


Snapshot plus events is a robust pattern

flowchart TD
    A["1. GET /api/snapshot
(Baseline state at t0)"] --> B["2. Connect Event Stream / WebSocket
(Subscribe to mutations)"]
    B --> C["3. Buffer In-Flight Events
(Queue events arriving during fetch)"]
    C --> D["4. Apply Ordered Updates
(Discard events older than snapshot version)"]

The snapshot provides a baseline.

Events provide changes after that baseline.

The protocol must define the race between reading the snapshot and subscribing.


Ordering cannot be assumed

Messages can arrive:

  • late;
  • out of order;
  • duplicated;
  • after a reconnect;
  • from an earlier connection.

Use sequence numbers, versions, timestamps with care, or a server reconciliation step.


Duplicates are normal in reliable systems

At-least-once delivery can repeat a message.

Make handlers idempotent where possible:

flowchart TD
    Ev["Incoming Live Event (version, eventId)"] --> Check{"Version > Current Local Version?"}
    Check -->|No: Already applied or obsolete| Ignore["Drop / Acknowledge Safely (Idempotent)"]
    Check -->|Yes: Exact next version| Apply["Apply update to UI state & cache"]
    Check -->|Gap detected: Version > Current + 1| Resync["Buffer event & trigger snapshot reconciliation"]

Exactly-once behavior is usually an application-level illusion built from IDs and state.


Backpressure protects the client

If messages arrive faster than the UI or storage can process them, define a policy:

  • coalesce updates;
  • drop obsolete intermediate states;
  • pause subscriptions;
  • apply batches;
  • request a fresh snapshot;
  • show degraded status.

Unbounded queues turn a temporary burst into a memory and responsiveness problem.


WebSocket reconnect needs bounded backoff

flowchart TD
    Drop["Socket Disconnected"] --> Wait1["Wait Base Delay (e.g. 500ms + Jitter)"]
    Wait1 --> Try1["Attempt Reconnect"]
    Try1 -->|Failure| Wait2["Wait Exponential Delay (1000ms + Jitter)"]
    Wait2 --> Try2["Attempt Reconnect"]
    Try2 -->|Failure| Max["Cap at Max Delay (e.g. 10s) & Surface Disconnected Banner"]

Add jitter and stop retrying when the failure is permanent, such as invalid credentials or forbidden access.


Polling, SSE, or WebSocket?

RequirementGood starting point
occasional freshnesspolling
server-to-client streamSSE
high-frequency two-way interactionWebSocket
intermittent network with durable workHTTP plus local outbox
peer media/dataWebRTC

The product requirement should choose the transport.


WebRTC is a different kind of real-time

WebRTC is designed for peer media and data.

It still needs:

  • signaling;
  • identity and authorization;
  • NAT traversal;
  • relay infrastructure;
  • connection lifecycle;
  • application-level message rules.

It is not a replacement for ordinary server events.


Signaling establishes a connection

sequenceDiagram
    autonumber
    participant PeerA as Peer A
    participant Sig as Signaling Server (HTTP/WS)
    participant PeerB as Peer B

    PeerA->>Sig: Send SDP Offer + ICE Candidates
    Sig->>PeerB: Forward Offer
    PeerB->>Sig: Send SDP Answer + ICE Candidates
    Sig->>PeerA: Forward Answer
    Note over PeerA,PeerB: Direct P2P Media / DataChannel Established
    PeerA<<-->>PeerB: Direct P2P DataChannel / Media

The signaling channel helps peers exchange connection information.

The actual media or data path may then be direct or relayed.


NAT traversal and relay

Some network environments prevent direct peer connectivity.

STUN can help discover a reachable address.

TURN can relay traffic when direct connection fails.

Real-time architecture must budget for the cases where the ideal path is unavailable.


Client persistence starts with meaning

Ask what must survive:

  • reload;
  • tab close;
  • browser restart;
  • route change;
  • offline time;
  • service-worker update.

Choose storage after defining lifetime, size, sensitivity, and access pattern.


Cookies

Cookies are sent with requests according to domain, path, and policy rules.

They are useful for server-managed sessions but require careful security attributes:

Secure + HttpOnly + SameSite + limited scope

Do not use cookies as a general client database.


localStorage

localStorage.setItem("theme", "dark");
const theme = localStorage.getItem("theme");

It is convenient for small string preferences and simple drafts.

It is synchronous, string-only, quota-limited, and not a secure store.


localStorage is synchronous

Large reads and writes can block the main thread.

Avoid using it for large datasets, frequent updates, or high-volume logs.

For larger asynchronous data, consider IndexedDB or an application-specific persistence layer.


localStorage stores strings

localStorage.setItem("settings", JSON.stringify(settings));

const raw = localStorage.getItem("settings");
const value: unknown = raw === null ? null : JSON.parse(raw);

Validate versions and shape on read. Data written by the application is still an external runtime boundary after reload.


sessionStorage has a shorter lifetime

It is scoped to a browser tab session.

It can suit temporary per-tab state that should survive reload but not a new tab or later session.

The same string-only and synchronous limitations still apply.


Storage events coordinate tabs

window.addEventListener("storage", event => {
  if (event.key === "theme") applyTheme(event.newValue);
});

Storage events can notify other documents, but they are not a durable event bus or a complete synchronization protocol.


IndexedDB is an asynchronous local database

Use it for:

  • larger structured data;
  • offline records;
  • drafts and outboxes;
  • indexes and queries;
  • data that should not block rendering.

The asynchronous API adds complexity, but supports a more appropriate data model.


IndexedDB is transactional

flowchart LR
    Tx["db.transaction(['inspections', 'outbox'], 'readwrite')"] --> Ops["Execute Reads & Writes"]
    Ops --> Success["All operations succeed
→ Automatic Commit"]
    Ops --> Fail["Any error thrown
→ Automatic Complete Abort (Rollback)"]

Group related updates so a draft and its outbox record cannot silently diverge.

Design transaction boundaries around invariants the application must preserve.


IndexedDB is asynchronous

Plan for:

  • request errors;
  • aborted transactions;
  • blocked upgrades;
  • unavailable storage;
  • concurrent tabs;
  • schema migration.

The database is local, but it is still a failure-prone boundary.


IndexedDB versioning is a migration contract

flowchart TD
    Open["indexedDB.open('MunicipalApp', 2)"] --> Check{"Requested Version > Current DB Version?"}
    Check -->|No| Ready["onsuccess: Database ready for transactions"]
    Check -->|Yes| Upgrade["onupgradeneeded: Run migrations
- createObjectStore()
- createIndex()
- transform existing records"]
    Upgrade --> Ready

Test upgrades from realistic previous versions.

Do not assume users start with an empty database after a release.


Cache Storage stores request/response pairs

const cache = await caches.open("app-assets-v1");
await cache.put(request, response.clone());

It fits resources retrieved through the Fetch model, especially assets and selected HTTP responses.

It is not interchangeable with a domain database.


Cache Storage versus IndexedDB

Cache StorageIndexedDB
request/response pairsstructured application records
asset and HTTP-like retrievalqueries, indexes, transactions
service-worker friendlydomain/offline data friendly
cache strategypersistence and synchronization model

Choose based on data semantics, not only size.


Storage is not guaranteed forever

Data may be evicted because of:

  • quota pressure;
  • user settings;
  • browser policy;
  • private browsing;
  • device storage constraints;
  • application cleanup.

Offline architecture must tolerate missing local data and rehydrate from the server when possible.


Storage quotas affect product behavior

Large offline datasets need:

  • size limits;
  • eviction policy;
  • user-visible storage status;
  • cleanup rules;
  • recovery when writes fail.

Do not promise indefinite offline history unless the platform and product support that promise.


Choose storage by data semantics

flowchart TD
    subgraph BrowserStorageTaxonomy["Browser Storage by Purpose & Scope"]
        S1["Session Identity → Cookies (HttpOnly, Secure)"]
        S2["Small User Preferences → localStorage (<5MB, sync API)"]
        S3["Structured Offline Data → IndexedDB (Async, indexed, large quota)"]
        S4["HTTP Asset Responses → Cache Storage API (Request/Response pairs)"]
        S5["Durable Queued Operations → IndexedDB Outbox (Transactional)"]
    end

One application can use several stores with explicit ownership.


Service workers run outside the page

flowchart LR
    Page["Active Browser Window / Page"] <-->|fetch() / Navigation| SW["Service Worker
(self.addEventListener('fetch'))"]
    SW <-->|Cache Match / Put| Cache["Cache Storage API"]
    SW <-->|Network Request| Net["Remote Network Server"]

The worker has a different lifecycle and execution context.

It can intercept fetches and support caching, but it does not automatically understand domain data or user intent.


Secure context is required

Service-worker features generally require a secure context such as HTTPS, with localhost development commonly treated specially.

Deployment configuration is part of the offline architecture.


Service-worker lifecycle

stateDiagram-v2
    [*] --> Installing: navigator.serviceWorker.register()
    Installing --> Waiting: self.skipWaiting() / new SW downloaded
    Waiting --> Activating: Old SW clients closed / skipWaiting
    Activating --> Active: clients.claim()
    Active --> Active: Intercepting network requests
    Active --> Redundant: Replaced by updated script
    Redundant --> [*]

An updated worker may not control the current page immediately.

Design update messaging and cache compatibility rather than assuming an instant replacement.


Installation prepares resources

self.addEventListener("install", event => {
  event.waitUntil(precacheAssets());
});

Installation should be bounded and versioned.

Do not make one missing optional asset prevent the entire application from installing unless that is intentional.


Activation cleans up and takes ownership

self.addEventListener("activate", event => {
  event.waitUntil(deleteOldCaches());
});

Activation is a migration boundary for caches and control behavior.

Coordinate old and new asset versions so the page does not combine incompatible files.


Fetch interception is policy code

self.addEventListener("fetch", event => {
  event.respondWith(handleRequest(event.request));
});

The worker must decide which requests it can safely handle and which should pass through.

Do not cache credentials, private data, or mutations accidentally.


A service worker does not mean offline automatically

Offline support requires:

  • a cache strategy;
  • data persistence;
  • a fallback UI;
  • offline write behavior;
  • synchronization;
  • conflict policy;
  • update and eviction handling.

Registration is only the start.


Cache-first strategy

flowchart TD
    Req["Incoming fetch(event.request)"] --> Cache{"Cache.match(request)?"}
    Cache -->|Hit| Fast["Return cached Response (Instant)"]
    Cache -->|Miss| Net["Fetch from Network"]
    Net --> Put["cache.put(request, clone)"]
    Put --> Res["Return fresh Response"]

Good for versioned static assets or content where immediate availability is more important than freshness.

Risk: stale content can persist if versioning and invalidation are weak.


Network-first strategy

flowchart TD
    Req["Incoming fetch(event.request)"] --> Net{"Network fetch()"}
    Net -->|Success| Put["cache.put(request, clone)"]
    Put --> Res["Return fresh server response"]
    Net -->|Failure / Offline| Cache{"Cache.match(request)?"}
    Cache -->|Hit| Stale["Return cached offline fallback"]
    Cache -->|Miss| Err["Return custom offline error page"]

Good for data that should be fresh when connectivity exists but remain readable offline.

Risk: slow networks can delay the fallback unless timeouts are defined.


Stale-while-revalidate

flowchart TD
    Req["fetch(event.request)"] --> Cache{"Cache.match(request)?"}
    Cache -->|Hit| Ret["Return cached response immediately"]
    Cache -->|Miss| WaitNet["Await network response"]
    Ret --> BG["Async background fetch()"]
    BG --> Put["Update Cache Storage for next load"]
    WaitNet --> Put

Good when fast display and eventual freshness are both useful.

The UI should communicate meaningful staleness.


Network-only and cache-only

flowchart LR
    subgraph NetworkOnly["Network-Only Strategy"]
        R1["fetch(request)"] --> N1["Server API"]
        N1 -->|Never Cache| UI1["Critical Mutations / Auth"]
    end
    subgraph CacheOnly["Cache-Only Strategy"]
        R2["fetch(request)"] --> C1["Cache Storage API"]
        C1 -->|Zero Network| UI2["Pre-cached Static Assets / Offline Fonts"]
    end

Do not apply one strategy to every request.

Mutation requests should not be silently treated like cacheable reads.


Strategy matrix

DataUseful strategy
versioned assetscache-first
current account datanetwork-first / query policy
readable offline cataloguestale-while-revalidate
mutation commandnetwork-only plus outbox when offline
private durable draftIndexedDB, not public cache

The right choice depends on authority, freshness, and recovery.


Offline fallback is a product surface

An offline screen should explain:

  • what remains available;
  • which work is saved locally;
  • what is waiting to sync;
  • what cannot be done now;
  • how the user can retry.

Offline is a mode with capabilities, not simply an error page.


Offline-capable versus offline-first

offline-capable → core paths work without connectivity when needed
offline-first    → local state is the primary interaction model

Offline-first is a larger product commitment involving conflict, identity, local storage, and synchronization.

Do not adopt it for a feature that only needs cached reading.


navigator.onLine is only a hint

It can indicate a network interface state.

It does not prove:

  • the server is reachable;
  • authentication works;
  • the request will succeed;
  • the route is available;
  • the network has useful bandwidth.

Treat actual requests and failures as stronger evidence.


Offline reads

flowchart LR
    IDB["IndexedDB Store"] --> Read["Read Cached Record"]
    Read --> Valid["Check freshness & validity"]
    Valid --> Render["Render view with Offline/Stale indicator"]

Offline reads should identify:

  • when data was last synchronized;
  • whether it is incomplete;
  • whether an update is pending;
  • when the next refresh can occur.

Offline writes need durable intent

flowchart LR
    User["User Submits Form"] --> Split["Atomic Transaction"]
    Split --> Rec["Update Local Record"]
    Split --> Box["Enqueue Outbox Operation"]
    Box --> Sync["Background Sync Engine"]

Do not keep the only copy of a user’s work in memory while waiting for the network.

Persist the draft or operation before telling the user it is safely queued.


The outbox pattern

flowchart TD
    Cmd["User Action: Submit Inspection"] --> Tx["Atomic IndexedDB Transaction"]
    Tx -->|Write| Domain["Write local 'inspections' store (Status: PendingSync)"]
    Tx -->|Write| Outbox["Write 'outbox' store (Operation: CREATE_INSPECTION)"]
    Outbox --> Sync{"Sync Trigger
(Online event, page load, visibility)"}
    Sync --> Post["POST /api/inspections with Idempotency-Key"]
    Post -->|200 Ack| Done["Remove outbox record, update local status to Synced"]
    Post -->|Network Drop| Retry["Increment attempt count, schedule backoff retry"]
    Post -->|4xx Fatal| Dead["Mark outbox item FAILED, notify inspector"]

The transaction keeps local work and its synchronization intent together.


Local identity versus server identity

An offline record may need:

classDiagram
    class InspectionRecord {
        +UUID localId "Generated client-side immediately (crypto.randomUUID())"
        +String serverId "Canonical ID assigned or confirmed by server (null while offline)"
        +UUID operationId "Unique idempotency token sent in outbox request"
        +String syncStatus "draft | pending_sync | syncing | synced | conflict"
        +Number version "Optimistic concurrency version tag"
    }

Do not use one identifier for three different meanings.


Pending synchronization is user-relevant state

Show whether work is:

stateDiagram-v2
    [*] --> Draft: User edits record
    Draft --> PendingSync: User commits inspection
    PendingSync --> Syncing: Network available & outbox flush
    Syncing --> Synced: Server returns 200 OK
    Syncing --> PendingSync: Transient 5xx / timeout (retry)
    Syncing --> Failed: 422 validation / 403 forbidden
    Failed --> Draft: User edits data to fix validation
    Synced --> [*]

Users need confidence about whether their work is safe, not only whether the browser currently has a connection.


Sync is not the same as retry

Retry repeats an operation after a transient failure.

Synchronization reconciles local intent with remote authority after time and possibly other changes.

It may need identity, ordering, conflict detection, and a new domain decision.


Conflict example

server inspection status: open
device A offline:         completed
device B online:          reassigned

When device A reconnects, blindly overwriting the server may destroy a newer decision.

The system needs a conflict policy.


Conflict policies

flowchart TD
    subgraph ConflictResolution["Conflict Resolution Policies"]
        C1["Last Write Wins (LWW)
Clock timestamp determines winner (Risky)"]
        C2["Server Wins
Canonical authority; client state overwritten"]
        C3["Client Wins
Local user decision always takes precedence"]
        C4["Field-Level 3-Way Merge
Combine non-overlapping field edits"]
        C5["Manual User Resolution
Side-by-side visual diff prompt"]
    end

Choose based on data meaning and harm, not implementation convenience.


Version-based conflict detection

{
  "id": "inspection-7",
  "version": 12,
  "status": "open"
}

The client submits the version it edited.

The server rejects or resolves the operation when its current version differs.


Operation logs make synchronization inspectable (Part 1)

FieldTypeDescription
operationIdUUIDUnique idempotency key for this synchronization action
entitystringTarget domain entity (e.g. 'inspection')
commandstringAction type (e.g. 'SUBMIT_REPORT')
payloadJSONComplete serialized mutation payload

Operation logs make synchronization inspectable (Part 2)

FieldTypeDescription
createdAttimestampClient timestamp when user performed action
attemptsnumberRetry counter with maximum threshold
statusenum'pending_sync' | 'syncing' | 'failed'

Eventual consistency is a user experience

After an offline action, local state may say “completed” while the server has not confirmed it.

Represent that distinction:

locally complete, awaiting server confirmation

Do not present unconfirmed local intent as permanent server truth.


Background Sync is an enhancement

Background Sync can help flush queued work when the browser decides conditions are suitable.

The application must still work when it is unavailable, delayed, or denied.

Provide a visible manual retry and foreground synchronization path.


Periodic background sync needs restraint

Periodic work affects:

  • battery;
  • data usage;
  • privacy;
  • server load;
  • freshness expectations.

Use it only when the product meaningfully benefits from background refresh.


Service-worker update risk

An old page and a new worker can temporarily coexist.

Plan for:

  • compatible cache names;
  • atomic asset versions;
  • schema migration;
  • user messaging;
  • safe activation timing.

Updating a worker is a deployment and data-migration concern.


Cache versioning

const STATIC_CACHE = "app-shell-v3";

Version names make cleanup and incompatible asset replacement explicit.

Do not let old bundles and new runtime assumptions share one unbounded cache.


Do not cache every API response forever

For each response, define:

  • sensitivity;
  • freshness;
  • user scope;
  • size;
  • invalidation;
  • offline usefulness;
  • eviction.

Caching private data without a lifecycle can create security and correctness problems.


User identity and offline data

When a user signs out or changes account, decide what happens to local data:

  • delete it;
  • encrypt or isolate it;
  • keep only non-sensitive preferences;
  • mark it for a specific identity;
  • prevent another account from reading it.

Local persistence must respect authorization boundaries.


Cache API is not authorization

A cached response existing on the device does not prove the current user may still access it.

Authorization remains a server and application decision.

Clear or revalidate private cached data when identity or permissions change.


Progressive web apps are a capability set

A PWA may include:

  • installability;
  • service workers;
  • offline assets;
  • local data;
  • push notifications;
  • background synchronization.

Installability does not automatically imply offline data or reliable synchronization.


The app shell is only one layer

flowchart TD
    subgraph OfflineFieldSystem["The Four Pillars of Offline Resilience"]
        P1["1. App Shell (Cache Storage)
HTML, CSS, JS bundles cached for 0ms offline boot"]
        P2["2. Local Data Store (IndexedDB)
Read-only municipal permits, checklists, inspector profiles"]
        P3["3. Durable Outbox (IndexedDB)
Queued operations preserved across reboots & browser closes"]
        P4["4. Synchronization Engine
Background queue consumer with exponential backoff & idempotency"]
    end

Caching HTML, CSS, and JavaScript does not solve domain data or mutation conflicts.


Offline-first interaction

local write is primary
server sync is eventual

This can produce excellent responsiveness, but it requires durable local models, conflict handling, and clear confirmation states.


Network-first interaction

server is primary
local cache is fallback or acceleration

This is often appropriate for authoritative records where stale edits are risky and connectivity is usually available.


Select the offline scope

Possible scopes:

flowchart TD
    L1["Level 1: Offline Shell Only
App boots to static frame with 'No Connection' banner"]
    L2["Level 2: Offline Read Cache
User can browse previously viewed permits and checklists"]
    L3["Level 3: Offline Drafts
Unsaved form inputs persist across reboots in IndexedDB"]
    L4["Level 4: Offline Queued Outbox
Inspectors complete inspections offline; synced upon reconnect"]
    L5["Level 5: Full Local-First Workflow
CRDTs / multi-device peer synchronization with zero central locks"]
    L1 --> L2 --> L3 --> L4 --> L5

Choose the smallest scope that solves the user problem.


Real-time plus offline together

stateDiagram-v2
    state Online {
        [*] --> Streaming
        Streaming: WebSocket / SSE live events update local store
    }
    state Offline {
        [*] --> Autonomous
        Autonomous: Reads from IndexedDB; writes queued in Outbox
    }
    state Reconnecting {
        [*] --> Heartbeat
        Heartbeat: Egress probe succeeds
        Heartbeat --> FlushOutbox: Send queued operations with Idempotency-Key
        FlushOutbox --> InvalidateQueries: Refresh canonical server state
    }
    Online --> Offline: Connection lost
    Offline --> Reconnecting: Network regained
    Reconnecting --> Online: All operations settled

One local data model can be the bridge between live updates and offline work.


A reconnect sequence

flowchart TD
    A["1. Connectivity Hint (online event / window focus)"] --> B["2. Heartbeat Ping (verify real internet egress)"]
    B --> C["3. Validate Auth Token / Refresh Session"]
    C --> D["4. Fetch Server Version Vector / Changes"]
    D --> E["5. Flush Pending Outbox Operations with Idempotency"]
    E --> F["6. Detect & Resolve Concurrent Conflicts"]
    F --> G["7. Reconcile UI & Invalidate Fresh Server Queries"]

The order matters. Sending stale operations before understanding current server state can create avoidable conflicts.


Avoid double-applying events

outbox command succeeds → local record updated
server event arrives     → same change announced

Use operation IDs or server versions to recognize that the local optimistic change and remote event refer to one operation.


Data freshness after reconnect

After a long offline period, cached data may be:

  • outdated;
  • structurally migrated;
  • revoked by permissions;
  • superseded by a conflict;
  • incomplete.

Revalidate and communicate the result rather than silently labeling it current.


Practical project: Offline-Capable Field Inspections

Build a field-inspection application that can read assignments, save drafts, queue completed inspections, and recover when connectivity returns.

The practical combines live communication, persistence, service workers, outbox sync, idempotency, and conflict decisions.


Practical stages 1–3: communication models

  1. Implement polling.
  2. Prevent poll overlap.
  3. Replace polling with SSE.

Measure freshness, request count, cancellation, and behavior when the tab is hidden or offline.


Practical stages 4–7: live connection design

  1. Compare SSE and WebSocket.
  2. Add connection state.
  3. Design reconnection.
  4. Explore WebRTC conceptually.

Write down which message guarantees the application actually needs: ordering, identity, duplication handling, and recovery.


Practical stages 8–10: persistence foundations

  1. Add local preferences.
  2. Add IndexedDB.
  3. Register a service worker.

Test reload, tab close, unavailable storage, upgrade, and a missing or malformed local record.


Practical stages 11–13: caching and offline reads

  1. Implement cache-first assets.
  2. Implement network-first data.
  3. Add an offline fallback.

Make the strategy visible in the UI and document which requests are safe to cache.


Practical stages 14–17: offline writes and sync

  1. Build an offline draft.
  2. Create the outbox.
  3. Synchronize on reconnect.
  4. Add conflict detection.

Verification: records survive reload, duplicate submission is prevented or detected, and permanent failures remain recoverable.


Practical stages 18–20: operational hardening

  1. Treat Background Sync as an enhancement.
  2. Version the service-worker cache.
  3. Draw the full architecture.

The app must remain useful without Background Sync and must explain what is waiting to synchronize.


Practical extension: conflict resolution

Simulate a record changed on the server while the client was offline.

Compare:

server wins
client wins
field merge
manual resolution

Record which policy is safest for inspection evidence and why.


Try this yourself

Design the local records for one inspection:

inspection
draft
outbox operation
server version
local ID
operation ID
sync status

Then define the transaction that persists the draft and the queued operation together.


Troubleshooting guide (Part 1)

SymptomLikely cause
Poll responses arrive out of orderRequests overlap without identity or cancellation
Reconnected socket misses updatesNo snapshot or event-version recovery
Same event changes data twiceHandler lacks idempotency identity
Offline work disappears on reloadOnly in-memory state was used
Service worker is registered but offline failsNo request strategy or local data model

Troubleshooting guide (Part 2)

SymptomLikely cause
Old assets break with new workerCache versioning and activation are unsafe
Duplicate inspection is createdNo idempotency key or server deduplication
Local data leaks across accountsPersistence is not scoped or cleared on identity change
Sync retries foreverPermanent failures lack a terminal state

Completion checklist

  • the transport matches the actual freshness and direction requirements;
  • polling and live connections have cancellation and reconnection policy;
  • messages have validation and identity where needed;
  • connection state is visible to users;
  • storage is chosen by semantics and lifetime;
  • service-worker caches are versioned and scoped;
  • offline writes are durable before being acknowledged locally;
  • outbox operations have IDs, retry limits, and terminal failures;
  • synchronization detects duplicates, ordering issues, and conflicts;
  • the app remains useful without optional background capabilities.

Misconceptions to leave behind (Part 1)

MisconceptionBetter mental model
Real-time means WebSocketChoose the simplest transport that fits
Polling is outdatedPolling is often clear and sufficient
SSE and WebSocket are the sameDirection and protocol responsibilities differ
Reopening a socket solves reconnectReconnect also needs resync and recovery
Messages arrive exactly onceDesign for duplicates and reordering
localStorage is a databaseIt is synchronous string storage

Misconceptions to leave behind (Part 2)

MisconceptionBetter mental model
App-written local data is trustedReloaded storage is a runtime boundary
Service-worker registration means offlineStrategies, persistence, and recovery are still needed
navigator.onLine proves connectivityIt is only a hint
Retry and synchronization are identicalSync reconciles local intent with remote state
Background Sync is guaranteedIt is an enhancement, not the foundation
Every application should be offline-firstChoose an offline scope from user need

The chapter in one sentence

Design communication, persistence, and synchronization as explicit stateful systems that remain safe when connectivity is slow, absent, duplicated, or restored.


Next: Chapter 11

The next chapter will build on resilient client architecture with:

  • security boundaries and threat modeling;
  • authentication and authorization;
  • browser security policies;
  • safe handling of untrusted content;
  • defensive application design.

Questions

When the network disappears during an important user action, what exactly is saved, what is queued, what is visible, and what happens when the server has changed meanwhile?

11 - Rendering Topologies: CSR, SSR, SSG & Beyond

Chapter 11: choose where and when rendering happens, manage hydration and handoff costs, and design hybrid routes deliberately.

Rendering Topologies: CSR, SSR, SSG & Beyond

Choose where the work happens

Chapter 11

Polla Fattah


Today’s goal

Understand rendering as a placement and timing decision rather than a framework label.

We will connect:

  • client-side rendering, server-side rendering, and static generation;
  • hydration, serialization, handoff, and browser work;
  • partial hydration, islands, streaming, and hybrid routes;
  • revalidation, edge rendering, server components, and resumability;
  • cacheability, personalization, authentication, and failure modes;
  • performance as server, build, network, and device cost;
  • a route-specific rendering decision matrix.

By the end of today you can

  • explain where rendering happens and when it occurs;
  • compare CSR, SSR, and SSG without slogans;
  • identify hydration mismatch causes and handoff costs;
  • isolate interactive regions with islands or client boundaries;
  • design streaming boundaries that preserve UX and accessibility;
  • distinguish server components from server-rendered HTML;
  • reason about revalidation, edge execution, and resumability;
  • select a rendering strategy per route;
  • keep secrets and server-only data on the server;
  • measure the rendering cost triangle instead of optimizing one metric.

The central principle

Rendering architecture is the deliberate placement of work across build time, request time, and browser time according to freshness, personalization, interaction, cacheability, and device cost.

CSR, SSR, and SSG are tools in a continuum - not application-wide identities.


The four main questions

Ask:

  1. Where does rendering happen?
  2. When does it happen?
  3. What is transferred to the browser?
  4. What must execute in the browser?

The answers reveal the actual topology beneath a framework’s terminology.


Build time, request time, and browser time

flowchart LR
    B["1. Build Time
(Static precomputation,
CI/CD compilation)"] --> R["2. Request Time
(On-demand server render,
data resolution per user)"]
    R --> BR["3. Browser Time
(Client JavaScript execution,
hydration, reactivity)"]

Every topology moves work among these three moments.


Rendering is work

Rendering may include:

  • fetching data;
  • transforming content;
  • producing HTML;
  • serializing state;
  • parsing JavaScript;
  • hydrating event handlers;
  • committing DOM updates;
  • running device-side interaction.

“Rendered” does not mean “free”. It means the cost moved somewhere.


Client-side rendering

flowchart TD
    A["1. Browser receives empty HTML shell (<div id='root'>) & JS bundle"] --> B["2. Browser downloads & executes JavaScript"]
    B --> C["3. Client fires fetch() to API server"]
    C --> D["4. Client calculates virtual DOM / reactive graph"]
    D --> E["5. DOM created and painted (Content visible at last)"]

CSR moves much of the initial rendering work to the device.


CSR strengths

CSR can provide:

  • rich interaction after startup;
  • simple browser-side state ownership;
  • app-like navigation;
  • direct access to browser APIs;
  • a consistent client runtime.

It is often appropriate for authenticated tools and highly interactive workspaces.


CSR costs

The browser may need to:

  • download a larger JavaScript bundle;
  • parse and execute before content appears;
  • fetch data after startup;
  • render on lower-powered devices;
  • manage loading and error states before the first useful view.

The network and device become part of the initial critical path.


CSR → where initial rendering happens
SPA → how navigation and application shell are organized

A site can use client rendering without being one large SPA.

An SPA can include server-rendered initial HTML.

Do not use the terms as interchangeable architecture decisions.


CSR and search engines

Search engines may execute JavaScript, but discoverability, timing, metadata, content quality, and operational behavior still matter.

Public content often benefits from HTML being available before client execution.

The correct choice depends on the content and delivery requirements.


Server-side rendering

flowchart TD
    A["1. Browser requests URL (GET /permits/104)"] --> B["2. Server fetches database data & executes components to HTML"]
    B --> C["3. Browser receives full HTML document (Instant FCP)"]
    C --> D["4. Browser downloads client JavaScript bundle"]
    D --> E["5. Client hydrates HTML: attaches listeners & reconciles state"]

SSR moves initial rendering work to request time and improves the first document’s content availability.


SSR changes the critical path

The request may now wait for:

  • server data;
  • server rendering;
  • serialization;
  • network transfer;
  • browser parsing;
  • hydration.

SSR can improve first content while adding server latency and operational complexity.


SSR does not mean no JavaScript

If the page must respond to:

  • clicks;
  • typing;
  • menus;
  • client navigation;
  • live state;

some browser code still needs to execute.

SSR determines how initial output is produced, not whether interaction exists.


SSR does not automatically improve every page

SSR can be a poor fit when:

  • content is highly personalized and uncacheable;
  • server data is slow;
  • the page is mostly an internal interactive tool;
  • hydration cost dominates;
  • deployment cannot support reliable request rendering.

Measure the whole request-to-interaction path.


Time to first byte matters

request → server work → first byte → HTML parsing → useful content

A server-rendered page with a slow first byte may feel worse than a carefully designed static or client-rendered page.

Rendering location alone does not determine speed.


Static site generation

flowchart LR
    A["Build Command
(Fetch data at compile)"] --> B["Generate Static HTML
(/docs/intro.html)"]
    B --> C["Deploy to Edge CDN
(Globally distributed)"]
    C --> D["Instant Edge Serve
(<20ms TTFB globally)"]

SSG moves rendering cost to deployment or build time.

It works especially well for public content with predictable freshness and high cacheability.


SSG strengths

SSG can provide:

  • fast delivery from a CDN;
  • simple serving infrastructure;
  • strong cacheability;
  • stable public HTML;
  • low request-time computation.

The trade-off is build-time freshness and build complexity.


Static generation moves cost to deployment

Large sites may pay through:

  • long builds;
  • many pages;
  • content invalidation;
  • preview workflows;
  • deployment coordination.

“Static” is operationally simple at request time, not necessarily cheap everywhere.


Static does not mean non-interactive

static HTML + client island = interactive page

A generated page can include search, menus, forms, and other browser behavior.

Static describes when initial output was produced, not the absence of JavaScript.


Static data can become stale

Define a freshness policy:

  • rebuild on content change;
  • revalidate periodically;
  • regenerate on demand;
  • add client refresh;
  • show the publication time.

SSG requires a plan for change, not an assumption that published data never changes.


CSR, SSR, and SSG describe initial rendering

After initial delivery, all three may:

  • fetch new data;
  • navigate without a full document;
  • run client components;
  • update the DOM;
  • synchronize with external systems.

Do not infer the entire runtime architecture from the first render label.


Hydration reuses existing HTML

flowchart LR
    HTML["Server-Rendered HTML
(Passive DOM structure)"] & JS["Client JavaScript Bundle
(Component definitions & handlers)"] --> Hydrate["Hydration Step
- Walk DOM tree
- Attach event listeners
- Initialize reactive state"]
    Hydrate --> Interactive["Fully Interactive Page"]

Hydration is a handoff from server-produced markup to an interactive browser runtime.

It is not simply a second independent render with no cost.


Hydration mismatches are signals

A mismatch can come from:

  • random values;
  • current time or timezone;
  • browser-only APIs;
  • different data on server and client;
  • unstable IDs;
  • conditional rendering based on client detection.

The warning often indicates an unclear rendering boundary or nondeterministic initial model.


Browser-only APIs need a boundary

const width = window.innerWidth;

This cannot run during server rendering.

Options include:

  • client-only component;
  • post-hydration effect;
  • server-safe default;
  • explicit capability check outside initial render.

Deterministic initial rendering

The server and client should agree on the initial representation.

Avoid using during initial render:

  • random values;
  • current time without a shared value;
  • locale-dependent formatting without a fixed locale;
  • browser state unavailable to the server.

Make the handoff data explicit.


Hydration has CPU cost

The browser may need to:

  • download JavaScript;
  • parse and compile it;
  • execute component code;
  • attach listeners;
  • reconstruct client state;
  • process event boundaries.

HTML existing on screen does not mean the page is ready for interaction.


The server-to-client handoff

flowchart TD
    A["Server fetches Database Entities"] --> B["renderToString(App) + JSON.stringify(data)"]
    B --> C["Network Transfer (HTML payload + <script id='__DATA__'>)"]
    C --> D["Browser parses HTML and JSON data"]
    D --> E["Client framework reconstructs identical Component State"]
    E --> F["Hydration completes (TTI reached)"]

Design the handoff as a contract.

Transfer only what the browser needs and keep secrets server-side.


Serialization has constraints

Not every server-side value can cross the boundary safely:

  • functions;
  • database connections;
  • secrets;
  • file handles;
  • cyclic structures;
  • framework-specific objects;
  • sensitive records.

Use explicit transport models rather than passing arbitrary server objects.


Secrets must stay server-side

flowchart TD
    subgraph ServerZone["Secure Server Boundary"]
        Sec["Database Credentials / API Secrets / Private Columns"]
        Comp["Server Component / Data Loader"]
        Sec --> Comp
    end
    Comp -->|Explicit Public DTO / Props ONLY| Wire["Network Wire (HTML + Client State)"]
    Wire --> BrowserZone["Browser Client Runtime"]
    Sec -.->|BLOCKED: Never pass credentials or unscrubbed DB rows| Wire

SSR does not make a secret safe if the secret is embedded in HTML or serialized props.

Rendering on the server is not the same as authorization.


Hydration is not free because HTML exists

Measure:

  • time to first byte;
  • first contentful paint;
  • time to interactive;
  • JavaScript transfer;
  • hydration CPU;
  • input delay;
  • post-hydration data work.

The right topology minimizes the cost that matters for the route and its users.


Partial hydration

static HTML
  + hydrate only selected interactive regions

Partial hydration reduces browser work when most of the page is content or static structure.

It requires clear boundaries for state and events.


Islands architecture

flowchart TD
    subgraph StaticDocument["Zero-JS Static HTML Page"]
        H["Header & Layout HTML"]
        Content["Article / Product Text HTML"]
        Footer["Footer HTML"]
    end
    StaticDocument -.-> I1["Island 1: Navigation Menu
(client:load)"]
    StaticDocument -.-> I2["Island 2: Search Autocomplete
(client:idle)"]
    StaticDocument -.-> I3["Island 3: Interactive Comments
(client:visible)"]

Each island can hydrate independently.

The default becomes HTML with isolated interactive regions rather than one full-page client runtime.


Islands change the default

Instead of asking “how do we render the whole app in the browser?” ask:

which regions truly need client behavior?

This can reduce JavaScript, but introduces coordination decisions when islands need shared state.


Island hydration timing

Possible policies:

load immediately
hydrate when visible
hydrate on idle
hydrate on interaction
hydrate only on a specific condition

Choose timing based on user value and interaction urgency.


Islands have boundaries

Cross-island communication may need:

  • URL state;
  • custom events;
  • shared server state;
  • browser storage;
  • a client coordinator.

Do not recreate a hidden global application just to make every island know about every other island.


Streaming changes delivery timing

sequenceDiagram
    autonumber
    participant Browser
    participant Server
    Browser->>Server: GET /dashboard
    Note over Server: Fast data ready immediately
    Server-->>Browser: Flush HTTP headers + App Shell + Fast HTML
    Note over Browser: Browser parses & renders App Shell (Fast FCP!)
    Note over Server: Slow analytics query resolves after 400ms
    Server-->>Browser: Streamed HTML chunk <template id='chunk-1'>
    Server-->>Browser: Inline script swaps placeholder with chunk
    Note over Browser: Analytics widget renders without page reload

Streaming can improve perceived progress and reduce waiting for the slowest dependency.

It does not automatically eliminate total server, network, or browser work.


Streaming boundaries are UX boundaries

Each boundary should define:

  • what appears first;
  • what loading state is shown;
  • what remains interactive;
  • what happens on failure;
  • how layout remains stable.

The stream is part of the product experience, not only a transport optimization.


Suspense as a boundary concept

shell
  ├─ ready navigation
  └─ pending product panel

A suspense-like boundary lets the system reveal available content while a slower region resolves.

Use meaningful fallbacks rather than arbitrary spinners everywhere.


Streaming and accessibility

As content arrives progressively:

  • preserve logical document order;
  • do not move focus unexpectedly;
  • announce important updates appropriately;
  • keep headings and landmarks coherent;
  • avoid making keyboard users wait for a needed control without feedback.

Delivery timing must preserve usable structure.


Streaming and layout stability

Reserve predictable space where possible.

Unexpected late content can:

  • shift reading position;
  • move a focused control;
  • cause accidental clicks;
  • increase cumulative layout shift.

Loading UI is a layout contract.


SSR and streaming can mix

server-rendered shell → streamed data-dependent regions → client interaction

Streaming is a delivery strategy layered onto server rendering.

It does not create a separate universal topology that replaces SSR or SSG.


Static generation and streaming can also mix

A statically generated shell can load or stream a dynamic region later through client or edge behavior.

The route may combine:

  • prebuilt public content;
  • request-time personalization;
  • browser-side interactivity.

Hybrid is often a precise choice, not architectural inconsistency.


Hybrid rendering is route-level strategy

home page      → SSG / revalidation
catalogue      → SSR or cached hybrid
account        → SSR with personalization
admin editor   → CSR or interactive hybrid

Choose per route and sometimes per region within a route.


Revalidation and incremental generation

serve existing output
  → determine it is stale
  → regenerate in background or on demand
  → publish new output

Incremental strategies reduce full rebuild cost while preserving static delivery for many requests.


Incremental static regeneration

An incrementally generated page has:

  • a cached representation;
  • a freshness window or invalidation event;
  • regeneration work;
  • a fallback while regeneration occurs;
  • failure behavior.

The important design is the stale and failure policy, not the product name of the mechanism.


Stale-while-revalidate rendering

sequenceDiagram
    autonumber
    actor User1
    participant CDN as Edge CDN / Cache
    participant Server as Origin Server / Static Generator
    actor User2

    User1->>CDN: GET /permits/104 (Cache Stale, past max-age)
    CDN-->>User1: Return cached stale HTML instantly (0ms latency)
    CDN->>Server: Trigger background regeneration
    Server->>Server: Re-fetch database & re-render HTML
    Server-->>CDN: Update cache with fresh HTML version
    User2->>CDN: GET /permits/104 (5 seconds later)
    CDN-->>User2: Return fresh pre-rendered HTML

This works well for content where slightly stale output is acceptable and fast delivery matters.


Edge rendering

Edge execution moves request-time work closer to users or data paths.

Potential benefits:

  • lower network distance;
  • request-aware personalization;
  • distributed cache integration.

It is still request-time server work with runtime and deployment constraints.


Edge runtime trade-offs

Edge environments may constrain:

  • Node-specific APIs;
  • native modules;
  • filesystem access;
  • connection lifetimes;
  • debugging and observability;
  • regional consistency.

“Edge” is not automatically faster or simpler.


Render near the data or near the user?

near data → lower backend latency / private access
near user → lower delivery distance / edge personalization

The best location depends on data source, cacheability, user distribution, and runtime capabilities.


Server components are not the same as SSR

SSR            → when initial HTML is produced
Server component → where a component's logic and data access execute

A server component may participate in different rendering and navigation strategies.

Do not collapse execution location and HTML delivery into one concept.


Server components can run at different times

Depending on the framework and route, server-side component work may happen:

  • at build time;
  • at request time;
  • during a client navigation that requests a new payload;
  • during revalidation.

The route’s data and cache policy still determine the actual behavior.


Server components can reduce client JavaScript

They can keep data access and noninteractive rendering on the server.

Only interactive boundaries need client behavior.

This can reduce browser bundle responsibility, but the handoff and client boundaries still have costs.


Client components own browser behavior

Use a client boundary for:

  • event handlers;
  • local interactive state;
  • browser APIs;
  • effects and subscriptions;
  • client-only libraries.

Keep the boundary as small as the interaction allows.


Server/client boundaries are architectural boundaries

flowchart TD
    subgraph ServerComponent["Server Component (PermitDetail.server.tsx)"]
        DB[("Direct SQL / Internal Microservice")] --> SC["Executes exclusively on Server
- Zero client bundle footprint
- Keeps SQL drivers & secrets on server"]
    end
    SC -->|Passes Serialized Props| CC["Client Component ('use client')
(PermitActionButtons.tsx)
- Handles onClick, hover, local UI state"]

The boundary controls:

  • what crosses the network;
  • what JavaScript ships;
  • where authorization occurs;
  • which state is reconstructed in the browser.

Data can be loaded near server rendering

Loading data close to server-rendered components can avoid duplicating the initial request in the browser.

But make the client handoff explicit:

  • is the data serialized?
  • is it cached?
  • does the client revalidate?
  • can the browser receive only a projection?

The server component payload has a cost

The browser may receive a structured description of server/client boundaries and data needed for the client tree.

Measure:

  • payload size;
  • serialization work;
  • parsing;
  • client boundary setup;
  • repeated requests during navigation.

Moving logic server-side does not make all transfer cost disappear.


Hydration still matters with server components

Client components still need browser execution and interaction setup.

Server components can reduce the amount that hydrates, but they do not remove the need to design client boundaries, identity, and state handoff.


Next.js as a case study

Next.js can combine:

  • static generation;
  • server rendering;
  • route-level revalidation;
  • server and client components;
  • streaming;
  • client navigation.

Treat these as mechanisms for route decisions, not as a substitute for the rendering mental model.


Nuxt as a case study

Nuxt can combine:

  • universal rendering;
  • static generation;
  • route rules;
  • server handlers;
  • client hydration;
  • hybrid deployment.

Again, understand where data and rendering work occur rather than memorizing framework labels.


Frameworks should not become the mental model

Ask framework-independent questions:

  • when is the HTML produced?
  • where is data loaded?
  • what crosses to the browser?
  • which region hydrates?
  • how is freshness controlled?
  • what happens on navigation and failure?

Then map the answers to framework configuration.


The “double data” oversimplification

SSR does not always mean the same data is fetched twice.

Possible designs include:

  • server output only for static content;
  • serialized data reused by the client;
  • a client revalidation request by policy;
  • server component payloads with selective transfer.

Inspect the actual network and handoff path.


Streaming does not eliminate handoff cost

Even if HTML arrives progressively, the browser may still need to:

  • parse client JavaScript;
  • hydrate interactive regions;
  • receive additional data;
  • attach event behavior.

Streaming improves timing of availability; it does not erase client responsibility.


Resumability changes the handoff shape

flowchart TD
    subgraph HydrationModel["Hydration Model (React / Vue)"]
        H1["Download all component JS"] --> H2["Execute all component functions"]
        H2 --> H3["Rebuild VDOM & attach listeners"]
    end
    subgraph ResumabilityModel["Resumability Model (Qwik)"]
        R1["Zero JS executed on initial load"] --> R2["DOM serialized with state & handler symbols"]
        R2 --> R3["User clicks button → Download & execute ONLY that event handler"]
    end

Resumability can reduce eager browser execution by preserving more execution context across the server boundary.


Hydration versus resumability

HydrationResumability
eagerly reconstructs client behaviorresumes behavior on demand
predictable component startup costmore serialization and constraints
familiar event setup modelfiner-grained lazy execution
browser executes more initiallybrowser may do less initial work

Both require stable boundaries and explicit data transfer.


Resumability has constraints

The system must preserve enough information to resume:

  • event handlers;
  • closures or state references;
  • component identity;
  • serialized data;
  • dependency relationships.

Reducing initial work can move complexity into build output and programming constraints.


Rendering topologies form a continuum

flowchart LR
    A["Pure Static
(SSG)"] <--> B["Incremental
(ISR)"]
    B <--> C["Server-Side
(SSR)"]
    C <--> D["Streaming
(SSR + Suspense)"]
    D <--> E["Islands
(Partial Hydration)"]
    E <--> F["Pure Client
(CSR / SPA)"]

Real applications can occupy several points at once.


Route decision: marketing homepage

Often:

  • public;
  • highly cacheable;
  • content changes on a publishing schedule;
  • limited interaction.

SSG or revalidated static output is often a strong starting point, with small client islands for interaction.


Route decision: product catalogue

Often:

  • public or semi-public;
  • filterable and paginated;
  • freshness matters;
  • URL state matters;
  • some interaction is client-side.

Hybrid SSR/SSG with URL-driven client behavior or cached server queries may fit better than one whole-site rule.


Route decision: account page

Often:

  • personalized;
  • permission-sensitive;
  • less cacheable publicly;
  • form and interaction heavy.

Request rendering or a client application with a secure data boundary may be appropriate.


Route decision: rich internal editor

Often:

  • authenticated;
  • interaction-dense;
  • stateful;
  • device-aware;
  • not valuable to search engines.

CSR or a hybrid shell with focused server data boundaries may be simpler and more effective.


A route decision matrix (Part 1)

QuestionPush toward
public and rarely changing?SSG / revalidation
personalized per request?SSR / client data
high interaction density?client boundaries / CSR
slow independent region?streaming

A route decision matrix (Part 2)

QuestionPush toward
mostly static with small interactions?islands
expensive browser startup?server rendering / partial hydration
strict freshness and authorization?request-time server boundary

Personalization reduces cacheability

Separate:

public cacheable shell
  + user-specific region

Do not make an entire page uncacheable when only one small region depends on identity.

The boundary can be server-rendered, client-loaded, or streamed according to risk and UX.


Authentication and rendering

Server rendering can check authentication before producing private output.

But ensure:

  • private data is not cached publicly;
  • redirects do not leak resource existence;
  • serialized props do not contain excess data;
  • client boundaries enforce appropriate behavior too.

Rendering location does not replace authorization policy.


Error handling across topologies

build failure       → deployment / publishing response
server request fail → route or region error boundary
hydration mismatch  → model and handoff fix
client request fail → interactive recovery state

Each topology moves failure to a different phase.


Route loading across topologies

CSR: client shell → data and route code
SSR: request → server data and HTML
streaming: shell → regions as ready
SSG: prebuilt output → client enhancement

Choose loading states that match the actual phase the user is waiting through.


Navigation after initial load

Client navigation can avoid a full document request, but it still may require:

  • route data;
  • server component payload;
  • code chunks;
  • cache lookup;
  • state preservation;
  • scroll and focus management.

Initial rendering strategy does not fully determine navigation behavior.


Full document navigation still has value

Full navigation can provide:

  • a clean server boundary;
  • lower client runtime assumptions;
  • reliable reset of page state;
  • progressive enhancement;
  • simpler failure recovery.

Do not remove it merely because client navigation is fashionable.


Progressive enhancement and rendering

semantic HTML and server behavior
  → enhanced navigation and interaction

The basic task should remain understandable and usable when client code is delayed or unavailable, when the product permits it.


Performance is multidimensional

Measure:

  • server latency;
  • build time;
  • HTML size;
  • JavaScript transfer;
  • hydration CPU;
  • device memory;
  • interaction readiness;
  • cache hit rate;
  • freshness and error recovery.

Optimizing one number can worsen the actual experience.


JavaScript budget

Client JavaScript has costs in:

  • transfer;
  • parse and compile;
  • memory;
  • event setup;
  • hydration;
  • battery;
  • later navigation.

Ship browser code because the user needs the behavior, not because the framework can render it there.


HTML is already a runtime format

HTML provides:

  • document structure;
  • links;
  • forms;
  • semantic controls;
  • accessibility relationships;
  • progressive behavior.

Do not replace platform capabilities with client code without a clear benefit.


Server work has operational cost

Request rendering consumes:

  • compute;
  • memory;
  • connection capacity;
  • data-source capacity;
  • observability and deployment complexity.

SSR is not free work merely because the browser does less.


Build work has operational cost

Static generation consumes:

  • build time;
  • CI resources;
  • content pipeline capacity;
  • deployment storage;
  • invalidation and preview complexity.

Choose build-time rendering when its freshness and delivery benefits justify the workflow.


Client work has device cost

Browser work varies with:

  • device CPU;
  • memory;
  • battery;
  • network conditions;
  • browser capability;
  • competing applications.

Do not treat a fast developer laptop as a universal client.


The rendering cost triangle

flowchart TD
    subgraph CostTriangle["The Web Rendering Cost Triangle"]
        SC["Server & Request Cost\n(CPU, DB connections, edge compute bill)"]
        BC["Build Cost\n(CI/CD minutes, deploy frequency, build queue)"]
        CC["Client & Device Cost\n(Battery, main-thread blocking, RAM, TTI latency)"]
        SC --- BC
        BC --- CC
        CC --- SC
    end

Moving work away from one corner usually adds cost or constraints to another.


Content freshness changes the choice

rarely changes      → build / cache
changes hourly      → revalidation / background regeneration
changes per request → SSR or client query
changes continuously→ live client boundary

Freshness is a route requirement, not a framework preference.


User-specific state changes the choice

Private data may require:

  • request-time authorization;
  • private caching;
  • client-side loading after a public shell;
  • isolated personalized regions.

Separate public and private regions when possible to preserve cacheability.


Interaction density changes the choice

content-heavy → more server/static work
interaction-heavy → more client responsibility
mixed route → isolate interactive regions

The amount and locality of interaction matter more than whether a page is “modern”.


Practical project: one platform, several topologies

Compare the same product catalogue requirements through CSR, SSR, SSG, streaming, and an island-like boundary.

Record data freshness, interaction, handoff, JavaScript, caching, and failure assumptions for each version.


Practical stages 1–4: establish the comparison

  1. Establish route data and interaction requirements.
  2. Build a CSR baseline.
  3. Create a minimal server-rendered version.
  4. Create a static version with explicit freshness assumptions.

Do not compare only screenshots. Compare the work and data path behind each result.


Practical stages 5–9: handoff and streaming

  1. Add streaming or a delayed independent region.
  2. Record the server-to-client handoff and hydration cost.
  3. Design streaming boundaries.
  4. Compare one large boundary with many tiny boundaries.
  5. Identify which regions truly need browser execution.

Measure timing and layout stability as well as total output.


Practical stages 10–14: hybrid framework concepts

  1. Compare Next.js conceptually.
  2. Compare Nuxt conceptually.
  3. Build a hybrid route plan.
  4. Isolate personalization.
  5. Explore islands.

The final choice should be route-specific rather than application-wide dogma.


Practical stages 15–19: reduce and justify browser work

  1. Delay island hydration.
  2. Explore resumability conceptually.
  3. Build a rendering decision matrix.
  4. Measure JavaScript responsibility.
  5. Draw the final hybrid architecture.

Verification: each topology has a stated freshness and interaction model, and server-only data stays server-side.


Practical extension: authenticated account route

Compare the public catalogue with an authenticated account route.

Explain why the decision differs in terms of:

  • cacheability;
  • authorization;
  • personalization;
  • interaction density;
  • data freshness;
  • handoff and failure behavior.

Try this yourself

For one application, classify these routes:

home page
product catalogue
product detail
account page
admin editor

For each, choose where rendering happens, when it happens, what crosses to the browser, and what must be interactive.


Common rendering smells

Watch for:

  • entire site forced into CSR because one page is interactive;
  • entire application SSR because “SSR is faster”;
  • every component marked client-side;
  • huge framework bundle for a static page;
  • personalized region preventing the whole page from caching;
  • slow widget blocking all server HTML;
  • same data fetched on server and immediately fetched again on client.

Hydration mismatch is usually a modeling signal

Instead of suppressing the warning, ask:

  • which value differs?
  • who owns it?
  • was it available on both sides?
  • should it be client-only?
  • should the server serialize a stable value?
  • is the initial output deterministic?

Fix the rendering model before hiding the symptom.


Avoid browser detection during render

const isMobile = window.innerWidth < 768;

The server and client may choose different trees.

Prefer CSS for presentation, a stable server-safe structure, or a client-only boundary when behavior truly depends on browser capability.


Dates, time zones, and random values

These are common mismatch sources:

server timezone ≠ browser timezone
server clock     ≠ browser clock
server random    ≠ browser random

Use a shared value, deterministic seed, explicit locale/timezone, or post-hydration update.


Streaming and failure isolation

A streamed region can fail after the shell is visible.

Define:

  • a local fallback;
  • retry behavior;
  • whether prior content remains;
  • logging and observability;
  • accessible announcement of the failure.

Progressive delivery requires progressive recovery.


SSG and content publishing

A publishing workflow should answer:

  • when is a page rebuilt?
  • how are previews generated?
  • how is stale content invalidated?
  • can a failed build leave the previous version live?
  • which content requires request-time freshness?

Rendering strategy includes the authoring and deployment system.


Rendering topology and deployment

CSR      → static shell + CDN + browser runtime
SSG      → build pipeline + static hosting
SSR      → request runtime + data services
edge SSR → distributed runtime + edge constraints

Deployment capabilities are part of the architecture, not an afterthought.


Rendering topology and failure modes

build failure      → previous deployment / publishing error
server data failure→ route or region fallback
client bundle fail → static content plus degraded interaction
hydration mismatch → deterministic handoff fix
stale output       → revalidation or freshness message

Design the failure path at the same time as the happy path.


Practical decision checklist

Ask:

  • Is the content public?
  • How often does it change?
  • Is it user-specific?
  • How interactive is it?
  • Does it need browser-only APIs?
  • Can the output be cached?
  • Are data dependencies slow?
  • Can interaction be isolated?
  • How much state must cross to the browser?
  • What happens after initial navigation?

Completion checklist

  • route strategy is chosen from requirements rather than slogans;
  • build, request, and browser work are identified;
  • hydration output is deterministic;
  • secrets and server-only data stay server-side;
  • interactive boundaries are no larger than necessary;
  • streaming boundaries have meaningful loading and error behavior;
  • cacheability and personalization are separated where possible;
  • freshness and revalidation are explicit;
  • JavaScript and device cost are measured;
  • deployment and failure behavior match the topology.

Misconceptions to leave behind (Part 1)

MisconceptionBetter mental model
CSR means no serverIt moves initial rendering work to the browser
SPA and CSR are identicalNavigation model and rendering location differ
SSR means no JavaScriptInteractive regions still need browser code
SSR is always fasterServer latency, hydration, and cacheability matter
SSG cannot be interactiveStatic output can include client islands
Static means always freshPublication and revalidation define freshness

Misconceptions to leave behind (Part 2)

MisconceptionBetter mental model
Hydration rebuilds the DOMIt attaches client behavior to existing output
Streaming removes handoff costIt changes delivery timing, not all work
Server components are just SSRExecution location and HTML timing differ
Islands are automatically betterCoordination and boundary costs still exist
Edge is always fasterRuntime, data location, and cache behavior matter
Resumability is just faster hydrationIt changes the server-client handoff model

The chapter in one sentence

Choose a rendering topology per route and region by balancing freshness, personalization, interaction, cacheability, handoff, deployment, and device cost.


Next: Chapter 12

The next chapter will build on rendering architecture with:

  • testing strategy and quality boundaries;
  • unit, integration, and end-to-end tests;
  • browser behavior and accessibility verification;
  • performance and failure testing;
  • confidence in evolving front-end systems.

Questions

For your home page, catalogue, account page, and admin editor: where should the first useful HTML come from, and what is the smallest region that truly needs browser execution?

12 - Modern Build Systems, Development Tooling & Team Workflows

Chapter 12: understand module graphs, transformation, bundling, code splitting, environment boundaries, and reproducible team workflows.

Modern Build Systems, Development Tooling & Team Workflows

Understand the path from source to delivery

Chapter 12

Polla Fattah


Today’s goal

Make the front-end toolchain visible as an architecture rather than a collection of commands.

We will connect:

  • package management, lockfiles, scripts, and modules;
  • resolution, transformation, development servers, and HMR;
  • production bundling, tree shaking, minification, and source maps;
  • asset fingerprints, CSS, environment variables, and code splitting;
  • Vite, Rollup, esbuild, Rolldown, and Turbopack;
  • linting, formatting, type checking, testing, Git, and CI;
  • workspaces, monorepos, package boundaries, and dependency direction;
  • an inspectable, reproducible build workflow.

By the end of today you can

  • explain why development and production optimize differently;
  • read a module graph and identify dependency direction;
  • distinguish resolution, transformation, bundling, and serving;
  • understand what tree shaking can and cannot remove;
  • choose code-splitting boundaries based on route behavior;
  • keep build-time configuration separate from runtime secrets;
  • design fast local feedback and shared CI verification;
  • use workspaces without confusing them with architecture;
  • expose package APIs without leaking private files;
  • inspect emitted assets instead of guessing about performance.

The central principle

A build toolchain is a delivery architecture: it transforms source modules into environment-specific artifacts while preserving reproducibility, boundaries, and useful feedback.

The tool is not only a compiler.

It resolves dependencies, serves development code, creates production artifacts, and shapes how teams work.


The source-to-delivery pipeline

flowchart LR
    A["1. Source Files
(TS, TSX, CSS)"] --> B["2. Package Resolution
(node_modules, exports)"]
    B --> C["3. Module Graph
(Static dependency DAG)"]
    C --> D["4. Transformation
(TypeScript/JSX stripping)"]
    D --> E["5. Bundler / Tree Shaking
(Chunk splitting & minification)"]
    E --> F["6. Production Artifacts
(Hashed JS, CSS, Source Maps)"]

Each stage answers a different question and has different failure modes.


Development and production optimize differently

flowchart TD
    subgraph DevelopmentMode["Development Environment (Inner Loop)"]
        D1["Unbundled Native ESM
(Instant server boot)"]
        D2["Hot Module Replacement (HMR)
(<50ms stateful updates)"]
        D3["Detailed Source Maps & Error Overlays"]
    end
    subgraph ProductionMode["Production Environment (Delivery Artifacts)"]
        P1["Aggressive Dead-Code Elimination (Tree Shaking)"]
        P2["Route-Level Code Splitting & Chunking"]
        P3["Content-Hashed Filenames for Immutable CDN Caching"]
        P4["Byte Minification & Production Tree Stripping"]
    end

A development server is not automatically a production server.

The two modes can share configuration while serving different goals.


Package management is the beginning of the toolchain

Packages define:

  • source dependencies;
  • executable scripts;
  • development tools;
  • transitive dependency graphs;
  • version constraints;
  • reproducible installation inputs.

The package manifest is an architectural document, not only an install list.


Runtime and development dependencies

runtime dependency   → needed by the shipped application
development dependency → needed to build, test, lint, or develop

Misclassifying a dependency can:

  • inflate production output;
  • break a library consumer;
  • make a production build depend on local tooling;
  • hide a missing runtime requirement.

Semantic version ranges are constraints

{
  "dependencies": {
    "library": "^3.2.0"
  }
}

The range describes acceptable versions.

The lockfile records the concrete resolution used by an installation.

Do not confuse the manifest promise with the installed graph.


Lockfiles belong in application repositories

A lockfile records:

  • concrete versions;
  • resolved locations;
  • integrity information;
  • transitive dependency choices.

It lets local development, CI, and deployment start from the same dependency graph.

Update it intentionally and use one package manager consistently.


Package scripts create a common interface

{
  "scripts": {
    "dev": "vite",
    "build": "tsc --noEmit && vite build",
    "lint": "eslint .",
    "test": "vitest run"
  }
}

Scripts hide machine-specific command details and give the team a shared workflow vocabulary.


ES modules are the structural foundation

import { formatPrice } from "./price";
export function renderProduct(product: Product) {
  return formatPrice(product.priceCents);
}

Static module syntax gives tools information about dependencies, exports, and boundaries.


Static imports are analyzable

import { loadProducts } from "./api";

The tool can use static imports to:

  • build a dependency graph;
  • detect missing modules;
  • identify unused exports;
  • split and optimize code;
  • pre-bundle dependencies.

Dynamic imports create asynchronous boundaries

const reports = await import("./reports");

Dynamic imports can create:

  • lazy routes;
  • optional features;
  • smaller initial chunks;
  • delayed dependency loading.

They also add loading, error, caching, and prefetch decisions.


CommonJS still exists

const packageValue = require("package");
module.exports = packageValue;

Interoperability can affect:

  • static analysis;
  • default exports;
  • tree shaking;
  • runtime loading;
  • package conditions.

Know the module format at the boundary you are integrating.


Module resolution answers “which file?”

Resolution may consider:

  • relative paths;
  • package exports;
  • file extensions;
  • aliases;
  • conditions for browser, node, import, or require;
  • workspace links;
  • type declarations.

The module graph begins only after resolution succeeds.


Aliases can improve architecture - or hide it

import { Button } from "@shared/ui/Button";

Aliases can make stable package boundaries readable.

They can also conceal deep dependencies and make imports appear more independent than they are.

Use aliases to communicate architecture, not to avoid designing it.


Transformation is not one operation

flowchart LR
    Src["Source Code
(TypeScript / JSX)"] --> Parse["Parser
(Generate AST)"]
    Parse --> AST["Abstract Syntax Tree
(AST Data Structure)"]
    AST --> Transform["Transforms
(Strip types, lower syntax)"]
    Transform --> Emit["Code Generator
(Standard ES2022 JavaScript)"]

Transformation may include:

  • TypeScript syntax removal;
  • JSX transformation;
  • language lowering;
  • macro or plugin transforms;
  • CSS processing;
  • asset URL rewriting.

TypeScript compilation has separate jobs

type checking    → verify static relationships
transformation   → produce executable JavaScript

A fast transpiler can remove types without proving the program correct.

Run type checking as its own explicit quality step when the build tool does not perform it.


Transpilation and polyfills are different

Transpilation can rewrite syntax:

new syntax → older syntax

It does not automatically provide missing runtime APIs such as a browser feature.

Polyfills, browser targets, and runtime support are separate decisions.


Development servers optimize feedback

A development server may provide:

  • native module serving;
  • on-demand transformation;
  • dependency pre-bundling;
  • source maps;
  • error overlays;
  • HMR or Fast Refresh.

It should make the smallest useful update quickly and explain failures clearly.


Hot Module Replacement

sequenceDiagram
    autonumber
    actor Dev as Developer
    participant FS as File System Watcher
    participant Server as Dev Server (Vite)
    participant Browser as Browser Client

    Dev->>FS: Saves Button.tsx
    FS->>Server: File change detected
    Server->>Server: Re-transform single module (15ms)
    Server-->>Browser: WebSocket: { type: 'update', path: '/src/Button.tsx' }
    Browser->>Server: HTTP fetch(/src/Button.tsx?t=171000)
    Server-->>Browser: Fresh module code
    Note over Browser: HMR Runtime replaces module without full page reload!
Component state preserved.

HMR shortens feedback loops, but it is not identical to a fresh application start.


Fast Refresh is framework-aware HMR

Framework-aware refresh can preserve component state when a change is safe.

It may reset state when:

  • module exports change shape;
  • boundaries are not refresh-safe;
  • initialization semantics change;
  • the framework cannot preserve identity.

Test both preserved and reset behavior when it matters.


HMR correctness matters

Hot updates can expose bugs hidden by full reloads:

  • duplicate subscriptions;
  • stale module state;
  • missing cleanup;
  • side effects at import time;
  • global registration repeated on every update.

Write initialization and cleanup so reload behavior remains predictable.


Vite as a reference development tool

Vite’s development model emphasizes:

  • fast startup;
  • native ESM-like module requests;
  • on-demand transforms;
  • dependency pre-bundling;
  • a separate production build pipeline.

The exact internals can evolve; the development versus production distinction remains useful.


Dependency pre-bundling

Third-party packages may contain many modules or formats that are expensive to request individually.

Pre-bundling can:

  • reduce browser request count;
  • normalize dependency formats;
  • improve development startup after caching.

It does not mean the production bundle and development graph are identical.


Native ESM is an important mental model

flowchart TD
    subgraph NativeESM["Unbundled Dev Server (Vite / Dev)"]
        B["Browser requests /src/main.ts"] --> S["Dev Server compiles on demand"]
        S --> M1["main.ts imports App.ts"]
        M1 --> M2["App.ts imports Button.ts"]
        M2 --> M3["Browser requests individual modules via native HTTP/2"]
    end
    subgraph ProductionBundling["Bundled Production Output (Rollup / esbuild)"]
        G["Module Graph Crawler"] --> T["Tree Shaking & DCE"]
        T --> C1["Chunk 1: app.8f2a.js (120KB)"]
        T --> C2["Chunk 2: vendor.3c1d.js (80KB)"]
    end

Development tools can preserve a module-oriented model while adding transforms, caching, and dependency optimization.

Do not assume that a development request corresponds one-to-one with a production asset.


Production bundling is graph transformation

many source modules
  → chunks with shared dependencies
  → optimized assets

Bundling includes:

  • module combination;
  • dependency ordering;
  • tree shaking;
  • code splitting;
  • asset rewriting;
  • minification;
  • source-map generation.

It is much more than concatenation.


Tree shaking removes unreachable exports

export function used() {}
export function unused() {}

If the graph and package semantics allow it, the unused export may be removed from a production chunk.

The tool needs analyzable module structure and correct side-effect information.


Tree shaking depends on analyzable code

Dynamic behavior can limit removal:

  • unpredictable property access;
  • CommonJS patterns;
  • runtime module discovery;
  • side effects hidden in imports;
  • package metadata that is too broad.

The tool can only prove what the code structure makes visible.


Side effects matter

import "./register-global";

An import may be valuable even when it has no exported binding.

Do not mark a package as side-effect-free unless removing such imports is safe.


Minification changes representation

Minification can:

  • shorten identifiers;
  • remove whitespace;
  • simplify expressions;
  • eliminate unreachable code;
  • compress repeated structures.

It improves transfer and sometimes execution, but it makes debugging harder without source maps.


Source maps connect artifacts to source

flowchart LR
    Err["Runtime Error in Browser
(vendor.8f31c.js:1:4210)"] --> Map["Source Map File
(vendor.8f31c.js.map
VLQ Mappings)"]
    Map --> Src["Original Source Code in DevTools
(src/services/permitApi.ts:42:15)"]

Source maps are an operational policy:

  • public or private?
  • uploaded to an error service?
  • exposed in production?
  • retained for which release?

Asset fingerprinting enables safe caching

app.js → app.8f31c.js

Content-based names let immutable assets use long cache lifetimes.

HTML or manifests must point to the current fingerprinted files.


Code splitting creates delivery choices

flowchart TD
    Entry["Main Entrypoint (main.ts)"] --> AppChunk["Initial Core Bundle
(Header, Nav, Router, Theme)
[app.8f31c.js - 45KB]"]
    
    AppChunk -.->|Static Import| Shared["Shared Vendor Chunk
(React, Query Cache)
[vendor.2e1a.js - 75KB]"]
    AppChunk -->|Dynamic import('./Reports')| RouteA["Async Route: Reports & Charts
[reports.6d4b.js - 180KB]"]
    AppChunk -->|Dynamic import('./Admin')| RouteB["Async Route: Admin Console
[admin.9c2e.js - 95KB]"]

Split points affect:

  • initial transfer;
  • later navigation latency;
  • cache reuse;
  • request count;
  • failure and loading behavior.

Route-level splitting is often a strong default

const ReportsPage = lazy(() => import("./reports/ReportsPage"));

Routes usually represent meaningful user journeys and can isolate code that is not needed initially.

They also provide a natural loading and error boundary.


More chunks are not always better

Tiny chunks can create:

  • request overhead;
  • scheduling complexity;
  • waterfall risk;
  • poor cache reuse;
  • difficult debugging.

Split around behavior and navigation, not every file.


Shared chunks are a trade-off

Shared dependencies can be downloaded once and reused.

But a large shared chunk can become part of every route’s initial cost even when only one feature uses it.

Analyze the actual graph and route traffic.


Preloading and prefetching split code

preload  → needed soon for the current route
prefetch → likely needed later when resources allow

Use hints based on confident navigation and device/network conditions.

Prefetching code the user never needs is still work.


CSS participates in the build pipeline

The pipeline may:

  • process imports;
  • scope or extract styles;
  • rewrite asset URLs;
  • split CSS by entry or route;
  • minify output;
  • preserve source maps.

CSS loading order and extraction affect rendering and layout stability.


Asset imports are module dependencies

import logoUrl from "./logo.svg";
import styles from "./Card.module.css";

The tool can fingerprint assets, generate URLs, and include only resources reachable from the graph.

The runtime receives a reference appropriate to the target environment.


Environment variables have a boundary

flowchart TD
    subgraph BuildTime["Build-Time Replacement (Vite / Bundler)"]
        E1["import.meta.env.VITE_API_URL"] --> R1["Replaced during build with string literal
'https://api.erbil.gov.krd'"]
        Note1["WARNING: Embedded into public client bundle!
Never place database passwords here."]
    end
    subgraph RuntimeConfig["Runtime Environment Configuration"]
        E2["process.env.DATABASE_PASSWORD"] --> R2["Read dynamically on server execution"]
        Note2["Safe: Stays inside secure container / server."]
    end

Browser-exposed variables are public.

.env naming does not make a value secret once it is included in client output.


Build-time and runtime configuration differ

Build-time values require rebuilding to change.

Runtime values can vary per deployment or request without producing new assets.

Choose based on:

  • environment portability;
  • caching;
  • deployment frequency;
  • secrecy;
  • per-request variation.

Vite production builds

A production build typically performs:

  • module resolution;
  • transformation;
  • dependency graph optimization;
  • code splitting;
  • asset fingerprinting;
  • minification;
  • source-map output according to policy.

Inspect the generated result instead of inferring it from development requests.


Rollup

Rollup is a graph-oriented bundler known for:

  • ES module analysis;
  • library and application bundling;
  • tree shaking;
  • plugin-based transformation;
  • controllable output formats.

Use it directly when its lower-level control matches the responsibility you need.


esbuild

esbuild emphasizes very fast native transformation and bundling.

It can be useful for:

  • fast development transforms;
  • custom build integration;
  • straightforward application builds;
  • tooling that needs quick parsing and emission.

Fast transformation does not remove the need for application architecture or type checking.


Rolldown and Turbopack

Modern tools evolve their internals to improve:

  • incremental builds;
  • graph computation;
  • parallelism;
  • native performance;
  • framework integration.

Learn the responsibility each tool solves instead of treating its implementation name as the architecture.


development graph and feedback
production graph and output
framework-specific incremental integration

Tools can differ in when they bundle, how they cache, and how they integrate with a framework.

The stable mental model is the source graph and emitted responsibility.


The framework often chooses the tool

Framework defaults can determine:

  • development server;
  • route build model;
  • server/client boundaries;
  • asset handling;
  • production output;
  • plugin lifecycle.

Developers still need to understand the underlying responsibilities when diagnosing output or performance.


Plugins extend the pipeline

Plugins may add:

  • syntax transforms;
  • virtual modules;
  • asset handling;
  • route generation;
  • environment integration;
  • development middleware.

Every plugin adds behavior to resolution, transformation, or output. Keep plugin purpose and order understandable.


Avoid toolchain configuration as a hobby

Configuration should answer a delivery or developer-experience need.

Before adding a plugin or custom transform, ask:

  • which problem does it solve?
  • who owns it?
  • what does it change in production?
  • how is it tested?
  • what is the fallback if it breaks?

Complexity without a user or team benefit is a liability.


Linting catches a class of problems

Linting can detect:

  • suspicious patterns;
  • unused bindings;
  • import restrictions;
  • unsafe APIs;
  • inconsistent architectural rules.

It is not formatting, type checking, or a substitute for tests.


Formatting reduces diff noise

A formatter creates a shared representation so reviews focus on behavior and design.

Format consistently and avoid repeatedly reformatting unrelated files in feature commits.

Formatting should support review, not dominate it.


Type checking verifies static contracts

flowchart TD
    subgraph QualityPillars["The Five Quality Gates of Front-End Delivery"]
        Q1["1. Linter (ESLint / Biome)
Flags bug-prone anti-patterns & security flaws"]
        Q2["2. Formatter (Prettier / Biome)
Guarantees deterministic code style across team"]
        Q3["3. Type Checker (tsc --noEmit)
Verifies compile-time type safety & API contracts"]
        Q4["4. Test Runner (Vitest / Playwright)
Verifies runtime behavioral correctness"]
        Q5["5. Production Bundler (Vite build)
Verifies module graph resolution & asset generation"]
    end

These gates overlap in value but do not answer the same question.


Local feedback should be fast

flowchart LR
    IDE["1. IDE / Editor
(<50ms inline feedback)"] --> GitHook["2. Pre-Commit Hook
(lint-staged on changed files)"]
    GitHook --> PR["3. Git Pull Request"]
    PR --> CI["4. CI Automation Pipeline
(Full clean install, typecheck, test, build)"]
    CI --> Prod["5. Verified Deployment Artifact"]

Fast feedback catches cheap mistakes close to the change.

CI provides a shared clean-environment check.


Git is part of the engineering toolchain

Git supports:

  • reviewable change boundaries;
  • reproducible history;
  • rollback;
  • release traceability;
  • collaboration across branches.

Treat commit and merge practices as part of delivery quality.


Keep commits reviewable

Prefer commits that separate:

  • mechanical formatting;
  • dependency updates;
  • tool configuration;
  • feature behavior;
  • generated artifacts.

Small coherent changes make toolchain failures easier to bisect and understand.


Generated files need a policy

Decide which artifacts are:

  • committed;
  • generated in CI;
  • published separately;
  • reproducible from source;
  • ignored locally.

Inconsistent policies create noisy diffs and uncertain releases.


Environment reproducibility

Reproducibility depends on:

  • lockfile;
  • package-manager version;
  • Node/runtime version;
  • build configuration;
  • environment inputs;
  • operating-system assumptions;
  • clean install behavior.

Document and automate the parts that affect artifacts.


CI is the shared verification environment

CI should run from a clean state and verify the contracts that matter:

install → lint → typecheck → test → build → artifact checks

The exact order can vary, but the environment should not depend on one developer’s machine.


Build once, deploy predictably

source + locked dependencies + config
  → one verified artifact
  → promote the artifact across environments

Rebuilding separately for each environment can produce different output and weaken release confidence.

Keep truly runtime-varying configuration outside immutable assets where possible.


Workspaces solve package coordination

Workspaces can:

  • install multiple packages together;
  • link internal packages;
  • share scripts and dependency policy;
  • coordinate builds;
  • make local package changes visible quickly.

They are a package-management feature, not automatically a monorepo architecture.


A workspace is not automatically a monorepo architecture

workspace → repository/package coordination mechanism
monorepo  → repository organized around multiple related projects/packages

A workspace can host one application plus tools.

A monorepo still needs package boundaries, ownership, and dependency direction.


Monorepo benefits

Potential benefits include:

  • shared code with local changes;
  • coordinated releases;
  • consistent tooling;
  • cross-package refactors;
  • shared CI infrastructure.

The benefits appear only when boundaries and workflows remain understandable.


Monorepo costs

Costs include:

  • larger dependency graph;
  • longer or more complex CI;
  • ownership ambiguity;
  • cross-package build ordering;
  • accidental coupling;
  • difficult versioning decisions.

Do not choose a monorepo solely because it is fashionable.


Package boundaries should represent architecture

packages/ui        reusable visual behavior
packages/domain    shared domain contracts
apps/storefront    product application
apps/admin         internal application

The directory is useful when it reflects a responsibility and public API.


Internal packages need public APIs too

import { Button } from "@workspace/ui";

Prefer an intentional package entry point over:

import { Button } from "@workspace/ui/src/components/Button";

Private file imports bypass encapsulation and make internal restructuring expensive.


Dependency direction matters

flowchart TD
    App["apps/citizen-portal (Applications)"] --> Feat["packages/feature-permits (Feature Libraries)"]
    Feat --> Domain["packages/domain-licensing (Domain Models & Rules)"]
    Domain --> Core["packages/ui-components (Design System & Primitives)"]
    Core --> Util["packages/utilities (Pure helpers & math)"]

Lower-level packages should not import higher-level application decisions.

Direction makes ownership and reuse possible.


Circular dependencies are design feedback

A → B → C → A

Even if the build succeeds, cycles can create:

  • partial initialization;
  • undefined exports;
  • confusing evaluation order;
  • difficult testing;
  • architecture that cannot be layered cleanly.

Break the cycle by clarifying ownership or extracting a stable lower-level contract.


Analyze the dependency graph

Look for:

  • unexpected large imports;
  • application code imported by shared packages;
  • cycles;
  • duplicate versions;
  • route chunks containing unrelated features;
  • a dependency that dominates initial transfer.

The graph turns performance and architecture discussions into evidence.


Development and production differ

development → source modules, overlays, HMR, readable maps
production   → optimized chunks, fingerprints, minification, deployment paths

Always run a production build before making claims about production artifacts.


dev is not a production server

Development servers may:

  • transform on demand;
  • tolerate missing optimizations;
  • expose source files;
  • use different caching;
  • allow permissive CORS or proxies;
  • rely on local filesystem behavior.

Use a production build and production-like serving environment for delivery verification.


Build targets define browser assumptions

Targets influence:

  • syntax transformation;
  • polyfill decisions;
  • output size;
  • supported APIs;
  • debugging expectations.

Choose a browser policy deliberately and keep it visible to the team.


Baseline and browser policy

A supported-browser policy should answer:

  • which browsers and versions are supported;
  • which features are assumed;
  • when a feature needs a fallback;
  • how support changes are reviewed;
  • how real devices are tested.

The build cannot invent a product support policy.


Library builds versus application builds

application → optimize one deployed experience
library     → preserve a reusable public API and externalize peers

Library output needs stable formats, declarations, exports, and consumer compatibility.

Application output can make more assumptions about its deployment.


Dependency externalization

Libraries may leave dependencies external so consumers provide them.

This avoids duplicating framework code but creates peer-version and runtime compatibility obligations.

Decide which code belongs in the library artifact and which belongs to the consuming application.


Tree shaking and package design

Packages are easier to optimize when they:

  • use analyzable ES modules;
  • expose focused entry points;
  • avoid import-time side effects;
  • declare side-effect behavior accurately;
  • avoid pulling a whole framework for one helper.

Package API design directly affects emitted application code.


Development proxying

A development proxy can route browser requests to an API and reduce local cross-origin friction.

It should not hide production differences in:

  • origin;
  • cookies;
  • headers;
  • HTTPS;
  • path rewriting;
  • error behavior.

Document what the proxy changes.


HTTPS in development

HTTPS may be required to reproduce:

  • secure cookies;
  • service workers;
  • browser APIs requiring a secure context;
  • mixed-content behavior;
  • realistic origin policy.

Use a consistent local certificate workflow when the application depends on these capabilities.


Environment parity

Compare development and production for:

  • origin and base path;
  • asset URLs;
  • API proxying;
  • environment variables;
  • compression;
  • source maps;
  • service-worker behavior;
  • caching headers.

Parity is a debugging tool, not a demand that every environment be identical.


Team tooling should be boring

Good tooling is:

  • documented;
  • reproducible;
  • fast enough to use;
  • hard to misconfigure;
  • easy to upgrade deliberately;
  • understandable by the whole team.

The goal is dependable delivery, not admiration for configuration cleverness.


Avoid global tool dependencies

Prefer project-local versions for:

  • bundlers;
  • formatters;
  • linters;
  • test runners;
  • type tooling;
  • code generators.

Global tools can silently differ from CI and another developer’s environment.


Editor integration is developer experience

Editor support can provide:

  • type feedback;
  • import navigation;
  • formatting;
  • lint diagnostics;
  • test discovery;
  • refactoring support.

The editor should use the repository’s configuration rather than a parallel personal toolchain.


Pre-commit hooks are a feedback boundary

Use hooks for fast checks that should block obviously broken commits.

Keep them bounded.

Long full builds in every commit encourage bypassing the hook; put broad verification in CI.


Dependency updates need policy

For updates, consider:

  • security advisories;
  • lockfile changes;
  • peer compatibility;
  • bundle impact;
  • behavior changes;
  • migration notes;
  • rollback path.

Toolchain dependencies can change emitted artifacts even when application code is unchanged.


Toolchain dependencies have supply-chain risk

Build tools execute code in a privileged development and CI context.

Reduce risk with:

  • lockfiles and review;
  • trusted registries;
  • minimal dependencies;
  • update monitoring;
  • restricted CI credentials;
  • artifact inspection.

The toolchain is part of the application’s attack surface.


A reference workflow

flowchart TD
    A["1. Developer edits component in IDE"] --> B["2. Instant HMR verification in browser"]
    B --> C["3. Local targeted lint and test execution"]
    C --> D["4. Git commit & push to pull request"]
    D --> E["5. CI executes clean install (npm ci with locked dependencies)"]
    E --> F["6. Automated verification: lint + tsc + test suite"]
    F --> G["7. Production build & asset budget check"]
    G --> H["8. Immutable content-hashed artifacts deployed to CDN"]

Each stage should add confidence without repeating every earlier cost.


Practical project structure

app/
  src/
  public/
  package.json
  tsconfig.json
  vite.config.ts
  eslint.config.js
  lockfile

The exact files vary.

The important point is that source, configuration, scripts, and lockfile form one reproducible project boundary.


Build pipeline for the storefront

flowchart LR
    S1["1. Resolve Imports"] --> S2["2. Transform TS / JSX"]
    S2 --> S3["3. Route Code Splitting"]
    S3 --> S4["4. Tree Shaking & DCE"]
    S4 --> S5["5. Minification"]
    S5 --> S6["6. Asset Fingerprinting
(app.8f31c.js)"]
    S6 --> S7["7. Emit Manifest & Maps"]

Trace a source module through this pipeline to understand what the browser receives.


Lazy loading reports

const ReportsRoute = () => import("./reports/ReportsRoute");

The reports route should not appear in the initial chunk when the build and route architecture permit splitting.

Verify with emitted assets and a network trace rather than trusting the source syntax alone.


Inspect the build instead of guessing

Useful evidence includes:

  • asset sizes;
  • chunk composition;
  • duplicate dependencies;
  • source-map module lists;
  • route request waterfalls;
  • cache headers;
  • compressed transfer sizes.

Build analysis is most useful when connected to a user journey.


Toolchain smells

Watch for:

  • build configuration nobody understands;
  • every project using different lint rules;
  • lockfiles regenerated by multiple managers;
  • production builds rarely run;
  • huge initial bundles despite route structure;
  • hundreds of tiny chunks;
  • secrets in front-end environment variables;
  • internal package imports through private paths;
  • circular dependencies everywhere.

Choose tools by responsibility

application dev server + build → high-level framework tool
specialized library output   → library bundler
fast transformation          → native transformer
multiple local packages      → workspace / monorepo tooling
framework-specific pipeline  → framework-integrated tool

Choose the smallest toolset that serves the responsibility.


Current tooling changes; architecture remains

Tool internals evolve.

The stable questions remain:

  • what is the module graph?
  • where are transformation boundaries?
  • what ships initially?
  • what loads later?
  • what is public configuration?
  • how is output verified?
  • who owns each package boundary?

Practical lab: Inspect a Modern Front-End Toolchain

Use a modern build tool to observe resolution, transformation, module graphs, code splitting, and deployment artifacts.

The practical compares source modules with development requests and production output.


Practical stages 1–4: create and inspect

  1. Create a small TypeScript application.
  2. Add a dynamically imported reports route.
  3. Inspect development requests and production chunks.
  4. Compare source modules with emitted assets and source maps.

Record which decisions belong to the toolchain and which belong to application architecture.


Practical stages 5–9: build and split

  1. Separate type checking.
  2. Create a production build.
  3. Explore tree shaking.
  4. Explore minification.
  5. Add route-level code splitting.

Verify that reports code is not in the initial route chunk when the build permits splitting.


Practical stages 10–14: optimize and verify

  1. Prefetch a lazy route.
  2. Analyze a large dependency.
  3. Add linting and formatting.
  4. Add CI-style verification.
  5. Compare build-time and runtime configuration.

Interpret bundle size alongside route behavior and field performance.


Practical stages 15–18: boundaries and architecture

  1. Explore package encapsulation.
  2. Compare Vite and low-level bundler responsibilities.
  3. Research Turbopack through a Next.js example.
  4. Create a build architecture diagram.

Do not present a dependency-graph visualizer as a production bundler.


Practical extension: educational graph visualizer

Build a deliberately limited visualizer showing:

entry → imports → dynamic split point → emitted chunk

Label it educational.

The goal is to make resolution and dependency direction visible, not to recreate a production bundler.


Try this yourself

Pick one initial route and answer:

  • which source modules does it need?
  • which dependency dominates its size?
  • which feature can split at a route boundary?
  • what can be tree-shaken?
  • what must remain a side effect?
  • which environment values are public?

Then verify every answer against the emitted build.


Troubleshooting guide (Part 1)

SymptomLikely cause
Dev works, production failsDifferent transforms, paths, or environment assumptions
Reports code is in the initial chunkImport is static or split boundary is ineffective
Unused package code remainsSide effects, module format, or graph opacity
CI differs from localLockfile, runtime, or global tool mismatch
Secret appears in client outputPublic build-time variable was treated as private

Troubleshooting guide (Part 2)

SymptomLikely cause
Package consumers import private filesPublic API is incomplete or undocumented
HMR behaves strangelyImport-time side effects or cleanup missing
Tiny chunks hurt navigationSplit points follow files rather than user journeys
Build graph contains cyclesDependency direction is unclear

Completion checklist

  • package versions and lockfile are reproducible;
  • resolution and module boundaries are understandable;
  • type checking is an explicit quality step;
  • development and production workflows are distinguished;
  • code splitting follows route or feature behavior;
  • tree-shaking assumptions are validated by output;
  • source maps and environment values have policy;
  • package APIs protect internal files;
  • CI verifies a clean production build;
  • emitted artifacts are inspected rather than guessed at.

Misconceptions to leave behind (Part 1)

MisconceptionBetter mental model
A build tool is just a compilerIt resolves, transforms, serves, bundles, and emits
Bundling means concatenationIt is graph transformation and asset design
TypeScript must emit JavaScript through tscType checking and transformation can be separate
Transpilation adds missing browser APIsPolyfills and runtime support are separate
More code splitting is always betterSplit around user journeys and costs
Tree shaking removes anything not calledIt depends on analyzable graphs and side effects

Misconceptions to leave behind (Part 2)

MisconceptionBetter mental model
Minification makes source maps unnecessaryDebugging still needs source policy
.env values are secretClient-exposed values are public
A workspace is a monorepo architectureCoordination tooling and boundaries differ
A successful circular build is healthyCycles are dependency-design feedback
The dev server is productionProduction artifacts and serving behavior differ
Quality gates are interchangeableLint, format, types, tests, and build answer different questions

The chapter in one sentence

Treat the toolchain as an inspectable delivery graph that transforms source into reproducible artifacts while preserving package, environment, and team boundaries.


Next: Chapter 13

The next chapter will build on tooling and architecture with:

  • testing strategy and confidence boundaries;
  • unit, integration, and end-to-end verification;
  • browser behavior and accessibility tests;
  • performance and failure testing;
  • quality workflows for evolving applications.

Questions

Can you trace one user-visible route from its source entry point to the emitted assets, and explain which toolchain decisions affect its first-load cost?

13 - Front-End Security, Authentication & Browser Isolation

Chapter 13: reason about origins, untrusted input, authentication, authorization, browser defenses, and layered security boundaries.

Front-End Security, Authentication & Browser Isolation

Design the boundary before the attack

Chapter 13

Polla Fattah


Today’s goal

Treat the browser as a security runtime with explicit trust boundaries.

We will connect:

  • origins, same-origin policy, and CORS;
  • XSS, encoding, sanitization, CSP, and Trusted Types;
  • CSRF, cookies, sessions, and logout;
  • authentication, authorization, bearer tokens, OAuth, and PKCE;
  • BFF architecture and token storage trade-offs;
  • secrets, SRI, dependencies, and third-party scripts;
  • clickjacking, framing, COOP, COEP, and CORP;
  • secure messaging, iframes, logging, and a review checklist.

By the end of today you can

  • identify the origin and trust boundary of a browser request;
  • explain what CORS does and does not protect;
  • trace untrusted input from source to dangerous sink;
  • prefer safe text rendering and reviewed sanitization;
  • use CSP as defense in depth;
  • distinguish XSS from CSRF and authentication from authorization;
  • choose cookie, BFF, or browser-token architecture deliberately;
  • explain OAuth authorization code flow with PKCE;
  • keep secrets out of browser bundles and URLs;
  • review third-party, framing, messaging, and isolation risks.

The central principle

Front-end security comes from layered trust boundaries, not from one framework feature, header, token format, or browser flag.

Ask where data comes from, who can read it, who can change it, and which layer must enforce the decision.


The security review progression

flowchart TD
    A["1. Identify Origins & Trust Boundaries"] --> B["2. Trace Untrusted Input (Sources)"]
    B --> C["3. Secure Rendering Sinks (XSS Defense)"]
    C --> D["4. Protect State-Changing Requests (CSRF)"]
    D --> E["5. Separate Authentication from Authorization"]
    E --> F["6. Secure Credentials & Token Storage (BFF)"]
    F --> G["7. Enforce Browser Isolation (CSP, COOP, COEP)"]
    G --> H["8. Audit Dependencies & Third-Party Scripts"]

Security is a system of related controls, not a checklist of isolated switches.


The browser is a multi-origin runtime

page origin
API origin
identity-provider origin
CDN origin
embedded iframe origin
third-party script origin

The browser applies different rules to interactions among these origins.

Your architecture must make those relationships intentional.


What is an origin?

An origin is the combination of:

flowchart LR
    subgraph OriginTuple["The Web Origin Definition (RFC 6454)"]
        S["Scheme (e.g. https://)"] --- H["Host (e.g. app.erbil.gov.krd)"] --- P["Port (e.g. :443)"]
    end

Examples:

https://app.example.com
https://api.example.com
http://app.example.com
https://app.example.com:8443

Small differences can produce different origins.


Same-origin examples

https://app.example.com/page
https://app.example.com/api

These share an origin when scheme, host, and port match.

Paths differ, but origin policy is not path policy.


Origin is not the same as site

Cookies and browser policies sometimes reason about a broader “site” concept based on registrable domains.

Do not use “same site” and “same origin” interchangeably.

They affect different browser mechanisms.


Same-origin policy

The same-origin policy restricts how a document can read or interact with resources from another origin.

It is a browser protection boundary.

It does not mean cross-origin activity is impossible.


Same-origin policy is not “no cross-origin activity”

Browsers may allow controlled cross-origin actions such as:

  • loading images;
  • submitting forms;
  • embedding frames;
  • sending requests;
  • loading scripts under policy rules.

The key question is often whether the initiating page can read the response or control the embedded context.


Cross-origin reads are the core concern

page → cross-origin request → response
                         ↘ browser may block script access

CORS controls whether browser JavaScript may read a cross-origin response.

It is not a universal firewall around the server.


Browser storage is origin-scoped

https://app.example.com storage
≠
https://other.example.com storage

Storage isolation helps protect data between origins.

It does not protect an origin from XSS that executes inside that origin.


CORS is a server permission

Access-Control-Allow-Origin: https://app.example.com

The server tells the browser which origin may read a response.

The browser enforces the permission for script access.


CORS is not authentication

CORS does not prove who the user is.

It does not issue a session, validate an access token, or decide whether a user may update a record.

Authentication and authorization remain server responsibilities.


CORS does not protect the server from non-browser clients

browser CORS policy ≠ server access control

Command-line clients, native apps, scripts, and attackers can send requests without browser CORS enforcement.

The server must authenticate, authorize, validate, and rate-limit independently.


Simple and preflighted CORS requests

Some cross-origin requests can proceed with a browser permission check.

Others trigger an OPTIONS preflight describing:

  • intended method;
  • requested headers;
  • requested origin.

The server must explicitly allow the operation before the browser sends the actual request.


Preflight is not a failure

sequenceDiagram
    autonumber
    actor Browser as Browser Client (https://app.erbil.gov.krd)
    participant API as Municipal API (https://api.erbil.gov.krd)

    Browser->>API: OPTIONS /permits/104 (Preflight)
Origin: https://app.erbil.gov.krd
Access-Control-Request-Method: PUT
Access-Control-Request-Headers: Content-Type
    Note over API: API verifies origin in allowed whitelist
    API-->>Browser: 204 No Content
Access-Control-Allow-Origin: https://app.erbil.gov.krd
Access-Control-Allow-Methods: GET, PUT, POST
Access-Control-Allow-Headers: Content-Type
    Browser->>API: PUT /permits/104 (Actual Mutation Request)
    API-->>Browser: 200 OK (Resource updated)

Preflight is a safety mechanism.

If it fails, inspect origin, method, headers, credentials, and server policy rather than disabling security blindly.


Credentials and CORS need alignment

Credentialed requests require coordinated policy:

  • explicit allowed origin;
  • allowed credentials;
  • cookie attributes;
  • server authentication;
  • CSRF protection where relevant.

Wildcard origins are not a substitute for a deliberate credential policy.


CORS errors often reveal architecture problems

A CORS failure may indicate:

  • API and app origins were not designed together;
  • development proxy hid production behavior;
  • credential policy is unclear;
  • a public/private boundary is ambiguous;
  • the browser is being asked to call a server that should be behind a BFF.

Fix the boundary, not only the console message.


Development proxies can hide CORS

development: browser → same-origin proxy → API
production:   browser → cross-origin API

Test production-like origins before shipping.

Otherwise the first real CORS behavior appears in deployment.


XSS is more than <script> tags

Cross-site scripting occurs when attacker-controlled data becomes executable or dangerous browser content.

Possible paths include:

  • HTML injection;
  • event-handler attributes;
  • dangerous URLs;
  • script-capable SVG;
  • template or expression injection;
  • DOM APIs that interpret strings as markup.

Safe rendering by default

element.textContent = userText;

Prefer APIs and framework bindings that treat values as text.

Escaping by default is useful, but still inspect escape hatches and URL/style contexts.


Dangerous escape hatches

Treat these as security-sensitive:

innerHTML
dangerouslySetInnerHTML
v-html
document.write
eval / Function
unreviewed HTML renderers

Every escape hatch needs a documented trust boundary and a reviewed input policy.


Encoding and sanitization are different

output encoding → display a value as text in one context
sanitization    → remove or constrain allowed markup and behavior

Encoding is context-specific.

Sanitization is a policy for permitting a restricted subset of content.

Neither should be applied blindly to every output context.


Prefer text over HTML

messageElement.textContent = message;

If the product requirement is text, do not create an HTML parsing problem.

Rich HTML should be an explicit feature with an explicit content policy.


DOM-based XSS

flowchart TD
    subgraph Sources["Untrusted Sources"]
        S1["location.search / hash"]
        S2["API JSON responses"]
        S3["localStorage / cookies"]
        S4["postMessage events"]
        S5["User form inputs"]
    end

    subgraph DangerousSinks["Dangerous DOM Sinks (Vulnerabilities)"]
        D1["element.innerHTML"]
        D2["dangerouslySetInnerHTML / v-html"]
        D3["eval() / new Function()"]
        D4["<a href='javascript:...'>"]
        D5["document.write()"]
    end

    Sources -->|Direct assignment without sanitization| DangerousSinks
    DangerousSinks --> XSS["Cross-Site Scripting (XSS)
Attacker script executes with full user privileges!"]

The server does not need to be involved.

Client code can create XSS by moving an untrusted value into a dangerous sink.


Sources and sinks

source: location.search, storage, API, form, postMessage
sink:   innerHTML, script URL, eval, unsafe style, HTML parser

Security review traces values from source to sink and asks what validation or encoding occurs between them.


textContent is safer for text

const status = document.querySelector("#status");
status?.textContent = externalValue;

It creates a text node rather than parsing markup.

Use the simplest API that matches the content requirement.


URL handling needs context

link.href = userProvidedUrl;

Validate:

  • allowed schemes;
  • allowed hosts when appropriate;
  • relative versus absolute behavior;
  • redirect policy;
  • display text separately from destination.

Text escaping alone does not make every URL safe.


Avoid eval()-style execution

Never turn untrusted strings into code through:

  • eval();
  • Function();
  • string-based timers;
  • dynamic script construction;
  • template expression interpreters without a trusted boundary.

If the product needs expressions, design a constrained language and parser rather than executing JavaScript.


Sanitizing rich HTML

If rich content is required:

  1. define the allowed elements and attributes;
  2. sanitize with a maintained, reviewed library;
  3. sanitize near the trust boundary;
  4. preserve the sanitized representation;
  5. render through one controlled component;
  6. test dangerous payloads and URL contexts.

Do not assume a generic “clean HTML” label communicates the policy.


Sanitization should happen near the trust boundary

external content → validate/sanitize → trusted rich-text model → renderer

Repeated ad hoc sanitization in every component creates inconsistent policies and missed sinks.


Content Security Policy is defense in depth

CSP can restrict what the browser may execute or load if an injection reaches the page.

It should support safe architecture, not justify unsafe rendering.

Content-Security-Policy: default-src 'self'

Start from the resources the application actually needs.


default-src establishes a baseline

More specific directives can control:

script-src
style-src
img-src
connect-src
font-src
frame-src
object-src

The policy should be reviewed with application dependencies and deployment origins.


Script policies matter most

Avoid broad script permissions where possible.

Prefer:

  • external scripts from known origins;
  • nonces or hashes for deliberate inline code;
  • removal of inline handlers;
  • restricted dynamic execution.

unsafe-inline and unsafe-eval should be treated as explicit trade-offs, not defaults.


Nonces authorize specific inline scripts

Content-Security-Policy: script-src 'nonce-random-per-response'

The server places the same unpredictable nonce on an intentionally permitted script.

Never reuse a predictable or long-lived nonce.


CSP report-only mode

Report-only mode helps discover violations before enforcement.

Use it to:

  • inventory real dependencies;
  • identify inline scripts;
  • find unexpected connections;
  • observe third-party behavior;
  • refine the policy.

Then move to enforcement intentionally.


CSP and third-party scripts

Third-party scripts expand the policy and trust surface.

For each script, ask:

  • what data can it read?
  • what can it send?
  • what happens if it changes?
  • can it be removed or isolated?
  • is its origin and integrity controlled?

Allowing a script is granting code execution in the page’s origin.


Trusted Types

Trusted Types can require dangerous DOM sinks to receive approved trusted values rather than arbitrary strings.

They can make unsafe paths harder to reach accidentally.

They are a design constraint and enforcement layer, not a substitute for understanding the content policy.


Trusted Types are not sanitization by themselves

A policy can create trusted HTML only after applying a reviewed sanitizer or construction rule.

string → reviewed policy → TrustedHTML → controlled sink

The policy is where the security decision lives.


CSRF and XSS are different

flowchart TD
    subgraph XSS_Threat["Cross-Site Scripting (XSS)"]
        X1["Attacker injects malicious script into trusted origin"]
        X2["Script runs with full DOM access: reads tokens, steals cookies, logs keystrokes"]
    end
    subgraph CSRF_Threat["Cross-Site Request Forgery (CSRF)"]
        C1["Attacker tricks authenticated browser into issuing request to target origin"]
        C2["Browser automatically attaches ambient credentials (cookies)"]
        C3["Attacker cannot read response, but executes unauthorized side-effects"]
    end

They can interact, but the defenses and threat paths differ.


CSRF tokens

For cookie-authenticated mutations, a server can require a token that an attacker site cannot read.

session cookie + request-specific CSRF proof

Validate the token on the server for state-changing operations.


CSRF applies to state-changing requests

Protect operations such as:

  • create;
  • update;
  • delete;
  • change email;
  • change password;
  • transfer or purchase.

Do not rely on a request method name alone; define which operations change state.


SameSite cookies reduce cross-site sending

flowchart TD
    subgraph SameSiteDirectives["SameSite Cookie Attribute Policies"]
        ST["SameSite=Strict
Cookie NEVER sent on cross-site requests
(Even clicking an external link to the portal)"]
        LX["SameSite=Lax (Modern Browser Default)
Cookie sent on top-level safe GET navigations
Blocked on cross-site POST / PUT / fetch mutations"]
        NN["SameSite=None; Secure
Cookie sent across all cross-site requests (Requires HTTPS)
High CSRF exposure without explicit tokens"]
    end

SameSite is valuable defense in depth, but consider legacy behavior, integrations, and the operation’s risk.


Origin and Referer checks

For sensitive mutations, the server may verify request metadata such as Origin or Referer according to a documented policy.

These checks complement, rather than replace, appropriate session and CSRF design.


Set-Cookie: session=…; HttpOnly; Secure; SameSite=Lax

Cookies can keep session credentials out of JavaScript.

They still require:

  • CSRF consideration;
  • expiration and rotation;
  • scope control;
  • logout and revocation behavior.

HttpOnly

HttpOnly prevents JavaScript from reading the cookie value.

It can reduce token theft through direct storage access.

It does not prevent XSS code from making requests as the user while the page is open.


Secure

Secure tells the browser to send the cookie only over HTTPS.

It protects transport confidentiality for the cookie, but does not solve application authorization or compromised client code.


Narrow scope where possible.

Broad domain cookies increase the set of subdomains and applications that participate in the session boundary.

Path is routing scope, not a complete security boundary for all cookie behavior.


Session fixation and rotation

Rotate session identity when privilege changes, such as after login.

This prevents an attacker from setting or learning a session identifier that remains valid after authentication.

Expire and revoke sessions according to risk and product requirements.


Logout is a security transition

Logout may need to:

  • revoke or expire the server session;
  • clear client state;
  • clear private caches;
  • stop live connections;
  • remove pending user-specific data;
  • prevent back-navigation from revealing sensitive content.

It is more than hiding a button.


Authentication versus authorization

authentication → who is this?
authorization   → what may this identity do here?

The front end can reflect permissions.

The server must enforce them.


Front-end authorization is not security enforcement

{canDelete && <DeleteButton />}

This improves user experience.

It does not protect the delete endpoint if an attacker sends the request directly.

Every sensitive server operation must enforce authorization independently.


Permissions should come from an authoritative model

Avoid deriving critical permissions solely from:

  • hidden UI controls;
  • route names;
  • client booleans;
  • decoded but unverified payloads;
  • stale local storage.

The server should return or enforce a permission model the client can safely present.


Bearer tokens

Authorization: Bearer access-token

Anyone who possesses a bearer token may use it within its scope and lifetime.

Protect issuance, storage, transport, scope, expiry, and revocation.


Browser token storage is a trade-off

flowchart TD
    subgraph LocalStorageOption["Option A: localStorage / Memory Bearer Token"]
        L1["Readable by JavaScript in same origin"]
        L2["Immune to CSRF (not ambiently sent)"]
        L3["CRITICAL RISK: A single XSS flaw exposes token to theft!"]
    end
    subgraph HttpOnlyCookieOption["Option B: HttpOnly, Secure, SameSite Cookie"]
        C1["Completely inaccessible to JavaScript (XSS cannot steal)"]
        C2["Ambiently sent by browser on matching requests"]
        C3["DEFENSE REQUIRED: Enforce SameSite=Lax + Anti-CSRF Token headers"]
    end

There is no universal slogan that replaces threat modeling and architecture.


Prefer architecture over token folklore

Choose based on:

  • XSS exposure;
  • CSRF exposure;
  • same-origin or cross-origin deployment;
  • backend control;
  • mobile/native clients;
  • refresh and revocation needs;
  • third-party integrations.

A BFF can change the browser’s credential boundary substantially.


Backend-for-Frontend topology

flowchart LR
    subgraph BrowserZone["Browser Runtime"]
        SPA["Single-Page App"]
    end
    subgraph InternalBoundary["Same-Origin Boundary"]
        BFF["Backend-for-Frontend (BFF)"]
    end
    subgraph SecureBackend["Internal Protected Network"]
        IDP["OAuth / OIDC IDP"]
        APIs["Microservices / APIs"]
    end

    SPA <-->|HttpOnly, Secure Cookie| BFF
    BFF <-->|Bearer Tokens| APIs
    BFF <-->|PKCE Exchange| IDP

A BFF acts as an application-specific gateway for the front-end.


Operational capabilities of a BFF

A Backend-for-Frontend can:

  • Hold server credentials: keep sensitive API keys and tokens out of browser memory;
  • Manage sessions: issue encrypted HttpOnly, SameSite=Strict cookies to the SPA;
  • Aggregate responses: combine multiple downstream microservice calls into one tailored payload;
  • Perform edge transformations: translate internal protocols without client complexity.

It replaces direct client token management with a hardened same-origin boundary.


OAuth is delegated authorization

OAuth allows a client to obtain access to resources through an authorization server.

It is not automatically a login protocol.

OpenID Connect adds an identity layer for authentication scenarios.


OAuth roles

resource owner
client
authorization server
resource server

Keep the roles distinct when reasoning about tokens and trust.


Authorization Code flow

browser → authorization server
        ← authorization code
browser → backend/token endpoint with code
        ← access token / session
browser → resource server through chosen architecture

The code is exchanged rather than delivering an access token directly through the browser redirect.


PKCE Phase 1: Authorization request

sequenceDiagram
    autonumber
    actor User as Citizen / User
    participant App as Browser SPA
    participant Auth as Authorization Server (IDP)

    App->>App: 1. Generate code_verifier (random secret)
    App->>App: 2. Compute code_challenge = SHA256(verifier)
    App->>Auth: 3. Redirect to /authorize?code_challenge=...
    User->>Auth: 4. User logs in & grants consent
    Auth-->>App: 5. Redirect with auth_code

The browser receives an authorization code, not a sensitive token.


PKCE Phase 2: Token redemption & verification

sequenceDiagram
    autonumber
    participant App as Browser SPA
    participant Auth as Authorization Server (IDP)
    participant API as Protected API Server

    App->>Auth: 1. POST /token with auth_code + code_verifier
    Note over Auth: 2. Verifies SHA256(verifier) == code_challenge!<br/>Prevents code interception attacks.
    Auth-->>App: 3. Emits Access Token (+ ID Token)
    App->>API: 4. GET /api/data with Bearer token
    API-->>App: 5. Returns protected resource

PKCE ensures an intercepted code cannot be redeemed without the original secret verifier.


Avoid legacy implicit token delivery

Delivering tokens directly in a browser redirect has a broader leakage and handling surface.

Use modern authorization code patterns with PKCE where appropriate, following the identity provider and platform guidance.


Browser-based OAuth has its own threat model

Consider:

  • redirect interception;
  • authorization-code injection;
  • open redirects;
  • state and nonce validation;
  • browser history and referrer leakage;
  • token exposure to scripts;
  • malicious extensions or compromised dependencies.

The browser is not a confidential client environment.


OpenID Connect

OIDC adds identity claims and an ID token to OAuth-style authorization.

Use it when the application needs to authenticate the user through an identity provider.

Still validate issuer, audience, signature, nonce, time claims, and flow context on the server or trusted verifier.


ID token versus access token

ID token     → information for the client about authentication
Access token → authority presented to a resource server

Do not send an ID token to an API as if it were an access token.

Do not use a client-decoded payload as proof of authorization.


JWT is a format, not an architecture

A JWT can be:

  • signed or unsigned by a chosen algorithm policy;
  • short-lived or long-lived;
  • intended for one audience or another;
  • used in different trust models.

The string shape does not make the authentication design secure.


Token validation belongs at the resource server

The resource server must validate:

  • signature and key policy;
  • issuer;
  • audience;
  • expiry and not-before;
  • scopes or permissions;
  • token type and context.

The front end may display decoded information, but display is not verification.


Front-end token decoding is not verification

const payload = JSON.parse(atob(token.split(".")[1]));

This can read a payload.

It does not prove that the token is authentic, current, intended for this API, or authorized for this action.


Refresh tokens need stronger protection

Refresh tokens can create long-lived access.

Consider:

  • whether the browser should receive them;
  • rotation and reuse detection;
  • secure cookie or BFF storage;
  • revocation;
  • device and session binding;
  • logout behavior.

Do not treat refresh credentials like ordinary UI state.


OAuth state and nonce

state → binds the response to the initiating client flow
nonce → binds identity claims to the authentication request

Validate both in the correct flow and preserve them through redirects safely.


Redirect URI validation

Authorization servers should require registered redirect URIs.

Avoid:

  • wildcard redirect patterns;
  • open redirect chains;
  • accepting attacker-controlled return URLs;
  • mixing trusted and untrusted redirect targets.

Redirect handling is a credential boundary.


Open redirects

/login?returnTo=https://attacker.example

Unvalidated redirects can:

  • enable phishing;
  • leak codes or tokens through chains;
  • make trusted links misleading.

Allowlist internal destinations or use opaque server-side state.


Secrets do not belong in front-end bundles

Anything shipped to the browser can be inspected by:

  • users;
  • browser extensions;
  • automated tools;
  • copied source maps;
  • network observers under the user’s control.

Server credentials and private keys must remain on trusted server infrastructure.


Public API keys are a separate category

Some client identifiers are intentionally public and protected by:

  • origin restrictions;
  • quotas;
  • limited scopes;
  • server-side enforcement;
  • monitoring.

Calling a value a “public key” does not make every key safe to expose.


Environment variables are not a vault

build-time client variable → likely embedded in public output

Use server-side secret management for confidential values.

Review compiled assets and source maps for accidental leakage.


Subresource Integrity

<script
  src="https://cdn.example/script.js"
  integrity="sha384-…"
  crossorigin="anonymous">
</script>

SRI lets the browser verify that a fetched resource matches an expected cryptographic digest.

It is useful for fixed external resources with stable content.


What SRI protects against

SRI can detect a changed resource at the browser boundary.

It does not:

  • make a trusted third-party script safe by itself;
  • protect inline code;
  • validate dynamic resource selection;
  • replace CSP or dependency review;
  • prevent a trusted script from doing harmful things.

SRI and CORS

Cross-origin integrity checks require the resource to be delivered with appropriate CORS behavior.

Coordinate:

  • integrity attribute;
  • crossorigin mode;
  • resource response headers;
  • CDN deployment policy.

Dependency supply-chain risk

Dependencies can execute code during:

  • installation;
  • build;
  • development;
  • application runtime.

Review packages by capability, maintenance, provenance, update behavior, and access to credentials.


Build-time dependencies can be highly privileged

A build plugin may read:

  • source files;
  • environment variables;
  • CI credentials;
  • generated artifacts;
  • deployment configuration.

Keep CI secrets scoped and avoid installing unnecessary tooling into privileged environments.


Reduce supply-chain exposure

Use:

  • minimal dependencies;
  • lockfiles and review;
  • trusted registries;
  • vulnerability and behavior monitoring;
  • restricted scripts where appropriate;
  • separate build and deploy credentials;
  • reproducible artifacts.

Pinning helps reproducibility but does not eliminate malicious or compromised code.


Third-party scripts are full trust grants

A script executing in the page origin may read:

  • DOM content;
  • accessible application state;
  • non-HttpOnly storage;
  • user input;
  • API responses visible to the page.

Load third-party code only when its capability and risk are justified.


Clickjacking

Clickjacking tricks a user into interacting with a framed page or disguised control.

Protect sensitive pages with framing policy and deliberate embedding rules.

Do not rely on visual design alone to prevent deceptive framing.


Prevent unauthorized framing

Relevant controls can include:

Content-Security-Policy: frame-ancestors 'self'

and appropriate legacy-compatible headers where needed.

Allow only known embedding origins when framing is an actual product requirement.


Browser isolation

Isolation policies help control interactions among browsing contexts and cross-origin resources.

They are useful for high-risk capabilities, cross-origin data, and protection from opener or embedding relationships.

They can also break integrations, so test before enforcement.


COOP

Cross-Origin-Opener-Policy controls whether a document shares a browsing context group with cross-origin documents.

It can reduce window.opener relationships and support stronger isolation.


window.opener

Opening a new window can create an opener relationship.

That relationship can enable unexpected navigation or cross-window interaction.

Use safe link behavior and appropriate opener policy for untrusted destinations.


COEP

Cross-Origin-Embedder-Policy controls whether cross-origin resources can be embedded under the document’s isolation requirements.

It may be needed for cross-origin isolated capabilities.

It can also require every dependency and integration to provide compatible headers.


CORP

Cross-Origin-Resource-Policy lets a resource express which origins may load it in certain cross-origin contexts.

It is a resource-side policy, not the same as CORS or COEP.


COOP, COEP, and cross-origin isolation

Together, appropriate COOP and COEP policies can establish a cross-origin isolated context for specific browser capabilities.

Check:

  • workers;
  • images and fonts;
  • analytics;
  • iframes;
  • third-party libraries;
  • CDN headers.

CORS, CORP, and COEP are different

PolicyMain question
CORSmay browser script read this response?
CORPmay this resource be loaded cross-origin in this context?
COEPwhich embedded resources may this document accept?

Use the policy that matches the boundary being controlled.


Security headers need testing

Test headers in:

  • production-like origins;
  • authenticated and unauthenticated flows;
  • embedded and popup scenarios;
  • worker and asset loading;
  • third-party integrations;
  • error and redirect paths.

A header that “looks secure” but breaks recovery or silently disables a feature is not a finished design.


Security architecture for a typical SPA

flowchart LR
    A["Browser UI Client"] -->|1. Same-Origin Cookie Session| B["BFF / Gateway Server"]
    B -->|2. Server-Enforced Role & Scope Authorization| C["Core Business API"]
    C -->|3. Validated Database Queries| D[("Municipal PostgreSQL")]
    
    A -.->|NEVER trust client claims for authorization!| C

The front end presents permissions and handles UX.

The server owns enforcement, secrets, and trusted token validation.


browser ⇄ same-origin app/API
        cookie session

Review:

  • HttpOnly, Secure, SameSite;
  • CSRF tokens or equivalent controls;
  • session rotation;
  • logout and cache clearing;
  • CORS avoidance through same-origin deployment where practical.

OAuth SPA

browser → authorization server → code + PKCE
       → chosen token/session boundary

Document:

  • redirect URIs;
  • state and nonce;
  • token audience and scope;
  • storage and refresh policy;
  • logout and revocation;
  • server-side validation.

BFF architecture

browser → BFF session
BFF     → access tokens and downstream APIs

The BFF can reduce browser token exposure and normalize multiple APIs.

It adds a service to deploy, observe, scale, and secure.


Security UX matters

Good security behavior should tell users:

  • what happened;
  • whether their data was saved;
  • whether they need to sign in;
  • whether they lack permission;
  • how to recover;
  • whether an action is pending or rejected.

Do not reveal sensitive details merely to make an error sound precise.


Error messages must not leak sensitive detail

Avoid exposing:

  • whether a private account exists;
  • stack traces;
  • database structure;
  • tokens;
  • internal URLs;
  • authorization details that aid enumeration.

Log diagnostic context securely and show a useful, bounded user message.


Security and logging

Logs should support investigation without becoming a data leak.

Define:

  • what identifiers are safe;
  • what must be redacted;
  • retention period;
  • access controls;
  • correlation IDs;
  • incident response ownership.

Never log secrets simply because a request failed.


Sensitive data in URLs

URLs can appear in:

  • browser history;
  • referrer headers;
  • server logs;
  • analytics;
  • screenshots;
  • copied links.

Do not place passwords, access tokens, private records, or sensitive form values in query strings or fragments without a very deliberate design.


postMessage is an explicit cross-origin channel

window.postMessage(message, "https://trusted.example");

Use an exact target origin where possible.

Treat both outgoing and incoming messages as untrusted protocol data.


Validate postMessage origin and shape

window.addEventListener("message", event => {
  if (event.origin !== "https://trusted.example") return;
  const message: unknown = event.data;
  handleValidatedMessage(message);
});

Check origin, source window when relevant, message type, and payload schema.


Iframes are security boundaries

For an iframe, decide:

  • which origin it uses;
  • whether it needs sandboxing;
  • which capabilities it receives;
  • whether it may navigate or submit forms;
  • how it communicates;
  • who may frame your application.

Embedding is an architecture decision, not only a layout choice.


Framework escaping is good - but not sufficient

React, Vue, and other frameworks make common text rendering safer by default.

They cannot decide:

  • whether a URL is allowed;
  • whether rich HTML should be sanitized;
  • whether a third-party script is trustworthy;
  • whether an API call is authorized;
  • whether a token belongs in the browser.

Framework safety is one layer in a larger boundary model.


Security review checklist: inputs and rendering

Inputs
  URL, API, forms, storage, messages, third parties
Rendering
  text by default, reviewed HTML, safe URL and style contexts

Trace every untrusted source to its eventual sink.


Security review checklist: requests and credentials

Requests     validation, authorization, CSRF, retry behavior
Credentials  cookie/token scope, expiry, rotation, logout

Ask what happens when the request is replayed, delayed, cross-origin, or sent by a non-browser client.


Security review checklist: browser policies

CSRF, CORS, CSP, framing, COOP, COEP, CORP, SRI

Each policy should have:

  • a threat it addresses;
  • a scope;
  • an owner;
  • tested integrations;
  • a failure and rollout plan.

Security review checklist: authentication and authorization

Verify:

  • identity is established through a supported flow;
  • tokens are validated by the correct server;
  • permissions are enforced server-side;
  • client UI reflects but does not enforce authority;
  • logout clears or revokes relevant state;
  • redirects and callbacks are constrained.

A layered security model

browser policy
  → safe rendering
  → request protection
  → authentication
  → authorization
  → dependency and deployment controls
  → monitoring and recovery

No layer is perfect.

The value comes from reducing the impact when one assumption fails.


Practical lab: Secure a Front-End Application Boundary

Review and harden a small application boundary involving untrusted input, cross-origin requests, cookies, OAuth-style redirects, and dangerous sinks.

The practical turns each security concept into an observable browser behavior and documented design decision.


Practical stages 1–4: origins and CORS

  1. Map the origins.
  2. Observe a CORS failure.
  3. Add explicit CORS permission.
  4. Trigger a preflight.

Distinguish browser read permission from server authentication and authorization.


Practical stages 5–10: rendering defenses

  1. Demonstrate safe text rendering.
  2. Identify a dangerous sink.
  3. Add rich-text sanitization.
  4. Add CSP in report-only mode.
  5. Enforce a basic CSP.
  6. Explore Trusted Types.

Document which content is text, which is approved rich HTML, and which sinks remain intentionally available.


Practical stages 11–14: cookies and permissions

  1. Implement a cookie session.
  2. Add CSRF protection.
  3. Experiment with SameSite.
  4. Separate authentication and authorization.

Verification: hiding a control does not grant permission, and cookie-authenticated mutations have an explicit CSRF decision.


Practical stages 15–18: identity architecture

  1. Draw a BFF architecture.
  2. Model OAuth authorization code plus PKCE.
  3. Compare ID and access tokens.
  4. Search the bundle for secrets.

Keep token validation and confidential credentials on the appropriate server boundary.


Practical stages 19–24: browser isolation

  1. Add SRI to a fixed external resource.
  2. Inventory third-party trust.
  3. Secure postMessage.
  4. Explore framing policy.
  5. Diagram COOP, COEP, and CORP.
  6. Build the security architecture map.

Document affected integrations and failure behavior for every security change.


Practical extension: compare session and bearer designs

Compare:

same-origin session / BFF
browser-held bearer token

Evaluate XSS, CSRF, token exposure, refresh, logout, cross-origin APIs, and operational complexity.

Avoid declaring one architecture universally correct.


Try this yourself

Trace one value from:

URL → parser → component → DOM sink

Then repeat for:

API → state → request → server authorization

Mark where validation, encoding, authentication, authorization, and logging occur.


Troubleshooting guide (Part 1)

SymptomLikely cause
CORS fails in production onlyDevelopment proxy hid real origins
Browser blocks a request but curl succeedsCORS governs browser reads, not server access
Escaped text still creates a dangerous linkURL context was not validated
CSP breaks analytics or workersDependencies were not inventoried before enforcement
CSRF token is missing on a mutationCookie session and request protection were not designed together

Troubleshooting guide (Part 2)

SymptomLikely cause
Hidden button is treated as authorizationServer enforcement is missing
Token payload looks validDecoding is not signature or audience validation
Logout leaves private data visibleClient caches, connections, or local storage were not cleared
iframe integration breaks after isolation headersPolicy was copied without testing dependencies

Completion checklist

  • origins and cross-origin relationships are documented;
  • CORS is treated as browser read permission, not server security;
  • untrusted inputs and dangerous sinks are mapped;
  • text rendering is preferred over HTML;
  • rich HTML has a reviewed sanitization policy;
  • CSP and Trusted Types are defense in depth;
  • cookie mutations have CSRF protection decisions;
  • authentication and authorization are separate;
  • tokens and secrets stay within appropriate boundaries;
  • third-party and browser-isolation policies are tested.

Misconceptions to leave behind (Part 1)

MisconceptionBetter mental model
CORS protects the API from attackersIt controls browser script read access
CORS failure means the server did not receive the requestBrowser policy and server receipt differ
XSS means only <script> tagsAny attacker-controlled executable context can matter
Framework escaping makes XSS impossibleEscape hatches and context-specific sinks remain
Sanitization and encoding are the sameThey solve different content problems
CSP prevents XSS by itselfIt is defense in depth
CSRF and XSS are the sameThey exploit different trust paths

Misconceptions to leave behind (Part 2)

MisconceptionBetter mental model
HttpOnly prevents XSSIt limits direct cookie reads, not same-origin actions
Hidden controls enforce permissionServer authorization must enforce it
JWT means secure authenticationJWT is a format, not an architecture
ID and access tokens are interchangeableThey serve different audiences and purposes
.env values are secretClient-exposed build values are public
Lockfiles eliminate supply-chain riskThey improve reproducibility, not trust
COOP, COEP, and CORP are interchangeableEach controls a different browser boundary

The chapter in one sentence

Secure the browser application by tracing trust across origins, content, requests, identity, dependencies, and embedded contexts - and enforce each decision at the authoritative boundary.


Next: Chapter 14

The next chapter will build on secure architecture with:

  • front-end testing strategy;
  • confidence boundaries and test levels;
  • behavior, integration, and browser verification;
  • performance and resilience testing;
  • quality workflows for production systems.

Questions

Which value in your application crosses the most trust boundaries, and which layer currently makes the security decision about it?

14 - Scaling Front-End Architecture: Design Systems, Monorepos & Micro-Frontends

Chapter 14: scale shared UI, packages, teams, and deployments without turning reuse into hidden coupling.

Scaling Front-End Architecture: Design Systems, Monorepos & Micro-Frontends

Scale the boundaries, not the confusion

Chapter 14

Polla Fattah


Today’s goal

Understand how front-end architecture changes when products, packages, teams, and deployments grow.

We will connect:

  • reusable components and design systems;
  • tokens, themes, accessibility, documentation, and versioning;
  • package APIs, ownership, workspaces, and monorepos;
  • task graphs, affected builds, and caching;
  • platform teams and golden paths;
  • micro-frontend composition and runtime independence;
  • Module Federation, contracts, failure containment, and observability;
  • migration strategy and architecture decisions at scale.

By the end of today you can

  • distinguish a component library from a design system;
  • define raw and semantic design tokens;
  • design stable package APIs and release policies;
  • decide when a monorepo or separate repositories help;
  • model dependency direction and affected builds;
  • explain why micro-frontends are organizational and deployment architecture;
  • choose server, client, build-time, or runtime composition;
  • design contracts, failure boundaries, and observability;
  • evaluate Module Federation trade-offs;
  • recognize when a modular monolith is the better answer.

The central principle

Scale through explicit contracts, ownership, and failure boundaries - not through maximum sharing or maximum distribution.

Every shared abstraction creates a coordination obligation.

Every independently deployed unit creates an integration obligation.


Scale changes the nature of front-end problems

At small scale, a change may affect one application.

At larger scale, the same change may affect:

  • several products;
  • multiple teams;
  • package consumers;
  • release trains;
  • shared tokens and accessibility;
  • deployment and runtime contracts.

The architecture must make change impact visible.


Technical scale and organizational scale differ

technical scale → code, builds, bundles, runtime complexity
organizational scale → teams, ownership, coordination, release authority

A micro-frontend does not solve a team-ownership problem automatically.

A monorepo does not solve a dependency-design problem automatically.


Reuse is not automatically good

Reuse can reduce:

  • duplicated fixes;
  • inconsistent behavior;
  • accessibility defects;
  • visual drift.

Reuse can also create:

  • shared release coupling;
  • generic APIs;
  • cross-domain assumptions;
  • difficult migration;
  • a global dependency nobody can change.

Reuse is a decision with a cost.


Prefer stable shared concepts

Good shared concepts often include:

Button
Dialog
Field
Tabs
tokens and themes

Be careful sharing a component whose behavior is actually product-specific:

ApprovalWorkflowPanel
RegionalTaxEditor
CataloguePricingCard

Share the stable concept, not the current visual coincidence.


Component library versus design system

component library → reusable implementation
design system     → product of principles, tokens, components, guidance, and governance

A design system includes how people choose, use, test, document, version, and evolve the components.


A design system is a product

It has:

  • users: product teams and customers;
  • a roadmap;
  • documentation;
  • support expectations;
  • release notes;
  • migration paths;
  • quality and accessibility goals;
  • adoption feedback.

Treating it as “just a package” underestimates its operational work.


Design-system architecture

flowchart TD
    subgraph DesignSystemPyramid["The Design System Architecture Pyramid"]
        L1["1. Foundations: Color theory, typography scales, spacing grids"]
        L2["2. Design Tokens: Global raw tokens & semantic intent tokens"]
        L3["3. UI Primitives: Accessible Headless/Compound components (Button, Modal, Input)"]
        L4["4. Composed Patterns: Form layouts, data tables, alert banners"]
        L5["5. Documentation & Guidelines: Usage examples, do's/don'ts, accessibility notes"]
        L6["6. Governance & Versioning: SemVer release policies, deprecation lifecycles"]
        L1 --> L2 --> L3 --> L4 --> L5 --> L6
    end

The layers should have intentional dependency direction.


Foundations define shared vocabulary

Foundations may include:

  • color and contrast;
  • typography;
  • spacing;
  • elevation;
  • motion;
  • layout;
  • iconography;
  • interaction states.

They should support product consistency without hiding meaningful product differences.


Design tokens are named decisions

{
  "color": {
    "brand": { "primary": "#2b5cff" }
  },
  "space": { "md": "1rem" }
}

Tokens make recurring design decisions inspectable and transformable across platforms.


Raw and semantic tokens

raw/foundation: blue-600, space-4, radius-2
semantic:       action-primary, surface-muted, focus-ring

Raw tokens describe values.

Semantic tokens describe meaning and can change when a theme or brand changes.


Why semantic tokens matter

--color-action-primary: var(--color-blue-600);

Components consume meaning:

background: var(--color-action-primary);

The component need not know which palette value currently represents the primary action.


Token interoperability

Tokens may be consumed by:

  • CSS variables;
  • JavaScript or TypeScript;
  • native mobile platforms;
  • design tools;
  • documentation and testing tools.

Define naming, units, themes, fallbacks, and transformation rules so the token system remains coherent across outputs.


A token pipeline

flowchart LR
    A["Source Tokens (JSON / W3C Format)"] --> B["Token Transformer (Style Dictionary)"]
    B --> C1["CSS Custom Properties (:root { --color-action: #... })"]
    B --> C2["TypeScript Types & Constants"]
    B --> C3["Figma / Design Tool Sync"]

The pipeline should preserve semantic meaning, not only copy strings.


CSS variables provide runtime tokens

:root {
  --surface-page: white;
  --text-primary: #121212;
}

[data-theme="dark"] {
  --surface-page: #121212;
  --text-primary: white;
}

Runtime variables enable themes without rebuilding every component.


Tokens are not a complete design system

Tokens do not define:

  • keyboard behavior;
  • accessible names;
  • component states;
  • content guidance;
  • composition rules;
  • error behavior;
  • ownership or release policy.

They are a foundation, not the whole system.


Component API stability matters more at scale

Every public prop, event, slot, CSS hook, and token name can have many consumers.

Before changing it, ask:

  • who uses it?
  • what behavior is relied upon?
  • is the change additive?
  • what migration path exists?
  • can the old API be deprecated safely?

Public versus internal component APIs

public API  → documented, supported, versioned
internal API → implementation detail, changeable with local ownership

Do not expose every internal component just because it exists in the repository.

Small public surfaces are easier to keep reliable.


Escape hatches need boundaries

Examples:

custom class
render prop
slot
unstyled mode
as / component override

Escape hatches support real variation, but can undermine consistency and accessibility if unbounded.

Document what the consumer becomes responsible for when using one.


Headless components at scale

Headless components can centralize:

  • keyboard behavior;
  • focus management;
  • ARIA relationships;
  • selection state;
  • interaction lifecycle.

The consuming product controls presentation while inheriting a tested behavior contract.


Accessibility is part of the component contract

For a dialog, tabs, menu, or combobox, the contract includes:

  • roles;
  • names;
  • keyboard behavior;
  • focus movement;
  • disabled states;
  • announcements;
  • escape and dismissal rules.

Visual similarity is not enough for a shared component.


Design-system documentation is an interface

Documentation should show:

  • when to use the component;
  • when not to use it;
  • accessible behavior;
  • controlled and uncontrolled modes;
  • composition examples;
  • supported variants;
  • migration and deprecation notes.

Examples are part of the package’s practical API.


Component explorers create feedback loops

A component explorer can provide:

  • isolated rendering;
  • visual examples;
  • interaction states;
  • accessibility checks;
  • visual regression targets;
  • documentation source.

Keep examples representative of supported contracts rather than every possible boolean combination.


Design-system versioning

Versioning must cover more than TypeScript signatures.

A release may change:

  • visual output;
  • keyboard behavior;
  • ARIA structure;
  • CSS variables;
  • token meaning;
  • generated assets;
  • browser support.

Consumers need clear impact information.


Semantic versioning is a communication contract

PATCH → compatible fix
MINOR → compatible capability
MAJOR → consumer action may be required

Visual or accessibility changes can be breaking even when the TypeScript type still compiles.


Visual changes can be breaking changes

A “small” change can break:

  • layout assumptions;
  • screenshot baselines;
  • contrast;
  • click targets;
  • responsive behavior;
  • user workflows.

Describe visual and interaction impact in release notes.


Deprecation is a lifecycle

announce → document replacement → measure usage → codemod / migrate → remove

Leaving deprecated APIs forever increases package complexity and makes the preferred path unclear.


Codemods make migrations repeatable

A codemod can transform predictable usage:

old prop name → new prop name
old import path → public entry point
old token → semantic token

Use codemods where the transformation is mechanically reliable, then review semantic edge cases.


Release notes and migration guides

Good notes explain:

  • what changed;
  • why it changed;
  • who is affected;
  • what to search for;
  • how to migrate;
  • how to verify behavior;
  • when removal occurs.

Migration work is part of the design-system product.


Shared packages should represent responsibility

shared/ui       reusable interaction and visual primitives
shared/tokens   semantic design vocabulary
shared/config   agreed tooling policy
domain/catalogue product-specific vocabulary

Do not make a shared package the default home for anything that might be reused.


Package boundaries are stronger than folders

A package boundary can enforce:

  • public exports;
  • dependency direction;
  • versioning;
  • ownership;
  • build and test scope;
  • consumer expectations.

Folders organize files. Packages communicate architecture.


Avoid the “shared everything” package

One giant shared package creates:

  • unclear ownership;
  • broad rebuilds;
  • accidental imports;
  • difficult releases;
  • hidden domain coupling.

Split by stable responsibility, not by arbitrary file count.


Shared domain packages need caution

Sharing a domain package can reduce duplicated contracts.

It can also force products to adopt one domain model before their workflows are truly aligned.

Share stable concepts and schemas; keep product-specific orchestration local.


Shared types can create false confidence

shared type → compile-time agreement
runtime data → still needs validation

Sharing a TypeScript type does not make two independently deployed systems compatible at runtime.

Add versioned contracts, validation, and compatibility tests where needed.


Internal package dependency direction

flowchart TD
    Apps["Applications Layer (apps/citizen-portal, apps/inspector-app)"] --> Feat["Feature Packages (packages/feature-licensing)"]
    Feat --> Domain["Domain Logic & Models (packages/domain-permits)"]
    Domain --> UI["Design System UI (packages/ui-primitives)"]
    UI --> Tokens["Foundations & Tokens (packages/design-tokens)"]

If a shared package imports an application feature, the boundary has reversed.

Make direction enforceable with lint rules, package exports, or graph checks.


Workspaces coordinate packages

Workspaces can provide:

  • local package linking;
  • shared install and lockfile;
  • cross-package scripts;
  • consistent tooling;
  • coordinated local development.

They do not decide whether packages are well-designed.


Monorepo benefits

atomic cross-app changes
shared refactoring
visible dependency graph
consistent tooling
local package feedback

These benefits are strongest when projects genuinely evolve together and ownership is clear.


Monorepo costs

Expect:

  • larger graph and CI coordination;
  • build ordering;
  • ownership disputes;
  • broad blast radius;
  • package version and release decisions;
  • accidental internal imports.

A monorepo makes coupling visible; it does not make coupling disappear.


Monorepo versus polyrepo

monorepo → shared visibility and atomic changes
polyrepo  → stronger repository independence and separate lifecycle

Neither is universally superior.

Choose based on release coordination, ownership, dependency contracts, and operational capacity.


Monorepo is not micro-frontend architecture

monorepo       → source and package organization
micro-frontend → runtime or deployment composition

A monorepo can contain one modular application or many separately deployed front ends.

Micro-frontends can be built across one or many repositories.


Task graphs expose build relationships

flowchart LR
    subgraph PipelineGraph["Monorepo Directed Acyclic Task Graph (DAG)"]
        T["tokens:build"] --> U["ui:build"]
        U --> S["storefront:build"]
        D["domain:build"] --> A["admin:build"]
    end

A task graph can determine:

  • build order;
  • affected tests;
  • cacheable work;
  • parallel execution;
  • what a change can impact.

Affected builds reduce unnecessary work

If only documentation changes, a full application rebuild may be unnecessary.

If tokens change, many consumers may be affected.

The graph should be accurate enough to make selective verification trustworthy.


Caching build tasks

Cache tasks when their outputs are determined by:

  • source inputs;
  • dependency versions;
  • tool versions;
  • environment inputs;
  • configuration.

Incorrect cache keys are worse than no cache because they create false confidence.


Remote build cache

A remote cache can share verified task outputs across developers and CI.

Protect it with:

  • correct input hashing;
  • access control;
  • artifact integrity;
  • retention policy;
  • invalidation rules.

Speed must not weaken reproducibility or trust.


Repository ownership

Ownership should answer:

  • who maintains the package;
  • who reviews changes;
  • who handles releases;
  • who responds to failures;
  • who approves breaking changes.

Ownership is coordination infrastructure, not a reason to prevent useful contributions.


Federated contribution

A platform team can provide:

  • foundations;
  • paved workflows;
  • documentation;
  • release automation;
  • support.

Product teams should be able to contribute improvements without routing every decision through one gatekeeper.


Platform teams

A platform team should reduce product team friction through:

  • reliable defaults;
  • clear extension points;
  • observability;
  • migration support;
  • sensible constraints.

The platform is successful when teams can move safely, not when it controls every implementation.


Platform as a product

Treat internal developers as users.

Measure:

  • adoption;
  • time to first useful feature;
  • migration cost;
  • support volume;
  • build reliability;
  • accessibility and quality outcomes.

Build what reduces real product friction.


Paved road versus mandatory road

paved road → recommended, supported default
mandatory   → enforced constraint for a real safety or compatibility need

Use mandatory rules sparingly and explain the reason.

Too many constraints encourage teams to work around the platform.


Golden path

A golden path can provide:

  • starter repository;
  • standard scripts;
  • secure deployment;
  • observability;
  • supported dependencies;
  • examples and migration guides.

It should make the safe path convenient without making every product identical.


Shared infrastructure must remain replaceable

Avoid APIs that expose every implementation detail of:

  • the build tool;
  • the deployment platform;
  • the state library;
  • the rendering runtime.

Stable capabilities allow infrastructure to evolve without forcing every product to migrate at once.


Versioning internal packages

Possible models:

lockstep versions
independent versions
source-consumed packages

Choose according to release coordination and compatibility, not habit.


Lockstep versioning

All packages release together.

Benefits:

  • coordinated compatibility;
  • simple version alignment;
  • atomic platform upgrades.

Costs:

  • unrelated consumers receive synchronized changes;
  • releases may become slower or noisier.

Independent versioning

Packages release on their own cadence.

Benefits:

  • smaller consumer updates;
  • local ownership;
  • separate release urgency.

Costs:

  • compatibility matrix;
  • upgrade coordination;
  • more release automation.

Source consumption in monorepos

Consuming internal source directly can speed local development and preserve one graph.

It can also blur package boundaries and make production behavior differ from published consumption.

Decide whether the package boundary is real and test the intended production mode.


Micro-frontends are not small components

component → local composition and behavior
micro-frontend → independently owned application capability

Micro-frontends involve:

  • team ownership;
  • deployment lifecycle;
  • runtime or route composition;
  • integration contracts;
  • failure and observability boundaries.

Why micro-frontends are considered

They may help when:

  • teams need independent deployment;
  • domains have different release cadence;
  • a platform must integrate separately owned products;
  • migration from a legacy application must be incremental;
  • organizational boundaries are stable and meaningful.

They introduce coordination costs of their own.


Vertical slicing

team owns a business capability
  → UI + data + workflow + deployment

Vertical slices can align technical ownership with user-facing capability better than splitting by framework layer.


Route-level composition

flowchart TD
    Gateway["Edge Reverse Proxy / Gateway (Cloudflare / NGINX)"]
    Gateway -->|/catalogue/*| App1["Catalogue Micro-App (Next.js / Team Commerce)"]
    Gateway -->|/reports/*| App2["Audit & Reports Micro-App (Vite / Team Analytics)"]
    Gateway -->|/settings/*| App3["Citizen Account Micro-App (Remix / Team Identity)"]

Full-page or route-level composition is often simpler than embedding several runtimes into one screen.

It preserves clear navigation and failure boundaries.


Server-side composition

The server can assemble fragments or route outputs before delivery.

Benefits:

  • clear browser payload;
  • centralized navigation and auth;
  • independent backend ownership.

Costs include server orchestration, shared layout contracts, and failure handling.


Client-side composition

shell → loads remote feature → mounts feature

Client composition can allow independent deployment and rich in-page integration.

It must manage loading, version compatibility, failure, routing, styling, and shared state.


Build-time composition

A shell can consume packages or generated artifacts during its build.

This gives strong integration testing and simple runtime behavior.

It reduces deployment independence because the shell must rebuild to receive changes.


Runtime independence is a spectrum

flowchart LR
    A["1. Shared Git Monorepo"] --> B["2. Versioned npm Package"]
    B --> C["3. Build-Time Module Split"]
    C --> D["4. Runtime Module Federation"]
    D --> E["5. Sandboxed <iframe>"]

Moving right can increase deployment independence and isolation while increasing integration complexity.


The micro-frontend shell

The shell often owns:

  • global navigation;
  • authentication context;
  • layout;
  • routing;
  • loading and failure boundaries;
  • observability;
  • shared design-system contract.

Avoid making the shell a hidden global store for every domain.


Shared state across micro-frontends

Prefer explicit contracts:

URL state
server state
custom events
shared identity/session boundary

Avoid sharing mutable in-memory domain state unless the integration truly requires it and versioning is controlled.


Prefer server and URL contracts

Server APIs and URLs are durable boundaries.

They can survive:

  • different framework versions;
  • independent deployments;
  • full-page navigation;
  • process restarts;
  • separate repositories.

In-memory object sharing is fast but fragile across runtime boundaries.


Custom events can cross local boundaries

window.dispatchEvent(new CustomEvent("cart.updated", {
  detail: { count: 2 },
}));

Define:

  • event names;
  • payload schema;
  • versioning;
  • ownership;
  • duplicate behavior;
  • failure behavior.

Events should be protocol contracts, not accidental DOM messages.


Micro-frontend contract design

A contract may include:

  • mount inputs;
  • lifecycle methods;
  • events;
  • URL behavior;
  • auth assumptions;
  • CSS scope;
  • data and error models;
  • compatibility window.

Write the contract before choosing the integration technology.


CSS isolation

Options include:

CSS Modules
scoped styles
shadow DOM
namespaced conventions
design-system tokens

No option is free. Test global resets, typography, overlays, z-index, and responsive behavior across boundaries.


The design system is the visual contract

Across micro-frontends, a shared design system can align:

  • tokens;
  • controls;
  • accessibility;
  • spacing and typography;
  • interaction states.

It does not make domain workflows identical.


Design-system version drift

Independent applications may consume different versions.

Plan for:

  • compatible token evolution;
  • deprecation windows;
  • visual regression;
  • accessibility review;
  • migration ownership;
  • runtime and bundle duplication.

Independent deployment does not remove visual coordination.


Shared runtime dependencies

Sharing a framework or library can reduce duplicate downloads.

It can also create:

  • version coupling;
  • singleton assumptions;
  • incompatible runtime copies;
  • upgrade coordination.

Measure the cost of duplication against the cost of shared runtime coupling.


Module Federation: host and remote

host   → loads and integrates remote
remote → publishes an independently built capability

The mechanism can enable runtime module loading.

It does not automatically solve contracts, security, routing, CSS, or failure.


Runtime module loading

sequenceDiagram
    autonumber
    actor User as Citizen / Inspector
    participant Host as Shell Host Application (port 3000)
    participant Remote as Remote Micro-App (port 3001)

    User->>Host: Navigates to /licensing/audit
    Note over Host: Host encounters dynamic import('licensingRemote/AuditPanel')
    Host->>Remote: Fetch remoteEntry.js (Manifest of exposed modules & shared singletons)
    Remote-->>Host: Returns module federation metadata
    Note over Host: Host verifies shared React version (18.3.1 === 18.3.1)
Re-uses host React runtime without downloading second copy!
    Host->>Remote: Fetch AuditPanel.[hash].js
    Remote-->>Host: Emitted chunk code
    Note over Host: Host mounts Remote Component into shell DOM tree!

Every step can fail or become incompatible.

Represent loading, timeout, version, and fallback states in the host.


Runtime failure is normal

A remote can fail because of:

  • deployment outage;
  • network failure;
  • incompatible API;
  • missing asset;
  • authentication issue;
  • broken initialization.

The shell should preserve useful navigation and explain which capability is unavailable.


Shared dependencies and singleton risk

Sharing one runtime instance can reduce duplication.

It can also make independent remotes depend on:

  • one framework version;
  • one context implementation;
  • one runtime initialization;
  • one upgrade schedule.

Singletons are a coordination contract.


Deployment independence versus runtime coupling

independent deploy + shared runtime
  → release freedom with compatibility obligations

The more a remote relies on host internals, the less independent it truly is.

Measure independence by the contracts a team can change without synchronized deployment.


Micro-frontend observability

Trace:

  • host route and remote identity;
  • remote version;
  • load time and failure;
  • mount and unmount;
  • network and API errors;
  • user journey across boundaries.

Without cross-boundary correlation, failures become “the page is broken” with no owner.


Micro-frontend testing

Use several levels:

remote independently
contract compatibility
shell integration
end-to-end user journey

No single level proves that independent deployment remains safe.


Micro-frontend deployment

Deployment should define:

  • version publication;
  • manifest or entry discovery;
  • rollback;
  • cache invalidation;
  • compatibility support window;
  • monitoring and alert ownership.

The deployment system is part of the runtime contract.


Rollback

Runtime composition should make it possible to:

  • revert a remote version;
  • disable a capability;
  • serve a fallback;
  • keep the host usable;
  • identify affected users and routes.

Fast rollback is more valuable than theoretical independence.


Canary releases

Canaries can expose a new remote to:

  • selected users;
  • one region;
  • one route;
  • a small traffic percentage.

Monitor errors, load time, contract failures, and user outcomes before broad release.


Micro-frontend security

Ask whether a remote:

  • shares the same origin;
  • can read host DOM or storage;
  • receives auth context;
  • can navigate the shell;
  • loads third-party code;
  • needs iframe isolation;
  • can affect the whole page.

Deployment independence is not security isolation.


Same-origin micro-frontends share trust

If multiple applications run in the same origin and page context, a compromised remote may have broad access to the host environment.

Use explicit trust review, CSP, dependency policy, and isolation where needed.


Iframe isolation

Iframes can provide stronger origin and DOM isolation.

They add costs:

  • sizing and layout;
  • navigation and focus;
  • communication protocol;
  • authentication;
  • accessibility;
  • performance;
  • mobile behavior.

Use them when isolation is worth the integration cost.


Independent frameworks are not automatically better

Different teams may choose different frameworks.

That can support local autonomy but creates:

  • duplicate runtimes;
  • inconsistent accessibility;
  • design drift;
  • larger bundles;
  • harder shared behavior.

Use multiple frameworks only when the boundary and trade-off are justified.


Conway’s Law effect

Systems tend to reflect the communication structures of the organization.

If teams own stable business capabilities, vertical architecture may align well.

If teams are split around temporary technical layers, the architecture may reproduce handoffs and integration friction.


Team topology and front-end architecture

Consider:

  • stream-aligned product teams;
  • platform teams;
  • enabling teams;
  • ownership of shared contracts;
  • on-call and incident response.

Architecture should support how teams actually coordinate and change the product.


When a monolith is better

A modular monolith may be preferable when:

  • one team can coordinate releases;
  • shared state and navigation are central;
  • deployment independence has little value;
  • runtime composition would add more failure modes than it removes.

“Monolith” does not mean “unstructured”.


Modular monolith first

one deployable application
  + explicit feature modules
  + package boundaries
  + dependency direction
  + independent tests and ownership

This can provide many architectural benefits before runtime distribution is necessary.


Signs micro-frontends may be justified

Consider them when:

  • teams need independent release cadence;
  • domains have stable boundaries;
  • the shell can define durable contracts;
  • failure containment has value;
  • migration needs incremental replacement;
  • observability and platform support exist.

Signs micro-frontends are probably unnecessary

Warning signs:

  • the only reason is “the app is large”;
  • teams need to share mutable state constantly;
  • one team owns the whole product;
  • independent deployment is not actually required;
  • integration testing is weak;
  • the same design system and framework are already tightly coupled.

Start with a modular monolith and re-evaluate.


The decision ladder

flowchart TD
    subgraph ArchitectureScaleLadder["The Front-End Architectural Scale Ladder"]
        L1["Level 1: Local Component (Start here: zero coordination overhead)"]
        L2["Level 2: Workspace Package (Internal monorepo library with typed contracts)"]
        L3["Level 3: Versioned Design System Package (Published npm library across repos)"]
        L4["Level 4: Modular Monolith (Cohesive codebase with strict directory boundaries)"]
        L5["Level 5: Route-Based Multi-Zone Apps (Independent deployments partitioned by URL path)"]
        L6["Level 6: Runtime Micro-Frontends / Module Federation (High cost: earned only by organizational friction)"]
        L1 --> L2 --> L3 --> L4 --> L5 --> L6
    end

Move right only when the lower level cannot satisfy a real requirement.


Architecture decision: design system

Ask:

  • which concepts are stable?
  • who are the consumers?
  • what accessibility contract is shared?
  • how are visual changes reviewed?
  • what can remain product-specific?

A design system should reduce recurring decisions without flattening domain differences.


Architecture decision: monorepo

Ask:

  • do projects change together?
  • are atomic changes valuable?
  • can the team operate the graph and CI?
  • is ownership clear?
  • are package APIs enforceable?

Choose a monorepo for coordination value, not only shared code.


Architecture decision: separate repositories

Separate repositories can provide:

  • independent lifecycle;
  • stronger ownership boundaries;
  • smaller local graphs;
  • separate release and access policy.

They require deliberate contract publication, compatibility testing, and cross-repository coordination.


Architecture decision: runtime federation

Ask:

  • is independent deployment truly required?
  • what happens when a remote is unavailable?
  • how are versions compatible?
  • who owns the shell contract?
  • how are security and observability handled?

Runtime federation should be a response to a specific operational need.


Migration strategy

Large changes are safer when incremental:

map current boundaries
  → establish contracts
  → extract one capability
  → observe and stabilize
  → repeat

Migration architecture should support rollback and coexistence.


The strangler pattern

flowchart TD
    Proxy["API Gateway / Edge Router"]
    Legacy["Legacy Monolith (JSP / ASP.NET / AngularJS)"]
    NewApp["Modern Micro-App (Next.js / Vite)"]

    Proxy -->|90% Legacy Traffic| Legacy
    Proxy -->|10% Migrated Route: /permits/apply| NewApp
    Note over Proxy: Gradually shift routes from Legacy to NewApp until Legacy is completely retired!

Use stable URLs, data contracts, and ownership rules to keep the old and new systems coherent during transition.


Shared navigation during migration

Navigation can be a durable contract across applications:

  • common route vocabulary;
  • active-state behavior;
  • permissions;
  • full-page versus client navigation;
  • analytics and focus.

Do not force all applications into one runtime merely to share a menu.


Full-page navigation can be a feature

Full navigation can provide:

  • clean application boundaries;
  • independent failure recovery;
  • simpler deployment;
  • less shared runtime coupling;
  • natural browser history.

It is not automatically an outdated experience.


Design-system governance at enterprise scale

Governance should clarify:

  • what belongs in foundations;
  • what belongs in shared patterns;
  • what remains product-specific;
  • who approves breaking changes;
  • how contributions are accepted;
  • how accessibility and visual quality are measured.

Contribution tiers

foundation → tokens, primitives, accessibility
shared pattern → repeated cross-product interaction
product-specific → domain behavior and local workflow

Promote upward only when usage, stability, and ownership justify it.


Measure design-system value

Possible signals:

  • adoption of supported components;
  • accessibility defect reduction;
  • time to build common workflows;
  • migration effort;
  • visual consistency;
  • consumer satisfaction;
  • release and support cost.

Count outcomes, not only number of components.


Measure monorepo health

Useful measures include:

  • affected-build accuracy;
  • CI duration;
  • dependency cycle count;
  • package boundary violations;
  • ownership response time;
  • cache hit rate;
  • cross-app change reliability.

Repository size alone says little about health.


Measure micro-frontend health

Watch:

  • remote load success;
  • fallback frequency;
  • contract failures;
  • duplicate runtime cost;
  • cross-boundary latency;
  • deployment rollback time;
  • team independence in practice.

If every change still requires synchronized coordination, the distribution may be mostly ceremonial.


Avoid architecture by organization chart alone

Team boundaries are evidence, not the entire design.

Also examine:

  • business capability boundaries;
  • data ownership;
  • user journeys;
  • release frequency;
  • failure containment;
  • security and performance requirements.

Architecture should reflect durable responsibilities, not temporary reporting lines.


Stable business capabilities

Good boundaries often align with capabilities such as:

catalogue
orders
identity
reporting
workflow

They can be understood by users, teams, APIs, and deployment owners.


Contracts over implementation sharing

Two teams do not need to share implementation to coordinate.

They can share:

  • URL contracts;
  • API schemas;
  • events;
  • design tokens;
  • accessibility expectations;
  • versioned interfaces.

Implementation sharing is one option, not the default definition of collaboration.


Build-time safety versus runtime safety

build-time → types, package graph, contracts, tests
runtime    → loading, versions, network, permissions, failures

Micro-frontends and independently deployed packages need both.

Compilation cannot prove that a remote URL will be available tomorrow.


Failure containment

A boundary is valuable when a failure can remain local:

flowchart TD
    Shell["App Shell Navigation & Header (Healthy)"]
    
    subgraph ViewContainer["Page Content Layout"]
        Catalog["Permit Catalogue Widget (Healthy)"]
        
        subgraph ErrorBoundary["React / Vue Error Boundary"]
            FailedRemote["Remote Analytics Widget (Crash / 500 Network Drop)"]
            Fallback["Fallback UI: 'Analytics temporarily unavailable. Retry'"]
        end
    end

    Shell --- Catalog
    Shell --- ErrorBoundary
    FailedRemote -.->|Error caught by Boundary| Fallback
    Note over Shell: Entire application survives! User continues browsing catalogue.

If every remote failure breaks the shell, distribution has not created meaningful containment.


Global failure dependencies

Beware sharing:

  • one global runtime initialization;
  • one mutable store;
  • one required remote;
  • one unversioned design token package;
  • one authentication callback that blocks every route.

These can make many “independent” units fail together.


Versioned contracts

Version the things that cross boundaries:

  • package APIs;
  • event payloads;
  • route parameters;
  • remote manifests;
  • tokens;
  • server schemas.

Compatibility should be testable and documented.


Consumer-driven contract thinking

Consumers should express the behavior they rely on.

Providers can verify that they remain compatible before publishing.

This catches integration breakage earlier than waiting for a full end-to-end environment.


Security boundaries versus team boundaries

Two teams owning separate modules does not isolate their code from each other at runtime.

If security isolation is required, use actual origin, process, sandbox, or iframe boundaries and design the communication protocol accordingly.


Performance budgets across micro-frontends

Each remote may appear small alone but expensive together.

Track:

  • initial JavaScript;
  • duplicate frameworks;
  • remote request count;
  • hydration or mount cost;
  • CSS and font duplication;
  • runtime memory.

Measure the assembled user experience.


Shared performance budget

Define budgets for:

initial transfer
interactive readiness
route transition
remote failure recovery
device CPU and memory

Every team should know how its boundary contributes to the total.


Practical lab: Design-System Package and Ownership Map

Design a small shared UI package without turning every product-specific component into a design-system primitive.

Map tokens, public APIs, ownership, versioning, consumers, and release responsibility.


Practical stages 1–6: foundations and contracts

  1. Identify repeated foundations, primitives, and domain components.
  2. Define token ownership and semantic naming.
  3. Create a small package boundary and public API.
  4. Document compatibility, versioning, and affected consumers.
  5. Map team ownership.
  6. Add an accessibility contract.

Keep domain-specific behavior with the product team.


Practical stages 7–12: package and repository architecture

  1. Simulate a breaking change.
  2. Create workspace packages.
  3. Enforce package public APIs.
  4. Draw the package dependency graph.
  5. Add ownership.
  6. Add shared tooling.

Verification: the shared package has a small stable surface and does not become a hidden global dependency.


Practical stages 13–16: monorepo operations

  1. Simulate an atomic cross-app change.
  2. Simulate affected builds.
  3. Add task caching conceptually.
  4. Design a platform golden path.

Use the graph to explain what should rebuild, what can be cached, and who reviews each change.


Practical stages 17–22: micro-frontend contracts

  1. Model route-level app separation.
  2. Build a micro-frontend decision record.
  3. Build a small federation demo.
  4. Add remote failure handling.
  5. Explore shared dependencies.
  6. Add an integration contract.

Make runtime loading, version compatibility, fallback, and observability explicit.


Practical stages 23–29: boundary and final architecture

  1. Avoid shared in-memory domain state.
  2. Add a custom event.
  3. Explore CSS collision.
  4. Use the design system across host and remote.
  5. Draw the runtime architecture.
  6. Compare three architectures.
  7. Create the final scaling ADR.

Compare a modular monolith, route-level separation, and runtime federation.


Practical extension: package versus independently deployed remote

Compare:

monorepo package boundary
independently deployed micro-frontend

Explain why they solve different problems around source sharing, release cadence, runtime failure, ownership, and contracts.


Try this yourself

Take one repeated UI element and classify it:

design-system primitive
shared pattern
domain component
product-local component

Write the evidence for the classification: consumers, stability, ownership, accessibility contract, and expected change direction.


Troubleshooting guide (Part 1)

SymptomLikely cause
Shared package accepts every propThe concept is not stable or API is too generic
Design system contains domain workflowsProduct behavior was promoted too early
Every package rebuilds for one changeDependency graph or package boundary is too broad
Teams still coordinate every releaseRuntime independence is mostly nominal
Remote failure breaks the whole shellFailure containment was not designed

Troubleshooting guide (Part 2)

SymptomLikely cause
Different remotes look inconsistentToken, component, or version governance is weak
Micro-frontends share a giant storeDurable contracts were replaced by hidden coupling
Monorepo feels slow and opaqueTask graph, affected builds, or ownership is unclear
Security assumption follows team ownershipTeam boundary is not a runtime isolation boundary

Completion checklist

  • shared concepts are stable and intentionally owned;
  • tokens distinguish raw values from semantic meaning;
  • public package APIs are small and documented;
  • accessibility is part of the shared component contract;
  • versioning and deprecation have migration paths;
  • package dependency direction is enforceable;
  • monorepo tasks and affected builds are observable;
  • platform defaults are supported but not needlessly mandatory;
  • micro-frontends have durable contracts and failure fallbacks;
  • scaling decisions are measured against user and team outcomes.

Misconceptions to leave behind (Part 1)

MisconceptionBetter mental model
A design system is a component libraryIt is a product of foundations, behavior, docs, and governance
Tokens are just CSS variablesTokens are semantic cross-platform decisions
More abstraction means more reuseShared abstractions create coordination cost
A monorepo means one applicationIt is a repository and package organization choice
Workspaces and monorepos are identicalWorkspaces coordinate packages; architecture needs boundaries
Micro-frontends are tiny componentsThey are independently owned application capabilities
Micro-frontends are always more scalableThey trade local autonomy for integration complexity

Misconceptions to leave behind (Part 2)

MisconceptionBetter mental model
Independent deployment means no coordinationContracts, versions, and runtime failures still coordinate
Module Federation prevents remotes from breaking hostsRuntime loading requires compatibility and fallback
Shared dependencies are always betterThey reduce duplication but increase coupling
Different teams require different frameworksTeam boundaries do not automatically justify runtime diversity
Micro-frontends provide security isolationSame-origin code often shares trust
Full-page navigation is outdatedIt can provide valuable isolation and recovery

The chapter in one sentence

Scale front-end systems by making shared contracts, package ownership, deployment independence, and failure boundaries explicit - and distribute only when the benefit exceeds the coordination cost.


Next: Chapter 15

The next chapter will build on scaling architecture with:

  • performance engineering and budgets;
  • measurement and profiling;
  • loading, rendering, and interaction cost;
  • production observability;
  • performance decisions grounded in user journeys.

Questions

Which boundary in your current system is shared because it is genuinely stable, and which one is shared only because it was convenient to put in one package?

15 - Core Web Vitals & Performance Engineering

Chapter 15: measure real user experience, diagnose LCP, INP, and CLS, and optimize from evidence rather than instinct.

Core Web Vitals & Performance Engineering

Measure the experience before changing the code

Chapter 15

Polla Fattah


Today’s goal

Treat performance as a user-experience property and an engineering discipline.

We will connect:

  • field data, lab data, and representative measurement;
  • LCP, INP, and CLS;
  • network, JavaScript, rendering, memory, and layout cost;
  • images, fonts, lists, workers, caching, and third parties;
  • rendering topology and route-transition performance;
  • DevTools traces, User Timing, PerformanceObserver, and RUM;
  • budgets, regression detection, ownership, and performance culture.

By the end of today you can

  • explain what Core Web Vitals measure and what they do not;
  • use p75 and segmentation instead of hiding slow users in averages;
  • diagnose LCP resource discovery and server delay;
  • trace INP through input, main-thread work, rendering, and paint;
  • prevent CLS through reserved space and stable layout;
  • identify request waterfalls and unnecessary JavaScript;
  • optimize large lists and CPU work only when evidence supports it;
  • use browser tools and production signals together;
  • define route and user-journey budgets;
  • turn an optimization into a hypothesis, measurement, and regression guardrail.

The central principle

Performance work begins with a measured user problem, forms a hypothesis about its cause, and ends with a verified improvement under representative conditions.

An optimization is not successful because it sounds sophisticated.

It is successful when a user journey improves without unacceptable trade-offs.


The performance journey

flowchart TD
    A["1. User Request (Navigation / URL enter)"] --> B["2. Time to First Byte (TTFB - Server Response)"]
    B --> C["3. First Contentful Paint (FCP - Initial Typography/DOM)"]
    C --> D["4. Largest Contentful Paint (LCP - Hero Image/Header Rendered)"]
    D --> E["5. Interaction to Next Paint (INP - Responsive Main Thread)"]
    E --> F["6. Cumulative Layout Shift (CLS - Visual Stability Maintained)"]

Performance is a sequence of experiences, not one score.


Core Web Vitals

flowchart LR
    subgraph CoreWebVitals["The Three Core Web Vitals (Google Web Standards)"]
        LCP["Largest Contentful Paint (LCP)
Target: <= 2.5s (p75)
Measures: Perceived Loading Speed"]
        INP["Interaction to Next Paint (INP)
Target: <= 200ms (p75)
Measures: Overall Page Responsiveness"]
        CLS["Cumulative Layout Shift (CLS)
Target: <= 0.1 (p75)
Measures: Visual Stability & Jitter"]
    end

Together they cover important parts of the user journey.

They are not the complete product experience or business metric.


Why the 75th percentile matters

p75 = a value that 75% of observed experiences meet or beat

The worst quarter of experiences still matters.

An excellent average can hide slow devices, difficult networks, or specific routes with real users.


Field data versus lab data

Field dataLab data
real users and devicescontrolled environment
shows what happenedhelps explain why it may happen
segmented by route and conditionsrepeatable experiments
noisy but representativelimited but diagnostic

Use both; do not treat them as interchangeable.


Field data answers “what is happening?”

Field data can reveal:

  • a slow region or device class;
  • a route with poor p75;
  • a regression after release;
  • real-user impact of a third-party script;
  • differences by network, locale, or interaction.

It rarely identifies the exact code line by itself.


Lab data answers “why might it be happening?”

Lab tools can reveal:

  • request waterfalls;
  • long tasks;
  • layout calculations;
  • LCP resource delay;
  • JavaScript execution;
  • memory and rendering traces.

Use controlled experiments to test a specific hypothesis.


Synthetic testing needs representative conditions

Vary:

  • CPU speed;
  • network latency and bandwidth;
  • device memory;
  • viewport;
  • locale and direction;
  • route state;
  • cache state.

A single laptop and fast connection do not represent the user population.


Real User Monitoring

RUM can collect:

  • Web Vitals;
  • route and release;
  • device and connection class;
  • user journey markers;
  • errors and long tasks;
  • interaction context.

Collect only what is needed, protect privacy, and make the data actionable.


Segment performance data

Useful segments include:

route / release / device / network / locale / viewport / user state

A global p75 can improve while an important route or region regresses.


Largest Contentful Paint

LCP measures when the largest relevant content element becomes visible in the viewport.

It can be:

  • an image;
  • a text block;
  • a poster or background-like visual;
  • another large content element selected by the browser.

The metric is about the critical content experience, not only one file.


LCP is not just image download time

server response
  → resource discovery
  → request
  → download
  → decode
  → render and paint

An image can download quickly but still become LCP late because the page discovers or renders it late.


LCP subparts

Diagnose:

flowchart LR
    A["1. Time to First Byte
(Server & Network TTFB)"] --> B["2. Resource Load Delay
(Time until browser discovers LCP image)"]
    B --> C["3. Resource Load Duration
(Time to download image asset)"]
    C --> D["4. Element Render Delay
(Time spent decoding, layout & paint)"]

Each subpart suggests a different fix.

Optimizing the wrong subpart can change little.


Time to First Byte

TTFB includes:

  • connection setup;
  • redirects;
  • server processing;
  • time until the first response bytes.

Slow TTFB can delay every downstream rendering milestone.

Investigate server, deployment, cache, and network paths before optimizing browser code.


Redirects add delay

request A → redirect → request B → response

Remove unnecessary redirects from critical navigation and asset paths.

Check protocol, host, trailing slash, authentication, and locale redirects.


LCP resource discovery

The browser cannot request what it does not know exists.

An LCP image discovered only after:

  • JavaScript execution;
  • CSS evaluation;
  • a late component mount;

may arrive too late even if its file is optimized.


Prioritize the LCP resource selectively

Use:

  • semantic HTML;
  • early discoverability;
  • correct fetch priority;
  • selective preload;
  • appropriate caching.

Priority hints are not a substitute for a correct document and should not be applied to every image.


Do not lazy-load above-the-fold LCP images

Lazy loading tells the browser that a resource is not immediately important.

If that resource is the primary visible content, the hint contradicts the user experience.

Lazy-load content that is genuinely below the initial view.


Preload selectively

Preload can help when:

  • the resource is certain to be needed;
  • discovery is otherwise delayed;
  • it is part of the critical path.

It can hurt when it competes with CSS, fonts, scripts, or a different actual LCP resource.


Optimize LCP images

Consider:

  • correct dimensions;
  • modern format;
  • responsive sources;
  • compression quality;
  • CDN transformation;
  • caching;
  • avoiding unnecessary oversized files.

Image optimization must preserve visual quality and correct art direction.


Responsive images reduce waste

<img
  src="product-800.jpg"
  srcset="product-400.jpg 400w, product-800.jpg 800w"
  sizes="(max-width: 600px) 100vw, 50vw"
  alt="..."
>

Send an image appropriate to the viewport and layout rather than the largest available file to every device.


Image dimensions help CLS too

Reserve the rendered aspect ratio:

img {
  aspect-ratio: 4 / 3;
}

Known dimensions let the browser allocate space before the image arrives.


LCP text can be delayed by fonts

A text LCP may wait for:

  • font discovery;
  • font download;
  • font blocking behavior;
  • style calculation;
  • fallback-to-webfont swap.

Treat fonts as part of the critical rendering path when they affect the largest content.


Interaction to Next Paint

INP reflects how quickly the page produces the next visual update after user interactions across a session.

It considers more than the event handler’s own duration.


INP is a lifecycle metric

input → handler → rendering → paint

The slowest meaningful interaction can influence the session metric.

Optimize the user journey, not only one fast demo click.


Anatomy of an interaction

flowchart LR
    subgraph INP_Anatomy["Anatomy of an Interaction (INP Breakdown)"]
        I1["1. Input Delay
(Queued behind prior main-thread long tasks)"] --> I2["2. Processing Time
(Execution duration of event handlers)"]
        I2 --> I3["3. Presentation Delay
(Style recalc, layout, compositing & GPU paint)"]
    end

Any phase can dominate the perceived response.


Input delay

Input delay occurs when the main thread is busy before the event can be handled.

Possible causes:

  • startup JavaScript;
  • long tasks;
  • synchronous storage;
  • parsing or layout work;
  • another interaction’s work.

The handler code may be fast while the user still waits.


Main-thread contention

flowchart TD
    HTML["HTML Parsing"] & JS["Heavy JavaScript Execution"] & Style["Style Recalculations"] & Input["User Click / Keystroke"] --> MT["The Single Browser Main Thread"]
    MT --> Blocked["Long Task (>50ms)
Main thread frozen! User input delayed = High INP!"]

Work competes for responsiveness.

Reduce unnecessary work before moving necessary work to more complex mechanisms.


Long tasks

Long tasks block the main thread long enough to delay input and visual updates.

Find:

  • the task’s initiator;
  • script evaluation;
  • event handler work;
  • rendering and layout;
  • third-party contribution.

Then optimize the cause, not only the symptom.


Yielding

Break long work into opportunities for the browser to process input and paint.

await schedulerYieldOrTimeout();
processNextBatch();

Yielding improves responsiveness only when work can be safely divided and scheduled.


Avoid work before optimizing work

First ask:

  • is this calculation needed?
  • is it repeated?
  • is its input larger than necessary?
  • can it happen later?
  • can it be moved off the critical path?

Making unnecessary work faster is usually less valuable than removing it.


Event handler work

Keep urgent interaction work small:

read intent → update minimal state → schedule heavier work

Avoid synchronous parsing, large filtering, layout reads, and unrelated analytics in the immediate path.


Rendering can dominate INP

The handler may finish quickly, but a large update can cause:

  • broad component rendering;
  • large DOM changes;
  • style recalculation;
  • layout;
  • paint.

Trace from the input through the committed UI.


Reduce render scope

Keep state close to the consumers that need it.

Use stable identities and explicit boundaries to avoid updating unrelated parts of the interface.

Do not split components blindly; make the dependency graph narrower.


Virtualize large lists

Virtualization renders only the visible window plus a buffer.

flowchart LR
    Data["Dataset in Memory
(10,000 municipal records)"] --> Virtual["Virtual Scroller Window
(Calculates scroll offset)"]
    Virtual --> DOM["Lightweight DOM
(Only 35 active rendered <tr> elements!)"]
    DOM --> Smooth["60fps Butter-Smooth Scrolling
Zero memory bloat!"]

It can reduce rendering and layout cost for genuinely large collections.


Virtualization has UX costs

Consider:

  • keyboard navigation;
  • find-in-page;
  • screen readers;
  • variable row height;
  • scroll position;
  • focus preservation;
  • copy and selection behavior.

Virtualize when the trace shows a need and test the full interaction model.


Web Workers move CPU work off the main thread

Workers can help with:

  • large parsing;
  • data transformation;
  • computation;
  • search indexing;
  • image or file processing.

They do not make an algorithm cheaper and do not remove communication or serialization cost.


Worker communication has cost

sequenceDiagram
    autonumber
    participant Main as Browser Main Thread (UI at 60fps)
    participant Worker as Dedicated Web Worker Thread

    Main->>Worker: postMessage({ type: 'CALCULATE_AUDIT', data: largeMatrix })
    Note over Worker: Worker executes heavy 400ms calculation in background!
Main thread remains 100% responsive to user clicks.
    Worker-->>Main: postMessage({ type: 'AUDIT_COMPLETE', result: summary })
    Note over Main: Main thread updates UI badge with zero frame drops

Move enough work to justify the boundary.

Use transferable data or shared strategies deliberately and safely.


Interaction feedback first

The user needs an immediate response such as:

  • pressed state;
  • focus change;
  • optimistic visual update;
  • progress indication;
  • disabled or pending state.

Defer nonessential work so feedback is not behind logging, formatting, or secondary requests.


Cumulative Layout Shift

CLS measures unexpected layout movement during page use.

Common causes:

  • images without dimensions;
  • late banners;
  • font swaps;
  • inserted content;
  • changing ads or embeds;
  • transitions that move unrelated content.

Layout shift sources

Trace each shift to:

what moved?
what caused the new size?
was the change user-initiated?
was space reserved?

The fix depends on the cause, not only the visual symptom.


Reserve space

Reserve space for:

  • media;
  • ads or sponsored regions;
  • async controls;
  • embeds;
  • validation messages;
  • banners and notifications.

Stable geometry improves reading, focus, and interaction even beyond the metric.


User-initiated changes are different

A layout change directly caused by a user action may not count the same way as an unexpected shift.

It can still be a poor experience if it moves the user away from the control they are using.

Metrics do not replace design judgment.


Avoid inserting content above the view

If a late response inserts a banner above the user’s reading position, the page can jump.

Reserve the region or place updates where they do not displace current content unexpectedly.


Fonts and layout shift

Font changes can alter:

  • text width;
  • line wrapping;
  • block height;
  • button size;
  • layout position.

Choose fallback metrics and loading behavior that preserve useful geometry.


Font loading strategy

Options include:

  • system fonts;
  • self-hosted subsets;
  • preload for truly critical fonts;
  • font-display policy;
  • metric-compatible fallbacks;
  • delayed noncritical families.

The fastest font is often the one the critical path does not need.


font-display

Font display controls how text behaves while a webfont loads.

Choose based on:

  • content importance;
  • brand requirements;
  • fallback compatibility;
  • readability;
  • layout stability.

There is no single setting that is correct for every text region.


System fonts can be excellent

System fonts can provide:

  • immediate text;
  • no font transfer;
  • good platform integration;
  • stable performance.

Brand typography should justify its loading and layout cost.


Network performance

The network path includes:

DNS → connection → TLS → request → server → transfer → parse

Optimize the path that is actually slow.


Request waterfalls

flowchart LR
    A["1. HTML Document
(index.html)"] -->|Downloads & Parses| B["2. Client JS Bundle
(app.js)"]
    B -->|Executes & Fetches| C["3. API JSON Data
(/api/permits/104)"]
    C -->|Reads Image URL| D["4. LCP Hero Image
(hero.webp - Loaded at last!)"]

Sequential discovery delays the final content.

Find dependencies that could be:

  • discovered earlier;
  • requested in parallel;
  • embedded in the initial response;
  • cached or prefetched.

Flatten unnecessary waterfalls

const [products, categories] = await Promise.all([
  loadProducts(),
  loadCategories(),
]);

Parallelize independent work.

Do not parallelize operations that depend on one another or that would overwhelm the server.


Connection reuse

Reuse can reduce setup cost for:

  • HTTP connections;
  • TLS;
  • pooled server connections;
  • persistent sessions.

Avoid unnecessary origins that prevent effective reuse and increase discovery work.


Compression helps transfer, not execution

Compression reduces bytes over the network.

It does not remove:

  • JavaScript parsing;
  • compilation;
  • execution;
  • hydration;
  • memory;
  • DOM work.

Optimize both transfer and device work.


JavaScript is expensive in several ways

transfer → parse → compile → execute → render → memory

A small compressed bundle can still be expensive on a slower device if it contains heavy startup work.


Reduce unnecessary JavaScript

Possible strategies:

  • remove unused dependencies;
  • delay noncritical features;
  • keep static content server-rendered;
  • use route and component boundaries;
  • avoid shipping server-only logic;
  • replace a library with a platform capability when appropriate.

Measure the user journey after each change.


Route-level code splitting

home chunk
catalogue chunk
admin-editor chunk

Route splitting aligns code delivery with navigation and often gives a strong return.

Define loading, error, and prefetch behavior for each route.


Component-level lazy loading

Use it for:

  • rarely opened dialogs;
  • expensive editors;
  • below-the-fold visualizations;
  • optional reports;
  • feature-specific integrations.

Avoid splitting tiny, always-used components into network overhead.


Third-party JavaScript is shared product cost

Third-party code can add:

  • transfer;
  • CPU;
  • layout;
  • network requests;
  • privacy and security review;
  • failure dependencies.

The product owns the cost even when another company wrote the script.


Lazy-load third-party features

Delay analytics, chat, experiments, maps, and widgets when they are not needed for the first useful interaction.

Load them after consent, idle time, visibility, or explicit interaction where appropriate.


CSS performance

CSS affects:

  • style calculation;
  • layout;
  • paint;
  • render-blocking behavior;
  • asset discovery;
  • layout stability.

Keep critical styles available and avoid shipping large unused style systems to every route.


Critical CSS

Critical CSS covers styles needed for the initial visible content.

Possible approaches include:

  • inline critical rules;
  • route-specific CSS;
  • efficient stylesheet delivery;
  • deferred noncritical styles.

Avoid making the critical path more complex than the benefit justifies.


Avoid flashing unstyled content

Coordinate:

  • stylesheet discovery;
  • font behavior;
  • class application;
  • theme initialization;
  • server and client markup.

A fast first paint that changes dramatically is not a stable experience.


DOM size and depth

A large or deeply nested DOM can increase:

  • style calculation;
  • layout;
  • accessibility tree work;
  • memory;
  • query and update scope.

Do not reduce markup mechanically; identify the subtree causing measured work.


Batch reads and writes

Avoid alternating layout reads and writes:

flowchart TD
    subgraph LayoutThrashing["The Layout Thrashing Anti-Pattern"]
        R1["Read: elem1.offsetWidth"] --> W1["Write: elem1.style.width = '...'"]
        W1 -->|Forces immediate sync layout recalc!| R2["Read: elem2.offsetWidth"]
        R2 --> W2["Write: elem2.style.width = '...'"]
        W2 -->|Forces second layout recalc!| Bad["Result: 30 dropped frames / massive jank"]
    end
    subgraph BatchedSolution["The Batched Solution (Fast)"]
        BR1["Batch All Reads First:
r1 = elem1.offsetWidth;
r2 = elem2.offsetWidth;"] --> BW1["Batch All Writes in next frame:
elem1.style.width = ...;
elem2.style.width = ...;"]
        BW1 --> Good["Result: Exactly ONE clean layout recalculation!"]
    end

Group reads, then writes where possible to reduce forced layout and synchronization.


Paint and composite

Visual updates can cost through:

  • large paint regions;
  • expensive shadows or filters;
  • image decoding;
  • transparency and layers;
  • layout-triggering properties.

Use DevTools traces to identify actual paint and composite costs.


will-change is not a speed button

will-change can reserve resources or create layers.

Use it only for a measured, anticipated change and remove it when the change ends.

Overuse increases memory and can make rendering worse.


Animation frame budget

At a common 60Hz target, a frame arrives roughly every 16.7ms.

That time includes:

  • JavaScript;
  • style;
  • layout;
  • paint;
  • composite;
  • browser overhead.

Animation and interaction work must share the frame budget.


requestAnimationFrame

Use it to coordinate visual writes with the browser’s rendering cycle.

requestAnimationFrame(() => {
  element.style.transform = nextTransform;
});

It does not make expensive work free; it schedules it at a meaningful time.


Caching is a performance policy

Define:

  • what is immutable;
  • what can be stale;
  • what is user-specific;
  • what invalidates it;
  • how long it lives;
  • how it behaves offline.

Caching can reduce latency and increase correctness risk if the policy is unclear.


Fingerprinted static assets

app.8f31c.js
styles.2a1d.css

Content-addressed assets can be cached for a long time because a new content version receives a new URL.


HTML cache policy differs from hashed assets

HTML points to the current asset names and may need more frequent revalidation.

Caching HTML as aggressively as fingerprinted assets can serve old entry points or stale deployment manifests.

Match cache lifetime to artifact meaning.


API caching

API caching must consider:

  • freshness;
  • user identity;
  • permissions;
  • query inputs;
  • invalidation after mutation;
  • privacy;
  • offline usefulness.

The fastest response is not useful if it is the wrong user’s data.


Cache hit rate is not the whole story

Also measure:

  • stale response rate;
  • revalidation cost;
  • memory and storage;
  • cache eviction;
  • incorrect reuse;
  • time to useful content.

A high hit rate with stale or mismatched data is not a success.


Prefetching

Prefetch when a next action is likely and the cost is bounded.

Good signals include:

  • visible navigation intent;
  • hover or focus with caution;
  • route prediction;
  • idle time;
  • sufficient connection and storage.

Speculation can waste resources

Prefetching and prerendering can consume:

  • bandwidth;
  • battery;
  • memory;
  • server capacity;
  • privacy budget.

Do not optimize one user’s transition by imposing invisible cost on every user.


Preconnect and DNS prefetch

These hints can reduce connection setup for origins known to be needed.

Use them selectively:

  • preconnect for a critical known origin;
  • DNS prefetch when a connection may be useful later.

Unnecessary origins still add work and complexity.


Performance and rendering topology

CSR     → browser startup and execution cost
SSR     → server response plus hydration cost
SSG     → build-time work and static delivery
islands → selective browser responsibility

The rendering choice changes where performance work appears.


Hydration cost

Measure:

  • JavaScript needed for the route;
  • components that hydrate;
  • event setup;
  • client data handoff;
  • main-thread time before interaction.

HTML already visible does not mean the page is interactive.


Streaming performance

Streaming can improve:

  • first useful output;
  • perceived progress;
  • independent slow regions.

It can also add:

  • boundary complexity;
  • layout shifts;
  • coordination and error states;
  • client handoff work.

Measure the user journey, not only the first chunk.


SPA and soft-navigation performance

After initial load, measure:

  • route request time;
  • code chunk loading;
  • data loading;
  • transition feedback;
  • rendering and layout;
  • scroll and focus restoration.

A fast initial page can still have slow navigation.


Route-transition metrics

Define markers:

navigation requested
route code ready
data ready
content committed
interaction ready

Use them to locate the phase that users are waiting through.


User Timing API

performance.mark("catalogue-start");
await loadCatalogue();
performance.mark("catalogue-ready");
performance.measure("catalogue-load", "catalogue-start", "catalogue-ready");

Custom marks connect browser traces to product journeys.

Name them consistently and avoid collecting sensitive data.


PerformanceObserver

Observers can collect browser performance entries such as:

  • largest contentful paint;
  • layout shifts;
  • long tasks;
  • navigation timing;
  • resource timing.

Use them to build useful signals, not to collect every possible event without a question.


DevTools Performance panel

Start with a user action:

click / type / scroll / navigate
  → record trace
  → inspect main thread
  → correlate rendering and network

The trace should answer a hypothesis about the delay.


Flame charts

Flame charts reveal:

  • which functions consume time;
  • call depth;
  • repeated work;
  • long tasks;
  • layout and paint boundaries.

Read from the interaction or milestone rather than scanning for the largest colorful block without context.


Bottom-up analysis

Bottom-up views aggregate cost by function or category.

They can reveal:

  • a dependency’s shared cost;
  • repeated parsing;
  • expensive event handlers;
  • a framework operation called from many paths.

Use call stacks and user journeys to interpret the total.


Network panel

Inspect:

  • request order;
  • queueing;
  • connection setup;
  • response time;
  • transfer size;
  • cache status;
  • priority;
  • initiator chain.

Network evidence often explains why an apparently optimized asset still arrives late.


Coverage and bundle analysis

Coverage can show unused code during one journey.

Bundle analysis can show:

  • large dependencies;
  • duplicate packages;
  • shared chunk cost;
  • route chunk composition;
  • accidental client inclusion.

Unused in one journey is a clue, not automatic proof that code should be removed.


Lighthouse is one lab view

Lighthouse can provide useful repeatable diagnostics.

It is not:

  • real-user data;
  • a complete accessibility audit;
  • a business metric;
  • proof that every route is healthy;
  • a substitute for traces and production monitoring.

Use it as one instrument in the measurement stack.


Repeat measurements

Control:

  • cache state;
  • CPU and network;
  • viewport;
  • route data;
  • browser version;
  • device conditions.

Compare enough runs to distinguish signal from noise.


CPU and network throttling

Representative throttling helps reveal:

  • startup execution;
  • input delay;
  • slow resource discovery;
  • dependency waterfalls;
  • layout and paint pressure.

Do not mistake one artificial throttle for the entire user population.


Memory performance

Watch for:

  • detached DOM nodes;
  • unbounded caches;
  • retained event listeners;
  • subscriptions without cleanup;
  • large closures;
  • worker and image memory.

Memory problems can become interaction and crash problems later.


Memory leaks

A leak is retained data that should no longer be reachable.

Reproduce a lifecycle repeatedly:

mount → interact → unmount → repeat

Compare heap snapshots and retained references rather than relying on one memory reading.


Detached DOM nodes

Detached nodes may remain because:

  • a listener retains them;
  • a closure stores a reference;
  • a library cache never releases them;
  • a subscription outlives the component.

Clean up at the owner boundary.


Unbounded caches

Every cache needs:

  • maximum size or expiry;
  • invalidation;
  • identity scope;
  • persistence policy;
  • behavior under storage pressure.

An optimization that grows without bound is a future outage.


Subscription cleanup

Clean up:

  • event listeners;
  • observers;
  • timers;
  • sockets;
  • workers;
  • requests;
  • media resources.

The cleanup should belong to the lifecycle that created the subscription.


Performance budgets

A budget turns an aspiration into a reviewable constraint.

Possible budgets include:

  • initial transfer;
  • JavaScript execution;
  • route transition;
  • LCP, INP, and CLS;
  • memory;
  • third-party cost.

Budgets need context

A single global number hides route and user differences.

Define budgets by:

route × device class × network × journey

Keep the budget small enough to guide decisions and flexible enough to reflect product reality.


Bundle budgets

Track:

  • initial compressed transfer;
  • initial uncompressed parse volume;
  • route chunk sizes;
  • third-party additions;
  • duplicate dependencies.

Bundle size is a useful constraint, not the complete performance outcome.


Metric budgets

Metrics can be reviewed by:

  • route;
  • release;
  • device segment;
  • user journey;
  • p75 or other percentile.

Define what action occurs when the budget is exceeded.


User-journey budgets

open catalogue → see products → filter → open detail → save

The journey budget can include:

  • useful content;
  • interaction readiness;
  • route transition;
  • server confirmation;
  • recovery after failure.

It connects performance to product value.


Performance regression

A regression is meaningful when:

  • a measured signal worsens;
  • under comparable conditions;
  • for a meaningful segment;
  • beyond expected noise;
  • with an owner and response path.

Do not fail a build over a noisy metric without understanding variance.


Performance ownership

Assign responsibility for:

  • budgets;
  • monitoring;
  • dependency review;
  • route regressions;
  • third-party approval;
  • incident response;
  • optimization follow-up.

Performance that belongs to nobody becomes cleanup work after users complain.


Performance review in pull requests

Ask when a change affects:

  • initial route imports;
  • large lists;
  • images or fonts;
  • third-party scripts;
  • rendering scope;
  • network waterfalls;
  • cache behavior;
  • memory lifecycle.

Not every PR needs a benchmark, but every costly change needs a reasoning path.


Performance and accessibility

Do not optimize by removing:

  • labels;
  • focus behavior;
  • semantic structure;
  • keyboard support;
  • announcements;
  • readable loading states.

A fast interface that cannot be used is not a performance success.


Performance and internationalization

Localization can change:

  • text length;
  • line wrapping;
  • font coverage;
  • direction;
  • date and number formatting;
  • layout and interaction size.

Measure representative locales and directions, not only the default language.


Font subsetting and multilingual products

Subset fonts by language or route when appropriate.

Verify:

  • fallback behavior;
  • glyph coverage;
  • layout stability;
  • preload scope;
  • cache reuse.

Reducing font bytes should not create missing characters or unstable text.


Performance and security

Security controls can affect performance:

  • CSP and third-party loading;
  • integrity checks;
  • authentication redirects;
  • encryption and headers;
  • sanitization;
  • logging and monitoring.

Do not remove a security boundary for a benchmark win. Find the right design trade-off.


Performance and offline architecture

Offline systems can improve perceived speed with local reads.

They also add:

  • storage work;
  • synchronization;
  • reconciliation;
  • memory and disk cost;
  • stale-data decisions.

Measure both online first load and offline/restore journeys.


Performance and micro-frontends

Micro-frontends can add:

  • remote requests;
  • duplicate runtimes;
  • mount work;
  • CSS and font duplication;
  • integration waterfalls.

They can also isolate route loading and reduce one giant application bundle.

Measure the assembled page and route transition.


Diagnose LCP

Ask in order:

  1. Is TTFB slow?
  2. Is the LCP resource discovered late?
  3. Is the resource too large or poorly prioritized?
  4. Is rendering blocked by fonts or JavaScript?
  5. Is the final element delayed by layout or hydration?

Match the fix to the slow subpart.


Diagnose INP

Ask:

  1. Was the input delayed by a long task?
  2. Is the event handler doing unnecessary work?
  3. Does state update too broad a region?
  4. Is rendering or layout expensive?
  5. Can work be deferred, batched, or moved?

The answer may be outside the handler itself.


Diagnose CLS

Ask:

  1. Which element moved?
  2. What changed its geometry?
  3. Was space reserved?
  4. Did a font, image, ad, or async message arrive late?
  5. Was the change necessary and user-initiated?

Fix the layout model, not only the visible jump.


Optimization should match cause

slow server → cache, data path, server work
late image  → discovery, priority, dimensions
slow input  → main-thread or render scope
layout jump → reserved geometry and font policy
large route → imports, splitting, dependency review

Avoid applying a favorite optimization to every symptom.


Performance triage

Prioritize by:

  • user impact;
  • affected population;
  • severity;
  • confidence in the cause;
  • cost and risk of the fix;
  • reversibility;
  • business importance of the journey.

The biggest measured cause is often more valuable than a long list of minor improvements.


Performance hypotheses

Use a structured statement:

We believe [cause] makes [journey] slow for [segment].
If we change [intervention], [signal] should improve without [trade-off].

Then measure before and after under comparable conditions.


Before and after measurement

Record:

  • baseline;
  • environment;
  • route and data;
  • change;
  • expected signal;
  • observed result;
  • side effects;
  • rollback or follow-up.

An optimization without a baseline is a story, not evidence.


Avoid benchmark theater

Do not:

  • optimize a synthetic score disconnected from users;
  • report one best run;
  • hide slow segments;
  • compare different route states;
  • claim universal frame rates;
  • remove accessible behavior to win a metric.

Performance evidence should improve decisions, not decorate a release.


Production verification

After deployment, verify:

  • field metrics;
  • route segments;
  • release comparison;
  • error and abandonment signals;
  • cache behavior;
  • device and network impact.

Lab success does not prove production success.


Core Web Vitals are not business metrics

They describe aspects of experience.

Business outcomes may include:

  • task completion;
  • search success;
  • conversion;
  • retention;
  • support contacts;
  • revenue or cost.

Connect performance improvements to meaningful user outcomes.


Core Web Vitals are not the whole experience

Also consider:

  • accessibility;
  • correctness;
  • content clarity;
  • responsiveness outside measured interactions;
  • resilience and offline behavior;
  • perceived progress;
  • privacy and security.

Optimize the experience, not only the dashboard.


A performance measurement stack

flowchart TD
    A["1. Browser Native Observers
(PerformanceObserver: LCP, INP, CLS)"] --> B["2. Custom User Timing API
(performance.mark / performance.measure)"]
    B --> C["3. Real User Monitoring (RUM) Beacon
(navigator.sendBeacon to Telemetry API)"]
    C --> D["4. Metric Aggregation & p75 Analysis
(Segmented by device tier, connection, country)"]
    D --> E["5. Performance Budget Alerts & Regression CI Gates"]

Each layer answers different questions.


Performance dashboard design

A useful dashboard shows:

  • route and release;
  • p75 and distribution;
  • device/network segment;
  • sample size;
  • metric subparts where available;
  • business journey;
  • regression threshold;
  • owner and next action.

Avoid a single score that hides all context.


Percentiles beyond p75

p75 is useful for common experience reporting.

p90, p95, and tail behavior can reveal severe experiences for a smaller group.

Choose percentiles based on the decision and sample size; do not compare tiny samples as if they were stable populations.


Sample size matters

A metric from a handful of sessions can move dramatically by chance.

Use:

  • confidence intervals or uncertainty awareness;
  • minimum sample thresholds;
  • longer windows for sparse routes;
  • segmentation that does not destroy statistical usefulness.

Performance budgets by route

Different routes may need different budgets:

marketing page → initial load and LCP
catalogue      → filtering and transition INP
editor         → input responsiveness and save flow
reports        → data and visualization cost

Route budgets align architecture with user journeys.


Component performance contracts

A shared component can define expectations for:

  • DOM size;
  • interaction work;
  • image behavior;
  • hydration or client code;
  • accessibility;
  • cleanup;
  • bundle contribution.

Contracts should guide design without making every component artificially identical.


Performance and architecture decisions

Before choosing:

  • CSR or SSR;
  • monolith or micro-frontend;
  • client or server component;
  • eager or lazy feature;
  • cache or refetch;
  • virtualized or full list;

state the user journey, bottleneck, measurement, and trade-off.


A performance culture

Healthy teams:

  • measure before optimizing;
  • share traces and context;
  • review performance impact early;
  • protect accessibility and security;
  • assign ownership;
  • learn from production data;
  • keep budgets actionable.

Performance is a continuous engineering property, not a final cleanup phase.


Practical lab: Measure and Improve a Slow Interface

Diagnose a slow catalogue using field-like and lab measurements, then test whether virtualization, scheduling, asset changes, or caching address the measured cause.

Every change must begin with a hypothesis and end with a representative re-measurement.


Practical stages 1–5: loading diagnosis

  1. Record LCP, INP, and CLS baselines.
  2. Capture a trace for a slow interaction.
  3. Identify long tasks, layout work, network delays, and rendering cost.
  4. Inspect LCP.
  5. Optimize hero delivery and test a resource-priority experiment.

Do not preload or optimize blindly.


Practical stages 6–9: interaction work

  1. Diagnose INP.
  2. Remove unnecessary work.
  3. Virtualize a large list only if the trace supports it.
  4. Move CPU work to a worker.

Compare responsiveness, memory, communication cost, accessibility, and implementation complexity.


Practical stages 10–13: layout and network

  1. Diagnose CLS.
  2. Test font strategy.
  3. Inspect network waterfalls.
  4. Remove an artificial waterfall.

Re-measure under representative CPU, network, locale, and direction settings.


Practical stages 14–18: code and navigation

  1. Inspect JavaScript.
  2. Audit third-party scripts.
  3. Add caching.
  4. Measure an SPA route transition.
  5. Add User Timing.

Connect each change to a specific journey and signal.


Practical stages 19–24: production discipline

  1. Investigate memory.
  2. Create performance budgets.
  3. Add a CI guardrail.
  4. Build a RUM design.
  5. Build a performance regression report.
  6. Draw the final performance architecture.

Verification: the optimization improves a user journey or validated signal; no universal “zero INP” promise is made.


Practical extension: budget and regression test

Add a performance budget and regression test.

Document why the budget is useful and why it is not a replacement for Core Web Vitals, field data, or user-journey measurement.


Try this yourself

Choose one slow interaction and write:

segment:
baseline:
hypothesis:
trace evidence:
change:
expected trade-off:
after measurement:

If you cannot fill in the evidence, measure before changing code.


Troubleshooting guide (Part 1)

SymptomLikely cause
Lighthouse is good but users are slowLab conditions or route differ from field data
LCP image is compressed but still lateDiscovery, TTFB, priority, or render delay
Click handler is short but INP is poorInput delay or expensive rendering follows
CLS occurs after font loadMetrics and fallback geometry differ
Virtualization improves speed but breaks keyboard useUX and accessibility contract was not tested

Troubleshooting guide (Part 2)

SymptomLikely cause
Worker adds no benefitTransfer and setup cost exceed CPU savings
Bundle shrinks but interaction is unchangedExecution, rendering, or network is the actual bottleneck
Cache hit rate rises but data is wrongFreshness, identity, or invalidation policy is weak
Budget fails randomlySample, environment, or threshold is not controlled

Completion checklist

  • field and lab data are used for different questions;
  • p75 and segments are considered;
  • LCP is diagnosed by subparts;
  • INP includes input, handler, rendering, and paint;
  • CLS sources have reserved geometry;
  • network and JavaScript waterfalls are inspected;
  • list and worker optimizations are evidence-driven;
  • memory, cache, and subscription lifecycles are bounded;
  • route budgets and performance ownership are explicit;
  • production verification follows lab experiments.

Misconceptions to leave behind (Part 1)

MisconceptionBetter mental model
Performance means Lighthouse scoreIt is a measured user experience
Fast on my machine means fastDevice, network, and route segments differ
Average performance is enoughPercentiles reveal slow experiences
LCP is image download timeDiscovery, server, download, and render all matter
Every above-fold image should be preloadedPrioritize the actual critical resource
Lazy loading always helpsIt can delay content the user needs now
INP is only handler durationInput, processing, render, and paint form the interaction
Memoization always improves speedCaches have cost and may not target the cause

Misconceptions to leave behind (Part 2)

MisconceptionBetter mental model
Virtualize every listVirtualization has UX and accessibility trade-offs
Workers make code fasterThey trade main-thread work for communication cost
Compression solves JavaScript bloatParsing and execution remain
More code splitting is always betterChunks have request and coordination cost
Prefetch everythingSpeculation consumes user and server resources
SSR guarantees good Web VitalsTopology moves work; it does not remove it
Optimization is final polishPerformance is an architectural constraint

The chapter in one sentence

Measure the real journey, diagnose the actual bottleneck, and make the smallest evidence-backed change that improves users without sacrificing accessibility, security, or correctness.


Next: Chapter 16

The next chapter will build on performance engineering with:

  • observability and production diagnostics;
  • logging, tracing, and error reporting;
  • reliability signals and incident response;
  • operational feedback for front-end systems;
  • architecture that remains explainable in production.

Questions

Which user journey is slow for a real segment of your users, and what evidence would distinguish network, JavaScript, rendering, layout, and server causes?

16 - Testing Strategies for Resilient Interfaces

Chapter 16: build layered confidence around user behavior, accessibility, network boundaries, failure recovery, and production workflows.

Testing Strategies for Resilient Interfaces

Build confidence around real behavior

Chapter 16

Polla Fattah


Today’s goal

Design a test strategy that finds important failures without coupling every test to implementation details.

We will connect:

  • risk, evidence, and the testing pyramid;
  • static analysis, unit, component, integration, and E2E tests;
  • semantic queries, accessible names, keyboard, and focus behavior;
  • async UI, network mocking, cancellation, and races;
  • Playwright journeys, visual regression, contracts, and browser matrices;
  • offline, cache, routing, permissions, security, and performance tests;
  • flakiness, CI selection, coverage, maintenance, and test architecture.

By the end of today you can

  • choose a test level from the risk and boundary involved;
  • test user-visible behavior without inspecting private state unnecessarily;
  • use semantic queries and accessible names effectively;
  • distinguish simulated DOM tests from real-browser tests;
  • test loading, empty, error, retry, cancellation, and optimistic rollback;
  • mock at stable boundaries rather than mocking every internal module;
  • design realistic browser journeys and failure artifacts;
  • detect flakiness instead of normalizing it;
  • use coverage as a map rather than a grade;
  • maintain a layered suite that remains useful as the UI evolves.

The central principle

A test is valuable when it provides trustworthy evidence about a risk the product actually has.

Fast tests, realistic tests, and broad tests each provide different evidence.

Quality comes from a balanced set of boundaries - not from a single test type or coverage percentage.


Testing is risk management

Ask:

  • what can fail?
  • who is affected?
  • how likely is it?
  • how costly is detection after release?
  • which test level observes the risk most directly?

The answer determines where to invest test effort.


Confidence comes from different evidence

flowchart TD
    SA["Static Checks<br/>(Types, Linters, Schema)"] --> UT["Unit Tests<br/>(Pure Logic & Reducers)"]
    UT --> CT["Component Tests<br/>(DOM Semantics & Events)"]
    CT --> IT["Integration Tests<br/>(Boundaries & State Flow)"]
    IT --> E2E["E2E Tests<br/>(Real Browser Journeys)"]
    E2E --> FS["Field Signals<br/>(RUM & Telemetry)"]

No one layer can prove the others.


Testing pyramid, trophy, and reality

The shape is less important than the reasoning:

  • fast checks should catch cheap mistakes early;
  • integration tests should cover important boundaries;
  • E2E tests should protect critical journeys;
  • production signals should reveal what test environments missed.

Choose the mix from risk, not from a diagram’s proportions.


The cost-confidence spectrum

flowchart LR
    A["Static Checks<br/>Lowest Cost / Narrow"] --> B["Unit Tests"]
    B --> C["Component Tests"]
    C --> D["Integration Tests"]
    D --> E["Browser Journeys<br/>Highest Cost / Broad"]

Use the narrowest test that provides enough confidence for the risk.

Put a failure at the closest useful boundary.


Static analysis is part of testing strategy

Static checks can catch:

  • type inconsistencies;
  • unreachable branches;
  • unsafe imports;
  • dependency direction violations;
  • accessibility lint issues;
  • unused or suspicious code.

They are fast evidence, but they do not observe runtime behavior.


Static analysis has limits

It cannot fully prove:

  • server responses;
  • browser timing;
  • focus behavior;
  • visual layout;
  • network failures;
  • authentication flows;
  • user comprehension;
  • correct business outcomes.

Use runtime tests where the risk exists.


Unit tests target deterministic logic

expect(formatPrice(1250, "USD")).toBe("$12.50");

Good unit targets include:

  • parsers;
  • reducers;
  • formatters;
  • validation;
  • URL serialization;
  • cache-key construction;
  • state-transition logic.

Avoid unit testing language syntax

Do not write tests to prove that:

  • Array.prototype.map maps;
  • an if statement branches;
  • a framework renders a known element;
  • a constant equals itself.

Test your decision logic and product behavior around the language feature.


Test behavior, not line count

user input → validation message → submit disabled or request sent

This is more valuable than asserting that every internal line executed.

Coverage can identify untested areas, but behavior determines confidence.


Component tests should resemble user interaction

find searchbox → type query → press Enter → observe loading/results

Prefer interactions and visible outcomes over calls to private helpers or internal state snapshots.


Why role-based queries are valuable

screen.getByRole("button", { name: /save/i });

Role-based queries:

  • reflect the accessibility tree;
  • survive many visual refactors;
  • encourage meaningful semantics;
  • match how assistive technologies identify controls.

They are a quality signal, not complete accessibility certification.


Accessible name is part of the contract

<button aria-label="Close dialog">×</button>

The role identifies the control type.

The accessible name identifies what it does.

Test both when the behavior depends on them.


Accessible name computation is a web concept

Accessible names can come from:

  • visible text;
  • associated labels;
  • aria-label;
  • aria-labelledby;
  • other standard mechanisms.

Testing-library queries reflect platform semantics; they do not invent a private testing concept.


Preferred query order

Often prefer:

  1. role and accessible name;
  2. label;
  3. placeholder or visible text where appropriate;
  4. semantic state or value;
  5. test ID as a fallback.

Choose the query that represents the user-facing contract.


getByLabelText for forms

screen.getByLabelText("Email address");

This tests that the label and control are associated in a meaningful way.

It can reveal form accessibility problems that a class selector would miss.


Role queries do not replace accessibility audits

A control can be findable by role and still have:

  • poor focus behavior;
  • incorrect state announcements;
  • bad color contrast;
  • keyboard traps;
  • confusing reading order;
  • missing error association.

Automated checks supplement human and assistive-technology evaluation.


Test IDs are legitimate fallbacks

Use a test ID when:

  • no meaningful user-facing semantic exists;
  • a structural target is required;
  • a complex visualization needs a stable anchor;
  • the stronger query would be misleading.

The issue is not the attribute; it is using it to avoid designing a semantic interface.


Avoid overfitting to CSS selectors

container.querySelector(".blue-primary-button");

This couples the test to styling and class names.

Use CSS selectors when CSS structure is genuinely the behavior under test, such as a visual or layout integration.


Avoid testing internal state

Prefer:

click tab → selected tab is announced and panel changes

Over:

expect(component.state.selected).toBe("reviews")

Internal state is an implementation choice; visible behavior is the contract.


Avoid testing private methods

Private methods can change during a refactor while the product behavior remains correct.

Test them indirectly through the public behavior unless the logic is extracted into a meaningful pure unit with its own contract.


Component test example

it("shows an error when the email is invalid", async () => {
  await user.type(screen.getByLabelText("Email"), "not-an-email");
  await user.click(screen.getByRole("button", { name: "Save" }));

  expect(screen.getByRole("alert")).toHaveTextContent("valid email");
});

The test follows interaction and outcome rather than implementation.


Integration tests connect boundaries

Integration may include:

  • component plus reducer;
  • form plus validation;
  • query layer plus cache;
  • route plus URL state;
  • API adapter plus parser;
  • shell plus remote contract.

The scope should reflect a real interaction boundary.


Integration is a spectrum

flowchart LR
    M["Two Collaborating Modules"] --> FB["Feature Boundary"]
    FB --> RT["Route & URL State"]
    RT --> AW["Application Workflow"]

Name the scope clearly.

The goal is not to make every test “large”; it is to verify the collaboration that matters.


Browser simulation versus a real browser

Simulated environmentReal browser
fast and focusedrealistic layout and platform behavior
easy unit/integration loopnetwork, focus, CSS, storage, workers
incomplete browser APIshigher cost and setup
good for most logicneeded for critical browser behavior

Use both intentionally.


Vitest and test-runner responsibilities

A runner typically provides:

  • test discovery;
  • assertions;
  • isolation;
  • mocks and timers;
  • watch mode;
  • coverage;
  • reporting.

It does not decide whether the tested behavior is valuable.


Watch mode supports rapid feedback

Watch mode should:

  • rerun affected tests;
  • preserve readable failure output;
  • make focused testing easy;
  • encourage small feedback loops.

Do not let a slow or noisy watch workflow push developers toward skipping tests.


Browser mode

Browser-based component tests can reveal:

  • real DOM behavior;
  • CSS and layout interaction;
  • focus and selection;
  • browser API differences.

Use them where simulated DOM limitations are relevant rather than moving every unit test into a browser.


Simulated DOM still has value

It is often sufficient for:

  • roles and labels;
  • event flows;
  • state transitions;
  • loading and error rendering;
  • form validation;
  • network boundary behavior.

Know what the environment does not implement before trusting it for a browser-specific claim.


User event simulation

Prefer realistic sequences over synthetic dispatch:

flowchart LR
    F["focus"] --> KD["keydown"]
    KD --> IN["input"]
    IN --> CH["change"]
    CH --> BL["blur"]

A single direct property assignment may skip behavior your product depends on.

Use a user-event tool when the interaction sequence matters.


Keyboard testing

Test:

  • Tab order;
  • Enter and Space activation;
  • arrow-key navigation;
  • Escape behavior;
  • focus after open and close;
  • disabled controls;
  • keyboard traps.

Mouse-only tests miss a major part of interaction architecture.


Focus testing

await user.click(screen.getByRole("button", { name: "Open" }));
expect(screen.getByRole("dialog")).toHaveFocus();

Focus is observable user state.

Test where focus goes, what happens on dismissal, and whether the trigger is restored.


Testing async UI

Cover asynchronous state transitions:

stateDiagram-v2
    [*] --> Idle
    Idle --> Loading: User Trigger / Fetch
    Loading --> Success: 200 OK Response
    Loading --> Empty: Zero Results Found
    Loading --> Error: Network / 500 Fail
    Error --> Loading: Retry Action
    Success --> [*]
    Empty --> [*]

Assert after the relevant state settles, not immediately after starting an asynchronous operation.


Avoid arbitrary sleeps

await new Promise(resolve => setTimeout(resolve, 1000));

Sleeps make tests slow and still do not prove the condition is ready.

Wait for a meaningful outcome, event, network response, or visible state.


Assertions should match user outcomes

Prefer:

retry button appears
previous results remain visible
error is associated with the field

Over asserting:

setTimeout was called once

Unless the timer itself is the contract, test what the user experiences.


Mock at stable boundaries

Good boundaries include:

  • network API;
  • clock;
  • storage adapter;
  • repository;
  • browser capability;
  • feature flag provider.

Mocking at a stable boundary keeps tests aligned with architecture.


Mocking internal modules can over-couple tests

If every internal import is mocked, a refactor changes hundreds of tests without changing behavior.

Mock only where the boundary is expensive, nondeterministic, or outside the test’s responsibility.


Network mocking

Network mocks should model:

  • status codes;
  • response shapes;
  • latency where relevant;
  • malformed payloads;
  • cancellation;
  • retries;
  • partial failure.

Mock the protocol the feature depends on, not a convenient internal helper.


Mock Service Worker style

An interception layer can let application code use real request paths while tests control responses.

This preserves more of the transport boundary than mocking fetch() in every module.

Use realistic response contracts and reset handlers between tests.


Test success, error, and empty states

For remote features, cover:

loading / success / empty / stale / error / retry

The happy path is one state among several that users will encounter.


Test cancellation and races where relevant

sequenceDiagram
    autonumber
    participant UI as Search Input
    participant Net as Network Boundary
    participant Srv as Server
    UI->>Net: Query "erb" (Request 1)
    UI->>Net: Query "erbil" (Request 2)
    Note over Net: Request 1 delayed (800ms)
    Net->>Srv: Process Request 2
    Srv-->>UI: Results for "erbil" (Arrives at 200ms)
    Srv-->>UI: Results for "erb" (Arrives at 800ms - LATE!)
    Note over UI: Race Bug: Stale results clobber newest search!

The test should prove that Query 2 remains authoritative and Request 1 was aborted.


Demonstration: Failing race condition

// User types rapidly: "erb" then "erbil"
// Without AbortController, Request 1 resolves after Request 2:
test("demonstrates race condition failure", async () => {
  render(<LiveSearch />);
  await user.type(screen.getByRole("searchbox"), "erb");
  await user.type(screen.getByRole("searchbox"), "il");
  // If the component lacks cancellation, delayed response for "erb"
  // overwrites the newer "erbil" results in the DOM!
  expect(screen.getByRole("searchbox")).toHaveValue("erbil");
  // FAILS: DOM displays items for "erb" instead of "erbil"
  expect(await screen.findByText("Erbil International Airport")).toBeInTheDocument();
});

A test that does not control network timing will never detect this intermittent race.


Demonstration: Setting up the race test

Simulate network latency on the first query using MSW:

const heldResolvers: Array<() => void> = [];
server.use(
  http.get("/api/search", ({ request }) => {
    const q = new URL(request.url).searchParams.get("q");
    if (q === "erb") {
      // Hold Request 1 until manually released
      return new Promise(r => heldResolvers.push(() => 
        r(HttpResponse.json([{ id: 1, name: "Old Erbil Entry" }]))
      ));
    }
    return HttpResponse.json([{ id: 2, name: "Erbil International Airport" }]);
  })
);

Demonstration: Asserting race resilience

Trigger rapid typing and verify late responses are ignored:

render(<LiveSearch />);

// User types "erbil" - Request 2 resolves quickly
await user.type(screen.getByRole("searchbox"), "erbil");
expect(await screen.findByText("Erbil International Airport")).toBeInTheDocument();

// Resolve delayed Request 1: must NOT clobber current UI
heldResolvers[0]?.();
expect(screen.queryByText("Old Erbil Entry")).not.toBeInTheDocument();

The test controls network timing at the transport boundary.


What the passing race test proves

Passing the mocked component race test validates:

  • Order-independence: delayed earlier requests do not overwrite newer state;
  • DOM accuracy: rendered search results reflect the active query;
  • Cancellation: AbortController dispatches signal on new keystrokes.

The contract between input and view is verified.


What the passing test still cannot prove

Even with the component test green, it cannot prove:

  • Backend capacity: server rate-limiting under concurrent queries;
  • Screen-reader cadence: aria-live speech queue congestion;
  • Device rendering: frame drops on low-tier mobile hardware;
  • Input methods: IME Arabic/CJK composition events.

Confidence requires complementary evidence across layers.


Fake timers

Fake timers help test:

  • debounce;
  • retry backoff;
  • polling;
  • expiration;
  • scheduled UI work.

Advance time deliberately and restore real timers after each test.

Do not use fake time to hide a missing synchronization condition.


Test data builders

const product = buildProduct({ priceCents: 1250, inStock: true });

Builders provide valid defaults and make the relevant variation visible.

They reduce repetitive fixture setup without hiding important input differences.


Avoid magic fixtures

A giant fixture can make every test depend on irrelevant fields.

Prefer small, intention-revealing builders and explicit edge-case data.

The test should show why the data matters.


End-to-end testing

E2E tests verify a deployed-like journey through:

  • browser;
  • route;
  • application code;
  • server or controlled backend;
  • network and storage;
  • visible user outcomes.

They provide broad confidence at higher cost.


Playwright locators

Prefer locators that reflect user semantics:

page.getByRole("button", { name: "Save" });
page.getByLabel("Email");

Use stable test IDs when semantic locators do not describe the target.


Auto-waiting

A browser tool can wait for:

  • visibility;
  • enabled state;
  • attachment;
  • navigation;
  • assertions.

Use built-in waiting and web-first assertions rather than adding arbitrary sleeps.


Web-first assertions

await expect(page.getByRole("heading", { name: "Catalogue" }))
  .toBeVisible();

Assertions should wait for the user-visible condition and produce useful failure output.


Test real user journeys

Good E2E candidates include:

  • sign-in and redirect;
  • search and filter;
  • edit and save;
  • error and retry;
  • offline draft recovery;
  • back and forward navigation;
  • critical purchase or submission.

Do not use E2E to cover every small rendering branch.


The critical-path suite

Keep a small, reliable suite for:

open → act → submit → confirm

It should run quickly enough to block unsafe releases and produce actionable artifacts on failure.


E2E data isolation

Tests should use:

  • isolated accounts or tenants;
  • unique records;
  • resettable data;
  • explicit cleanup;
  • deterministic server state.

Shared mutable test data creates order dependence and flakiness.


Test setup through APIs when appropriate

Create test state through a supported API or fixture layer when the behavior under test is not the setup flow.

Then use the UI for the actual journey.

Do not bypass the behavior you are trying to verify.


E2E authentication

Choose between:

  • UI login for the login journey;
  • API or storage setup for unrelated tests;
  • reusable authenticated contexts with isolated identities.

Document which boundary each test actually verifies.


Multi-browser and mobile testing

Test a matrix based on risk:

  • supported browser engines;
  • viewport classes;
  • touch and keyboard behavior;
  • locale and direction;
  • critical layout and input paths.

Do not run every test on every browser without a reason; do not test only one browser when compatibility matters.


Accessibility-oriented testing

Combine:

semantic queries
automated rules
keyboard journeys
focus checks
human review
screen-reader testing

Each catches different failures.


Automated accessibility checks

Automated checks can find:

  • missing labels;
  • invalid ARIA relationships;
  • contrast issues in some cases;
  • landmark and structural problems.

They cannot fully evaluate meaning, workflow, focus strategy, or real assistive-technology experience.


Visual regression testing

Visual tests are useful for:

  • design-system components;
  • themes;
  • responsive layouts;
  • typography;
  • overlays;
  • cross-route visual contracts.

They should complement behavior tests, not replace them.


Visual regression noise

Control:

  • browser version;
  • fonts;
  • viewport;
  • animations;
  • network content;
  • time and locale;
  • screenshot masking.

Review visual changes as product changes, not blindly approve every diff.


Snapshot testing

Snapshots can reveal broad output changes quickly.

They become low-value when:

  • output is huge;
  • reviewers approve without reading;
  • implementation details dominate;
  • snapshots replace focused assertions.

Keep snapshots small and meaningful.


Contract tests

Contract tests verify an agreement at a boundary:

  • API request and response schema;
  • event payload;
  • package public API;
  • micro-frontend mount contract;
  • generated client assumptions.

Test the boundary the consumer actually depends on.


Schema-driven testing

A runtime schema can support:

  • response validation;
  • generated examples;
  • malformed-input tests;
  • compatibility checks;
  • property-based variations.

Shared schemas improve agreement but do not remove the need to test real delivery and failure.


Test network conditions

Include scenarios such as:

  • slow response;
  • offline;
  • timeout;
  • aborted request;
  • malformed response;
  • 401/403;
  • 404;
  • 422;
  • 500;
  • partial dashboard failure.

Resilient UI is defined by recovery behavior.


Test offline behavior where promised

Verify:

  • draft survives reload;
  • queued work is visible;
  • retry has a limit;
  • permanent failure is recoverable;
  • reconnection does not duplicate work;
  • stale data is labeled appropriately.

Do not claim offline support from an offline screen alone.


Flaky tests are signals

Flakiness can indicate:

  • uncontrolled time;
  • shared state;
  • races;
  • missing awaits;
  • unstable selectors;
  • order dependence;
  • real product nondeterminism.

Treat it as a defect in the test or system until understood.


Never normalize flakiness

“It passes on retry” hides:

  • release risk;
  • missing confidence;
  • slow CI;
  • developer distrust;
  • real timing bugs.

Quarantine only with an owner, issue, scope, and removal plan.


Retries are not a fix

Retries may help identify intermittent infrastructure failures.

They can also:

  • hide races;
  • double mutations;
  • make failures slower;
  • produce false green builds.

Use retries as a diagnostic or bounded operational policy, not as proof of reliability.


Trace artifacts improve failure diagnosis

Keep useful artifacts such as:

  • screenshots;
  • video;
  • browser trace;
  • console output;
  • network logs;
  • server correlation IDs;
  • DOM snapshots where safe.

Artifacts should help reconstruct the user journey without leaking sensitive data.


Failure messages are part of test quality

A useful failure says:

  • which journey failed;
  • what state was expected;
  • what state appeared;
  • which request or boundary was involved;
  • where artifacts are stored.

“Expected true to be false” is rarely enough for a team to act quickly.


Test names are documentation

shows previous results while a refresh fails

is more useful than:

handles error

Name the behavior and the important condition.


Arrange–Act–Assert

flowchart LR
    ARR["Arrange<br/>Establish state & mock boundaries"] --> ACT["Act<br/>Perform meaningful user interaction"]
    ACT --> AST["Assert<br/>Verify user-visible outcomes & semantics"]

Keep the structure readable, even when a test uses several assertions.


Given–When–Then

Given an expired session
When the user submits the form
Then the app asks them to sign in and preserves the draft

This style can clarify product behavior for technical and nontechnical reviewers.


One test can have several assertions

Several assertions are appropriate when they describe one outcome:

error is visible
field is invalid
submit is blocked

Split tests when assertions represent independent behaviors with different setup or failure meaning.


Avoid mega tests

A mega test that covers an entire application can be:

  • slow;
  • hard to diagnose;
  • stateful;
  • impossible to run in parallel;
  • fragile under small changes.

Keep critical journeys focused and compose confidence across layers.


Test independence

Each test should control or reset:

  • data;
  • clock;
  • network;
  • storage;
  • authentication;
  • feature flags;
  • browser state.

Order should not decide whether a test passes.


Parallel execution

Parallel tests need:

  • isolated data;
  • independent ports or contexts;
  • deterministic seeds;
  • no shared mutable files;
  • bounded server resources.

Parallelism improves speed only when the environment preserves independence.


Determinism

Control or model:

  • time;
  • randomness;
  • IDs;
  • network order;
  • locale;
  • timezone;
  • animation;
  • async scheduling.

Do not remove realistic variability from the product merely to make tests easy.


Locale, RTL, and date testing

Test representative:

  • long translations;
  • plural forms;
  • right-to-left layout;
  • localized dates and numbers;
  • timezone boundaries;
  • daylight-saving transitions.

The default locale hides real layout and logic bugs.


Browser APIs need realistic tests

If a feature uses:

  • storage;
  • service workers;
  • workers;
  • notifications;
  • permissions;
  • clipboard;
  • media;
  • WebSockets or SSE;

use a test environment that models the relevant behavior or include a real-browser test.


Testing service workers and live connections

Test:

  • install and update;
  • cache strategy;
  • offline fallback;
  • message validation;
  • reconnect;
  • duplicate events;
  • missed-event recovery;
  • cleanup.

Do not test only the connected happy path.


Race testing

Make the race controllable at the boundary:

sequenceDiagram
    participant Test as Test Runner
    participant MSW as Mock Boundary
    participant App as Client UI
    Test->>MSW: Hold Response A
    App->>MSW: Dispatch Request A
    App->>MSW: Dispatch Request B
    Test->>MSW: Resolve Response B
    MSW-->>App: Render B
    Test->>MSW: Resolve Response A (Delayed)
    MSW-->>App: Discard or Ignore A
    Test->>App: Assert UI displays B

Assert that the current identity or version wins according to the product policy.


Error and loading boundary testing

Verify that a failure is contained at the intended boundary:

remote panel fails → panel error and retry
shell remains usable

Also test loading fallbacks for focus, layout, and accessibility.


Cache behavior testing

Test:

  • key separation;
  • fresh versus stale;
  • deduplication;
  • invalidation;
  • mutation reconciliation;
  • cache eviction;
  • user identity scope.

A cache test should prove what the user sees, not merely that a map received a value.


Optimistic update test

Cover the optimistic state lifecycle:

flowchart TD
    Init["User Submits Action"] --> Prov["Apply Provisional Value to UI"]
    Prov --> Net{"Server Response"}
    Net -->|200 Confirmed| Keep["Preserve or Reconcile"]
    Net -->|500 Rejected| Roll["Rollback to Previous State & Show Error Alert"]
    Net -->|Newer Update Exists| Guard["Avoid Overwriting Newer Truth"]

Optimism without rollback testing is only a happy-path demo.


Testing forms

Cover:

  • labels and names;
  • touched and dirty behavior;
  • field and cross-field validation;
  • disabled and pending states;
  • server validation mapping;
  • draft preservation on failure;
  • successful reset or navigation.

Forms are workflows, not just input snapshots.


Testing routing

Verify:

  • route parsing and invalid parameters;
  • redirects;
  • URL state serialization;
  • back and forward behavior;
  • scroll and focus;
  • route-level loading and errors;
  • code-split failure and recovery.

The browser history is part of the product contract.


Permission and security testing boundaries

Test:

  • unauthorized UI behavior;
  • server rejection despite hidden controls;
  • CSRF decisions;
  • CORS integration;
  • token and session expiration;
  • safe rendering;
  • open redirect prevention.

Do not treat a client-side permission branch as security enforcement.


Coverage is a map, not a grade

Coverage can show:

  • unexecuted branches;
  • missing error paths;
  • dead code candidates;
  • areas needing review.

High coverage can still miss wrong assertions, incorrect fixtures, inaccessible behavior, and integration failures.


Mutation testing

Mutation testing changes code deliberately to see whether tests fail.

It can reveal weak assertions and untested logic.

Use it selectively on high-risk pure logic; applying it everywhere may cost more than the confidence gained.


Property-based testing

Property-based tests generate many inputs and verify invariants:

serialize(parse(url)) remains valid
price never formats as a negative value
parser rejects malformed records

They are useful for parsers, state transitions, URLs, and domain rules with broad input space.


Themes and responsive visual testing

Visual coverage should include:

  • light and dark themes;
  • narrow and wide viewports;
  • high text scale where relevant;
  • RTL;
  • long content;
  • loading and error states.

A component that looks correct in one theme and viewport is not fully verified.


Test environment strategy

Define which environment supports each layer:

flowchart LR
    U["Unit<br/>Fast Node / Bun"] --> C["Component<br/>jsdom / Browser Runner"]
    C --> I["Integration<br/>Controlled Mock Services"]
    I --> E["E2E<br/>Real Headless Browsers"]
    E --> P["Production<br/>Canary / RUM Telemetry"]

Keep environment differences documented and intentional.


Smoke tests

Smoke tests provide a fast release signal:

  • app loads;
  • critical route works;
  • one key action succeeds;
  • health and assets are available.

They should be small, reliable, and safe to run against production-like systems.


Production tests must be safe

Use:

  • read-only checks;
  • isolated test accounts;
  • non-destructive identifiers;
  • explicit cleanup;
  • rate limits;
  • privacy-aware logging.

Do not test a production payment or deletion path by accident.


Feature flags and testing

Flags multiply states:

old behavior × new behavior × user role × route × device

Define flag ownership, expiry, test coverage, rollout, and removal.

An old flag is permanent untested architecture.


Test matrix explosion

You cannot test every combination of:

  • browsers;
  • locales;
  • roles;
  • feature flags;
  • network conditions;
  • route states;
  • data sizes.

Prioritize combinations by risk and use representative sampling plus targeted coverage.


CI test parallelism

Parallel CI can reduce feedback time when:

  • tests are independent;
  • data is isolated;
  • workers have resources;
  • artifacts remain attributable;
  • test selection is reliable.

Fast, nondeterministic CI is not a quality improvement.


Test selection

Use changed files, dependency graphs, and risk labels to select fast feedback.

Still run a broader suite at release boundaries or on a reliable schedule.

Selection is safe only when the graph and ownership assumptions are accurate.


Quarantine flaky tests carefully

A quarantine policy needs:

  • owner;
  • issue;
  • reason;
  • expiry or review date;
  • alternative coverage;
  • removal plan.

Quarantine is a temporary containment measure, not a second home for broken tests.


Delete low-value tests

Remove tests that:

  • assert implementation with no product risk;
  • duplicate stronger coverage;
  • fail noisily without useful evidence;
  • protect behavior intentionally being removed;
  • cost more to maintain than the confidence they provide.

Test suites need maintenance like production code.


Testing refactors

Good tests allow internal refactoring while preserving contracts.

If a visual extraction breaks dozens of tests, inspect whether those tests depend on private structure rather than user behavior.

Testing should support architecture evolution, not freeze every implementation detail.


Test architecture smells

Watch for:

  • every test queries a test ID;
  • every internal function is mocked;
  • refactoring markup breaks hundreds of tests;
  • E2E suite takes hours;
  • failures disappear on retry;
  • 100% coverage but critical bugs escape;
  • snapshots are approved without review;
  • production bugs cannot be reproduced.

These are signals to redesign the evidence strategy.


A balanced strategy by layer

flowchart TD
    S["Static Checks: Contracts & unsafe patterns"]
    U["Unit Tests: Pure rules & transformations"]
    C["Component Tests: Visible interaction & semantics"]
    I["Integration Tests: Boundaries & failure recovery"]
    E["E2E Tests: Critical user revenue journeys"]
    P["Production: Safe smoke, RUM & telemetry"]

    S --> U --> C --> I --> E --> P

The layers should reinforce rather than duplicate one another.


Testing a dialog

Verify:

  • accessible name;
  • initial focus;
  • Escape and close button;
  • focus restoration;
  • background interaction policy;
  • error and loading states;
  • keyboard navigation.

One user-facing behavior can need several complementary tests.


Testing a pricing function

Unit-test:

  • currency and rounding;
  • zero and negative handling;
  • locale formatting;
  • discounts and boundaries;
  • invalid input policy.

Then integration-test where the formatted value appears in a real product journey.


Cover the complete search user journey:

flowchart LR
    Type["User Types Query"] --> Debounce["Debounce Timer"]
    Debounce --> Req["Network Request"]
    Req --> Load["Loading Skeleton"]
    Load --> Res["Results Display"]
    Res --> URL["Sync URL Search Params"]

Search is a small feature with many asynchronous contracts.


Testing offline drafts

Verify:

  • draft persistence;
  • reload recovery;
  • outbox entry;
  • sync status;
  • retry limit;
  • duplicate prevention;
  • conflict behavior.

Do not test only the initial save click.


Testing micro-frontends

Use:

  • remote unit and component tests;
  • contract tests for mount and events;
  • host integration tests;
  • remote failure and fallback tests;
  • end-to-end cross-route journeys;
  • version compatibility checks.

The runtime boundary needs evidence at both sides.


Testing rendering topologies

Verify:

  • server and client output agreement;
  • hydration behavior;
  • static freshness;
  • stream loading and failure;
  • client-only boundaries;
  • route navigation and handoff;
  • secrets staying server-side.

Rendering architecture changes what the test environment must observe.


Testing performance contracts

Test or monitor:

  • route budgets;
  • asset and chunk limits;
  • key interaction timing;
  • list behavior at representative sizes;
  • memory cleanup;
  • loading and transition markers.

Use performance tests as signals with context, not brittle universal promises.


Testing security contracts

Verify:

  • safe text rendering;
  • sanitized rich content;
  • CORS and credential behavior;
  • CSRF protection;
  • server authorization;
  • redirect validation;
  • security headers;
  • no secrets in browser artifacts.

Security tests should target trust boundaries explicitly.


The test strategy document

Document:

  • risks;
  • test layers;
  • supported browsers;
  • data and environment strategy;
  • ownership;
  • CI selection;
  • flaky-test policy;
  • production verification;
  • review and deletion rules.

The document keeps the suite intentional as the system evolves.


Testing philosophy

flowchart TD
    C["Test the Contract, Not Implementation"]
    B["Choose the Nearest Useful Boundary"]
    A["Make Every Failure Actionable"]
    J["Protect Critical Revenue Journeys"]
    T["Keep the Suite Fast & Trustworthy"]

    C --- B --- A --- J --- T

More tests do not automatically mean higher quality.

Better evidence does.


Practical lab: Resilient UI Integration Suite

Test a catalogue workflow through user-visible behavior, accessibility semantics, network boundaries, and recovery paths.

The practical builds a layered strategy rather than one giant end-to-end test.


Practical stages 1–2: pure logic & accessible semantics

  1. Stage 1 (Pure Logic & Parser Unit Testing):
    • Test price formatting, pagination math, and schema parsers against valid, boundary, and corrupt input.
    • Pure fast Node runtime without DOM overhead.
  2. Stage 2 (Component Semantics & Accessible Names):
    • Query by getByRole and getByLabelText.
    • Test keyboard navigation (Tab, Escape) and focus retention.

Verification: Tests depend on platform accessibility contracts, not CSS classes or private component state.


Practical stages 3–4: boundary mocking & fault injection

  1. Stage 3 (Network Interception with MSW):
    • Intercept requests at the HTTP transport boundary.
    • Simulate network delays, 500 server crashes, and offline states.
  2. Stage 4 (Asynchronous Resilience & Fault Injection):
    • Inject deliberate faults (dropped AbortController, broken optimistic rollback, missing aria-invalid).
    • Confirm tests fail immediately and pinpoint the exact failure mechanism.

Verification: Tests prove loading, empty, error, retry, and cancellation behaviors without arbitrary sleep() delays.


Practical stage 5: Playwright critical-path journey

  1. Stage 5 (Playwright Critical-Path Browser Journey):
    • Run a real headless Chromium/Firefox/WebKit test.
    • Exercise the full user journey: search, filter, optimistic order drafting, and error recovery.
    • Capture trace artifacts, network waterfalls, and screenshots on failure.

Verification: Fast feedback in CI with zero flakiness; high confidence across real browser layout engines.


Practical extension: contract and mutation tests

Add:

  • a contract test for the runtime-validated API response;
  • a deliberate mutation test that changes an important pricing or validation rule.

Confirm that the suite fails for the right reason and reports an actionable difference.


Try this yourself

Choose one critical user journey and map:

flowchart LR
    R["Risk Identified"] --> B["Test Boundary"]
    B --> S["Setup / Seed"]
    S --> A["User Action"]
    A --> AS["Semantic Assertion"]
    AS --> AR["Diagnosable Artifact"]

Then remove one test that duplicates stronger evidence and explain why confidence remains adequate.


Troubleshooting guide (Part 1)

SymptomLikely cause
Tests break after harmless markup refactorAssertions depend on private structure
Everything uses test IDsUser-facing semantics are missing or ignored
Async tests need sleepsTests wait for time instead of conditions
Mocks hide integration failuresMock boundary is too deep
E2E suite is slow and flakyToo much setup and shared mutable data

Troubleshooting guide (Part 2)

SymptomLikely cause
Retry makes CI greenFlakiness is being normalized
Coverage is high but bugs escapeAssertions and risk mapping are weak
Visual diffs are always approvedReview policy and baseline ownership are weak
Offline test passes only onceStorage and cleanup are not isolated

Completion checklist

  • risks determine the test layers;
  • static, unit, component, integration, and E2E roles are clear;
  • tests use semantic user-facing queries where appropriate;
  • keyboard and focus behavior are covered;
  • async states and races are deterministic;
  • network mocks sit at stable boundaries;
  • critical browser journeys have failure artifacts;
  • accessibility and visual checks supplement behavior tests;
  • flakiness has ownership and a removal policy;
  • coverage and CI selection support, rather than replace, judgment.

Misconceptions to leave behind (Part 1)

MisconceptionBetter mental model
100% coverage means well testedCoverage is a map, not a confidence grade
Unit tests are always betterUse the nearest boundary that proves the risk
Everything should be E2EBroad tests are costly and should protect journeys
Component tests should inspect stateTest visible behavior and semantics
CSS selectors are forbiddenUse the selector that matches the contract
Test IDs are badThey are a legitimate fallback when semantics are unsuitable

Misconceptions to leave behind (Part 2)

MisconceptionBetter mental model
getByRole certifies accessibilityQueries are one quality signal, not an audit
Simulated DOM equals a browserReal browser behavior needs targeted tests
Mocks make tests reliableBoundary choice and realistic failures matter
Sleeps fix async testsWait for meaningful conditions
Retries solve flaky testsThey can hide nondeterminism
More tests always mean higher qualityBetter evidence and maintainability matter

The chapter in one sentence

Build a layered, user-centered test strategy that protects real risks, observes meaningful boundaries, and remains trustworthy under failure and change.


Next: Chapter 17

The next chapter will build on resilient testing with:

  • maintainability and refactoring;
  • technical debt and architectural evolution;
  • documenting decisions;
  • sustainable quality over a system’s lifetime.

Questions

Which important user journey currently has the most confidence from implementation details - and the least evidence from the behavior users actually experience?

17 - Continuous Delivery, Observability & Maintenance

Chapter 17: build a safe frontend delivery loop with reproducible artifacts, progressive releases, production signals, rollback, and maintenance.

Continuous Delivery, Observability & Maintenance

Close the loop from commit to production learning

Chapter 17

Polla Fattah


Today’s goal

Make delivery and production feedback part of front-end architecture.

We will connect:

  • CI, reproducibility, artifacts, and environments;
  • preview, staging, production, and deployment protection;
  • release strategies, feature flags, canaries, and rollback;
  • logs, metrics, traces, RUM, errors, and session context;
  • dashboards, alerts, SLOs, error budgets, and synthetic monitoring;
  • dependency, browser, API, service-worker, and technical-debt maintenance;
  • ownership, runbooks, incident response, privacy, and production readiness.

By the end of today you can

  • design a CI pipeline with fast and broad verification stages;
  • produce and promote one versioned artifact;
  • separate preview, staging, and production responsibilities;
  • use feature flags without confusing them with authorization;
  • plan canary, rollback, and kill-switch behavior;
  • distinguish monitoring from observability;
  • define useful frontend telemetry without oversharing data;
  • create SLOs and alerts with owners and response actions;
  • maintain dependencies, browser support, APIs, storage, and service workers;
  • write a production runbook and rehearse the delivery loop.

The central principle

A production system is complete only when it can be built reproducibly, released safely, observed meaningfully, rolled back deliberately, and maintained continuously.

Delivery is not the final step after architecture.

It is where architecture meets users, operators, and reality.


The complete production loop

flowchart TD
    Commit["1. Commit & Code Review"] --> Verify["2. Parallel CI Verification Gates"]
    Verify --> Build["3. Build Immutable Hashed Artifact"]
    Build --> Preview["4. Deploy Isolated Preview URL"]
    Preview --> Canary["5. Progressive Canary Rollout (5% → 25% → 100%)"]
    Canary --> Observe["6. Continuous Observability & SLO Monitoring"]
    Observe --> Mitigate{"Health Check"}
    Mitigate -->|Degraded / Spikes| Rollback["Automated Rollback / Kill Switch"]
    Mitigate -->|Healthy| Maintain["7. Technical Debt & Dependency Maintenance"]

Every arrow needs an owner, evidence, and a recovery path.


Delivery is part of architecture

Delivery decisions determine:

  • which code reaches users;
  • how quickly a fix can ship;
  • whether a release can be identified;
  • how a failure is contained;
  • whether rollback is possible;
  • how production evidence reaches developers.

A frontend that cannot be safely changed is not operationally complete.


Continuous integration

CI validates changes in a clean, shared environment.

It can verify:

  • installation;
  • linting and formatting;
  • type checking;
  • unit and integration tests;
  • production build;
  • artifact integrity;
  • smoke behavior.

CI is a feedback system, not merely a hosted command runner.


CI is more than “run tests”

flowchart LR
    Inputs["Source + Lockfile + Node Runtime"] --> Install["Clean 'npm ci' Install"]
    Install --> Checks["Type Checks, Linters, Tests"]
    Checks --> Build["Production Bundle with Content Hashes"]
    Build --> Artifact["Immutable Artifact + Release Metadata"]

The environment and produced artifact are part of what CI verifies.


CI events

Useful triggers include:

  • pull request;
  • push to a protected branch;
  • release tag;
  • scheduled dependency check;
  • manual promotion;
  • rollback or emergency fix.

Choose the checks and permissions appropriate to each event.


Fast feedback still matters

flowchart LR
    Ed["Editor / IDE"] --> Loc["Pre-commit Hooks"]
    Loc --> PR["PR Verification CI"]
    PR --> Merge["Merge to Main"]
    Merge --> Rel["Canary Release"]

Fast checks should catch common mistakes near the change.

Broad checks should protect integration and delivery confidence.

Do not make every local edit wait for the entire release pipeline.


Parallel CI

Independent jobs can run concurrently:

flowchart LR
    subgraph ParallelGates["Parallel CI Gates"]
        L["Lint & Formatting"]
        T["Type Checking (tsc)"]
        U["Unit & Component Tests"]
        S["Security Secrets Scan"]
    end
    L --> Decision["Build & Deploy Decision"]
    T --> Decision
    U --> Decision
    S --> Decision

Parallelism requires isolated environments, clear dependencies, and attributable artifacts.


CI caching

Cache:

  • package downloads;
  • transformed dependencies;
  • test artifacts;
  • build tasks;
  • browser binaries when safe.

Cache keys must include every input that affects correctness.

An incorrect cache can create false green builds.


Reproducibility

A reproducible build depends on:

  • locked dependencies;
  • known runtime version;
  • explicit environment inputs;
  • stable build configuration;
  • deterministic generation;
  • documented external services.

If two clean builds produce different artifacts, investigate why before relying on promotion.


Continuous integration versus delivery versus deployment

flowchart LR
    CI["Continuous Integration<br/>Automated verification of every push"] --> CD["Continuous Delivery<br/>Always maintains a deployable release candidate"]
    CD --> CDeploy["Continuous Deployment<br/>Automatic rollout to production upon passing gates"]

An organization can practice delivery without automatically deploying every commit.

Automation level should match risk and release confidence.


Build artifact

An artifact is the output intended for deployment:

  • static assets;
  • HTML;
  • server bundle;
  • manifests;
  • source-map references;
  • release metadata.

It should be identifiable and traceable back to source and dependencies.


Immutable artifacts

build once → store artifact → promote same artifact

Rebuilding separately for staging and production can produce different output.

Promotion should change environment configuration, not application bytes, whenever possible.


Static frontend artifact

For a static application, verify:

  • asset fingerprints;
  • manifest paths;
  • base URL;
  • cache headers;
  • source-map policy;
  • service-worker version;
  • public environment values.

The static artifact is still an operational release unit.


Server-rendered artifact

A server-rendered release may include:

  • server code;
  • client chunks;
  • route configuration;
  • environment references;
  • source maps;
  • migration compatibility.

Test the server and browser parts as one release contract.


Environments

Different environments have distinct operational responsibilities:

flowchart LR
    Dev["Development<br/>Local / Fast HMR"] --> Prev["Preview<br/>PR-Isolated Ephemeral URL"]
    Prev --> Stg["Staging<br/>Shared Pre-Production Environment"]
    Stg --> Prod["Production<br/>Global Edge CDN & Telemetry"]

Each environment should have a purpose rather than existing only because a template created it.

Document data, secrets, access, observability, and deployment differences.


Development environment

Optimize for:

  • fast feedback;
  • local debugging;
  • safe test data;
  • representative boundaries;
  • developer control.

Development convenience must not hide production-critical behavior.


Preview environments

A preview environment connects a change to a deployed, reviewable experience.

It can reveal:

  • asset path issues;
  • routing failures;
  • environment assumptions;
  • visual changes;
  • integration behavior.

Preview environments improve communication

They let reviewers discuss:

  • the actual UI;
  • loading and error states;
  • responsive behavior;
  • accessibility;
  • product intent.

Preview is evidence that complements code review; it does not replace it.


Preview environment data

Use data that is:

  • safe to expose to reviewers;
  • representative enough to reveal behavior;
  • isolated from production;
  • resettable or expiring;
  • privacy-aware.

Never copy sensitive production data into preview casually.


Staging

Staging can validate:

  • production-like infrastructure;
  • integrations;
  • deployment scripts;
  • migration compatibility;
  • smoke and end-to-end behavior.

It is not automatically identical to production and should not create false certainty.


Production

Production needs:

  • controlled access;
  • secrets management;
  • monitoring;
  • rollback;
  • incident response;
  • privacy controls;
  • support ownership;
  • documented change history.

The production environment is part of the product’s operating system.


Deployment protection

Protect production with:

  • required reviews;
  • environment approvals;
  • scoped credentials;
  • branch or tag rules;
  • concurrency control;
  • audit history;
  • automated verification.

Protection should reduce dangerous mistakes without making recovery impossible.


Environment secrets

Secrets should be injected through the deployment environment, not committed or embedded in public frontend assets.

Separate:

public configuration → browser may receive
private secret        → trusted server or CI boundary only

Rotate and audit access.


Principle of least privilege

Give each job and environment only the access it needs.

Examples:

  • preview cannot delete production data;
  • build cannot deploy unrelated services;
  • frontend runtime cannot read deployment credentials;
  • telemetry cannot expose raw user content.

Least privilege limits blast radius.


Deployment concurrency

Decide what happens when two deployments target one environment:

  • queue;
  • cancel older pending deploy;
  • serialize production;
  • allow parallel previews;
  • block promotion until verification.

Ambiguous concurrency creates releases that are difficult to identify and roll back.


Deployment history

Record:

  • release version;
  • commit;
  • artifact digest;
  • environment;
  • time;
  • actor or automation;
  • feature flags;
  • result and rollback.

History turns “something changed” into an actionable investigation.


Release version

Expose a release identity safely:

UI build version
API version
deployment timestamp
commit or artifact identifier

It helps correlate errors, performance, support reports, and rollbacks.


Release correlation

Telemetry should connect:

user session → route → release → request → error / performance event

Do not collect identifying data unnecessarily. Use bounded, privacy-aware correlation IDs.


Deployment strategies

Common strategies include:

  • rolling;
  • blue-green;
  • canary;
  • feature-flagged release;
  • full replacement with fast rollback.

Select based on compatibility, traffic, data migration, and operational capability.


Rolling deployment

Replace instances or assets gradually.

During the transition, old and new versions may coexist.

Ensure:

  • API compatibility;
  • shared asset availability;
  • session behavior;
  • migration safety;
  • observability by version.

Backward compatibility during deployment

For a period, support:

old client ↔ current server
new client ↔ current server

Long-lived browser tabs, cached assets, service workers, and delayed requests make compatibility a real requirement.


Blue-green deployment

flowchart LR
    Router["Global Edge Router / CDN"]
    subgraph Blue["Blue Environment (Active v1.4)"]
        B_App["Production App Servers"]
    end
    subgraph Green["Green Environment (Staged v1.5)"]
        G_App["New Release Ready"]
    end
    Router -->|100% Active Traffic| Blue
    Router -. Instant Cutover .-> Green

It can enable fast switching and rollback.

Costs include duplicate capacity, data compatibility, and confidence that the inactive environment is genuinely ready.


Blue-green trade-offs

Consider:

  • database and API state;
  • cache warmth;
  • background jobs;
  • live connections;
  • asset references;
  • traffic switching and rollback timing.

Switching traffic does not undo side effects already performed by the new version.


Canary release

Expose a release to a small cohort and compare:

  • errors;
  • performance;
  • task completion;
  • support signals;
  • conversion or domain outcomes.

Canaries require meaningful cohort identity and enough traffic to observe signal.


Architectural timeline of a production release

sequenceDiagram
    autonumber
    actor Dev as Engineer
    participant CI as CI Pipeline
    participant Reg as Immutable Registry
    participant Edge as Edge CDN Router
    participant Telemetry as RUM / Telemetry

    Dev->>CI: Push Git commit (Tag v2.4.0)
    CI->>CI: Run lint, types, unit, E2E gates
    CI->>Reg: Publish immutable assets (hash: a9f3c1)
    CI->>Edge: Deploy v2.4.0 to 10% Canary cohort
    Edge-->>Telemetry: Stream real user Core Web Vitals & errors
    Note over Telemetry: 15-Minute Observation Window
    Telemetry-->>Edge: SLO Healthy (error rate < 0.05%)
    Edge->>Edge: Promote v2.4.0 to 100% traffic

Concrete failure & recovery scenario: Safari crash

sequenceDiagram
    autonumber
    participant Users as Safari 16 Users (10% Canary)
    participant Edge as Edge CDN Router
    participant RUM as Error Telemetry
    participant OnCall as On-Call Engineer

    Edge->>Users: Serve v2.4.0 bundle
    Note over Users: Syntax Error: Unsupported Regex Lookbehind
    Users-->>RUM: Spikes unhandled exceptions (5.8% error rate!)
    RUM->>OnCall: PagerDuty Alert: Critical SLO breach on Safari
    OnCall->>Edge: Trigger Instant Rollback to v2.3.0
    Edge->>Users: Serve previous known-good bundle v2.3.0
    Note over Users: Errors cease immediately (Recovery Time: 90s)
    OnCall->>OnCall: Triage syntax target in tsconfig for v2.4.1

Release strategy versus feature strategy

release strategy → how code is delivered
feature strategy → who can see or use behavior

A release can ship code behind a flag.

A canary can expose a release to a cohort.

Do not treat the mechanisms as interchangeable.


Feature flags are operational controls

Flags can support:

  • gradual rollout;
  • experiments;
  • kill switches;
  • migration paths;
  • permission-aware presentation.

Each flag needs an owner, purpose, lifecycle, and removal plan.


Feature flags are not just booleans

A flag may depend on:

  • user or account;
  • percentage cohort;
  • environment;
  • route;
  • experiment assignment;
  • release version;
  • server-side policy.

Use a clear evaluation model and keep the result observable.


Vendor-neutral flag APIs

if (flags.isEnabled("newCatalogue")) {
  renderNewCatalogue();
}

Keep product code dependent on a small capability interface rather than one vendor’s SDK throughout the application.


Flag evaluation location

Flags may evaluate:

server → secure and consistent, request-time
client → flexible, but value may be visible and stale
edge   → close to user, operationally constrained

Choose based on security, consistency, latency, and rollout needs.


Feature flags are not authorization

{flags.isEnabled("delete-flow") && <DeleteButton />}

This controls exposure or experience.

The server must still authorize the delete operation.


Flag categories

release flag     → safely expose unfinished code
operational flag → disable or tune a capability
experiment flag  → assign variants and measure
permission flag  → reflect access, never enforce alone

Different categories need different owners, lifetimes, and audit rules.


Flag lifecycle

create → rollout → measure → complete or remove

If a flag remains after its decision is complete, it becomes permanent branching complexity.


Flag debt

Old flags create:

  • dead paths;
  • extra tests;
  • confusing support behavior;
  • inconsistent analytics;
  • migration risk.

Track flag removal like technical debt with an owner and due condition.


Kill switches

A kill switch disables a harmful or expensive capability quickly.

It should have:

  • safe fallback;
  • owner;
  • access control;
  • audit trail;
  • test coverage;
  • clear recovery process.

A kill switch is useful only if operators can trust it under pressure.


Rollback

Rollback restores a previous known release or behavior.

It requires:

  • an identifiable artifact;
  • compatible data and APIs;
  • deployment access;
  • tested procedure;
  • observability to confirm recovery.

Rollback is a normal safety mechanism, not an admission of failure.


Rollback must be practiced

A document is not evidence that rollback works.

Rehearse:

  • who decides;
  • what command or action occurs;
  • which version returns;
  • how flags change;
  • what users experience;
  • how verification confirms recovery.

Practice reveals missing access, compatibility, and monitoring assumptions.


Database and API compatibility

A frontend rollback may meet:

  • a changed API;
  • a migrated schema;
  • a new required field;
  • an invalidated cache;
  • an incompatible service worker.

Use expand-and-contract changes and backward-compatible periods where rollback matters.


Rollback versus forward fix

flowchart TD
    Incident["Production Incident Detected"] --> Q1{"Is root cause obvious<br/>and fix trivial (<5 min)?"}
    Q1 -->|Yes, zero risk| Forward["Forward Hotfix via CI Pipeline"]
    Q1 -->|No / High Risk| Q2{"Was client storage or<br/>DB schema mutated?"}
    Q2 -->|Backward Compatible| Rollback["Instant Artifact / CDN Rollback"]
    Q2 -->|Schema Broken| Kill["Activate Feature Kill Switch & Triage"]

Choose based on blast radius, data effects, detection certainty, and fix confidence.


Observability

Observability asks whether internal system state can be understood from emitted evidence.

For frontend systems, evidence can include:

  • errors;
  • performance events;
  • network failures;
  • route transitions;
  • release identity;
  • user journey markers;
  • feature-flag context.

Monitoring versus observability

monitoring   → known signals and expected thresholds
observability → investigate unexpected states from evidence

You need both:

  • dashboards and alerts for known failure;
  • logs, traces, context, and exploration for unfamiliar failure.

Three traditional telemetry signals

flowchart TD
    Logs["Logs<br/>Structured contextual events & breadcrumbs"]
    Metrics["Metrics<br/>Aggregated counters, rates, and Web Vitals percentiles"]
    Traces["Traces<br/>End-to-end distributed execution paths across client and API"]

    Logs --- Metrics --- Traces

Frontend observability adapts these signals to browser privacy, lifecycle, and network constraints.


Logs

Useful frontend logs are:

  • structured;
  • bounded;
  • correlated to release and journey;
  • privacy-aware;
  • actionable.

Do not turn the browser console into a production data lake.


Browser console is not production observability

Console output can disappear because:

  • the user closes the page;
  • the browser filters it;
  • no one can access the user’s console;
  • context is incomplete;
  • sensitive data was logged.

Send carefully designed events to a controlled telemetry system when investigation requires it.


Metrics

Frontend metrics can include:

  • error rate;
  • route load time;
  • Web Vitals;
  • interaction timing;
  • request failure rate;
  • feature adoption;
  • queue depth or sync status.

Define units, aggregation, segment, and action before collecting a metric.


High cardinality

Values such as raw URLs, user IDs, or arbitrary error messages can create too many metric dimensions.

Prefer bounded labels:

route_template=/products/:id
error_kind=validation
release=2026.09.24

Keep detailed context in traces or sampled events with privacy controls.


Traces and spans

flowchart TD
    Root["User Journey: Checkout Submission"]
    Root --> S1["Span 1: Form Validation & Client State Update"]
    Root --> S2["Span 2: HTTP POST /api/v1/checkout (traceparent)"]
    S2 --> S3["Span 3: Backend Gateway & Payment Provider"]
    Root --> S4["Span 4: DOM Paint & Confirmation View Render"]

Traces connect frontend work with backend and network events.

They are especially useful for slow or distributed journeys.


Frontend tracing has special challenges

The browser has:

  • intermittent sessions;
  • sampling constraints;
  • privacy boundaries;
  • offline periods;
  • multiple tabs;
  • limited background time;
  • user-controlled execution.

Design telemetry that remains useful without assuming server-like process lifetime.


OpenTelemetry

OpenTelemetry provides a vocabulary and instrumentation approach for traces, metrics, and logs.

Browser support and semantic coverage continue to evolve.

Use it as an interoperability tool, not a promise that every frontend detail is automatically standardized.


Real User Monitoring revisited

RUM should connect:

release + route + segment + journey + metric

Collect enough context to act while avoiding unnecessary personal or form data.


Error monitoring

Error monitoring should capture:

  • error type and message;
  • stack with private source maps;
  • release identity;
  • route and operation;
  • browser context;
  • breadcrumbs;
  • safe user and feature context.

Group similar failures so teams can prioritize rather than count noise.


Global error handlers

Global handlers can catch unexpected failures and report them.

They should not:

  • hide the failure without a recovery UI;
  • send secrets;
  • duplicate every error repeatedly;
  • replace local error boundaries;
  • create a false sense that all errors are observable.

Source maps in production monitoring

Source maps make minified errors actionable.

Prefer private upload to the error-monitoring system rather than public deployment when exposure is not needed.

Tie maps to exact release artifacts.


Error grouping

Group by stable cause signals:

  • error type;
  • normalized message;
  • stack location;
  • operation;
  • release.

Avoid grouping all failures into one generic event or splitting one bug into thousands of dynamic messages.


Error rate needs context

Define:

errors / meaningful sessions
errors / requests
errors / route visits

An error count alone does not tell whether users are affected more or fewer.


Network error monitoring

Capture categories such as:

  • offline;
  • timeout;
  • DNS or connection;
  • CORS;
  • 401/403;
  • 404;
  • 429;
  • 5xx;
  • schema failure;
  • cancellation.

Different categories need different owners and actions.


Client versus server error

client error → invalid state, rendering, browser boundary
server error → API, infrastructure, authorization, domain failure

The same user-visible failure may require evidence from both sides.

Correlate requests and releases safely.


User context without oversharing

Useful context might be:

  • anonymous session ID;
  • account tier or role category;
  • route template;
  • release;
  • feature flag assignment.

Avoid raw form values, tokens, passwords, private content, and unnecessary identity data.


Breadcrumbs

Breadcrumbs can record a bounded sequence such as:

route opened → filter changed → request failed → retry clicked

They should explain the journey without recording sensitive payloads or growing without limit.


Session replay

Session replay has strong privacy and compliance implications.

Define:

  • consent;
  • masking;
  • field exclusion;
  • retention;
  • access;
  • sensitive-route policy;
  • sampling.

Debugging value does not override user privacy.


Observability sampling

Sampling controls:

  • cost;
  • storage;
  • user impact;
  • signal volume.

Sample ordinary success heavily, but retain enough rare failures and high-severity journeys to investigate them.


Head-based versus tail-based sampling

head-based → decide at trace start
tail-based → decide after seeing outcome and duration

Tail sampling can retain slow or failed traces more effectively but requires infrastructure and buffering.


Telemetry has performance cost

Telemetry can add:

  • network requests;
  • serialization;
  • CPU;
  • memory;
  • storage;
  • privacy review.

Instrument critical signals and batch or sample responsibly.


Telemetry buffering and sendBeacon

When a page is closing, a beacon can send small analytics or diagnostic payloads without blocking navigation.

It is not a guarantee of delivery and should not carry sensitive or large data.


Observability naming

Use stable names:

catalogue.search.submit
route.catalogue.ready
api.products.failure
feature.new-catalogue.exposed

Consistent naming makes dashboards and searches usable across releases.


Operational dashboards

A dashboard should show:

  • current release;
  • errors and affected sessions;
  • performance signals;
  • traffic and sample size;
  • top routes and operations;
  • recent deployments;
  • feature flags and cohorts;
  • known incidents.

Dashboards should support a decision, not display every metric.


Alerting

An alert should specify:

  • condition;
  • severity;
  • owner;
  • response time;
  • runbook;
  • suppression or grouping;
  • recovery signal.

If nobody knows what to do after an alert, it is likely not ready to page.


Alert fatigue

Too many low-value alerts cause operators to ignore high-value ones.

Reduce fatigue with:

  • actionable thresholds;
  • grouping;
  • maintenance windows;
  • ownership;
  • severity tiers;
  • periodic review.

Static thresholds versus anomaly detection

Static thresholds are understandable for known limits.

Anomaly detection can reveal unusual changes relative to baseline.

Both need context, sample size, and a response plan. Sophisticated detection without action is noise.


Service-level indicators

An SLI is a measured aspect of service behavior:

successful route loads / route loads
successful saves / save attempts
acceptable interaction samples / interaction samples

Define the numerator, denominator, population, and measurement boundary.


Service-level objectives

An SLO sets a target for an SLI over a time window.

Examples:

  • 99.9% of save attempts receive a usable result;
  • 95% of catalogue routes meet a journey threshold;
  • 99% of releases pass smoke verification.

The objective should be meaningful to users and operators.


Error budgets

allowed unreliability = 1 − objective

The budget can inform release pace, reliability investment, and risk acceptance.

It should not be used to justify harming a small but important user group.


Frontend SLOs need care

Frontend metrics are affected by:

  • user devices;
  • networks;
  • browser extensions;
  • third-party scripts;
  • sampling;
  • route and feature mix.

Define scope and segments so the objective reflects what the team can influence and what users experience.


Health checks

A frontend health check may verify:

  • static asset availability;
  • route response;
  • API reachability;
  • configuration;
  • basic rendering;
  • deployment identity.

It should be safe, bounded, and not confuse “server responds” with “user journey works”.


Synthetic production monitoring

Synthetic checks can periodically perform a safe journey:

open → navigate → search → verify result → exit

They can detect outages before enough real users generate field data.


Synthetic monitoring trade-offs

Synthetic traffic can:

  • miss real device diversity;
  • create false confidence;
  • add load;
  • require test identities;
  • fail because of test data rather than product behavior.

Use it alongside RUM and clear ownership.


Deployment verification

After release, check:

  • artifact and version;
  • critical route;
  • authentication;
  • API requests;
  • assets and service worker;
  • errors and performance;
  • feature flag exposure.

Deployment is not complete until the user-facing path is verified.


Automatic rollback

Automatic rollback can be appropriate when:

  • the signal is reliable;
  • the failure threshold is clear;
  • the previous artifact is compatible;
  • rollback is safe;
  • operators are notified;
  • a forward fix path exists.

Do not automate a destructive or ambiguous recovery.


Progressive delivery

preview → internal cohort → small percentage → broad rollout

At each step, observe release health and decide whether to continue, pause, or revert.


Release cohorts

Cohorts should be:

  • stable enough to compare;
  • representative enough to reveal risk;
  • privacy-aware;
  • consistently evaluated;
  • identifiable in telemetry.

Random percentage alone is not enough if users move between cohorts unpredictably.


Percentage rollout consistency

Hash a stable identity or account key so a user does not switch variants on every request.

Handle anonymous users and identity changes deliberately.

Document how rollout interacts with caching and server/client evaluation.


Experimentation

Experiments need:

  • hypothesis;
  • assignment policy;
  • primary metric;
  • guardrail metrics;
  • sample and duration;
  • analysis plan;
  • stop condition;
  • cleanup plan.

An experiment is not a permanent feature flag.


Guardrail metrics

Alongside product outcomes, monitor:

  • errors;
  • latency;
  • accessibility failures;
  • support contacts;
  • abandonment;
  • memory or resource use;
  • security signals.

Stop an experiment quickly when it harms users even if the primary metric looks promising.


Maintenance is architecture

The system changes after launch:

  • dependencies update;
  • browsers change;
  • APIs evolve;
  • storage persists;
  • service workers remain installed;
  • teams and ownership change.

Maintenance paths must be designed, not improvised after years of drift.


Dependency maintenance

Maintain dependencies with:

  • regular review;
  • security triage;
  • compatibility testing;
  • controlled updates;
  • rollback or pinning path;
  • removal of unused packages.

Never updating is also a risk strategy, usually a poor one.


Update continuously, not once every three years

Small frequent updates reduce:

  • migration distance;
  • surprise incompatibilities;
  • debugging scope;
  • security exposure;
  • ownership uncertainty.

They still require review and release discipline.


Dependency update automation

Automation can open updates and run checks.

It should not merge every update blindly.

Use grouping, ownership, security priority, release notes, and artifact/performance review.


Grouping updates

Group compatible low-risk updates to reduce review noise.

Keep risky or architectural updates separate so failures and migration decisions remain understandable.


Security advisories need triage

A vulnerability report needs:

  • affected package and path;
  • whether the code ships or executes;
  • exposure and exploitability;
  • available fix;
  • mitigation;
  • owner and deadline.

Severity score alone does not decide urgency.


Transitive dependencies

You may not import a package directly and still ship it through a dependency chain.

Track:

  • why it exists;
  • who owns the direct dependency;
  • whether it is reachable in production;
  • how updates propagate;
  • whether it can be removed.

Removing dependencies

Removing a dependency can improve:

  • bundle size;
  • security surface;
  • build time;
  • maintenance;
  • licensing clarity.

Verify that a platform capability or small local implementation truly meets the requirement before replacing it.


Browser support maintenance

Browser policy affects:

  • transformation targets;
  • polyfills;
  • testing matrix;
  • CSS behavior;
  • support cost;
  • user access.

Do not drop support from global statistics alone if your product serves a meaningful affected population.


Deprecating browser support

Use a planned lifecycle:

measure → announce → warn / document → migrate → remove

Provide a clear reason and verify that the release does not strand critical users unexpectedly.


API maintenance

Maintain APIs with:

  • backward-compatible additions;
  • versioned breaking changes;
  • schema monitoring;
  • deprecation windows;
  • client compatibility;
  • migration communication.

Frontend and backend releases often overlap in time.


Schema monitoring

Monitor:

  • unknown fields;
  • missing required fields;
  • invalid types;
  • deprecated fields still used;
  • response version distribution.

Runtime validation turns silent drift into observable evidence.


Data migration and frontend compatibility

During migration, old clients may still send old shapes.

Use additive changes, compatibility periods, and server-side normalization where rollback and long-lived clients matter.


Local storage migration

Persisted client state needs:

  • version;
  • schema validation;
  • migration path;
  • reset behavior;
  • identity scope;
  • failure handling.

Never assume all users have the current local schema.


Service-worker maintenance

Service workers can outlive deployments.

Maintain:

  • cache versioning;
  • activation policy;
  • update notification;
  • old client compatibility;
  • cache cleanup;
  • offline data migration.

The installed worker is part of the deployed client population.


The stale-client problem

Users may keep a tab open while a new release ships.

The old client may call:

  • a changed API;
  • a removed route;
  • an incompatible event;
  • an old asset path.

Design compatibility and update behavior for long-lived sessions.


Update notification

An update message should explain:

  • what changed;
  • whether current work is safe;
  • when reload is appropriate;
  • whether the user can defer;
  • what happens to drafts and connections.

Do not reload unexpectedly while the user is editing important data.


Technical debt

Debt is the future cost of a current shortcut, including:

  • missing tests;
  • unsupported dependencies;
  • unclear ownership;
  • stale flags;
  • undocumented platform behavior;
  • brittle deployment assumptions.

Debt is not automatically bad; unmanaged debt is.


Debt register

Track:

  • debt item;
  • user or team impact;
  • trigger for action;
  • owner;
  • risk;
  • estimated effort;
  • expiry or review date.

An explicit register turns vague concern into prioritizable work.


Maintenance budget

Reserve capacity for:

  • dependency updates;
  • browser support;
  • observability changes;
  • flag cleanup;
  • API migration;
  • service-worker maintenance;
  • refactoring and test improvement.

If maintenance has no budget, it will compete with emergencies.


End-of-life policy

Define how the system retires:

  • old browser versions;
  • APIs;
  • packages;
  • feature flags;
  • service-worker caches;
  • preview environments;
  • telemetry schemas.

Retirement is part of lifecycle architecture.


Ownership and runbooks

A runbook should state:

  • what signal indicates the problem;
  • how to confirm it;
  • immediate mitigation;
  • rollback or kill switch;
  • communication path;
  • recovery verification;
  • follow-up owner.

Runbooks reduce decision load during incidents.


Incident response lifecycle

flowchart LR
    Det["1. Detect<br/>SLO Alert"] --> Tri["2. Triage<br/>Assess Blast Radius"]
    Tri --> Mit["3. Mitigate<br/>Rollback / Kill Switch"]
    Mit --> Com["4. Communicate<br/>Status Page Update"]
    Com --> Rec["5. Recover<br/>Verify Telemetry Normal"]
    Rec --> Lrn["6. Learn<br/>Blameless Post-Mortem"]

Keep the user impact and current system state visible throughout the incident.


Blameless learning

Post-incident review should ask:

  • what conditions made the failure possible?
  • which signal was missing or late?
  • which recovery step worked or failed?
  • what system change reduces recurrence?

The goal is stronger systems, not individual blame.


Post-incident actions

Actions should be:

  • specific;
  • owned;
  • prioritized;
  • measurable;
  • connected to the failure;
  • reviewed for completion.

“Be more careful” is not a system improvement.


Maintenance metrics

Useful signals include:

  • mean time to recovery;
  • deployment failure rate;
  • rollback time;
  • dependency age;
  • stale flag count;
  • unresolved vulnerability age;
  • test and build feedback time;
  • observability coverage.

Do not turn one metric into a target that damages the system.


Deployment frequency is not quality by itself

Frequent deployment can indicate healthy delivery or uncontrolled churn.

Pair it with:

  • change failure rate;
  • recovery time;
  • user impact;
  • reliability and performance;
  • maintenance health.

Mean time to recovery

MTTR asks how quickly the system returns to an acceptable state after failure.

Improve it through:

  • detection;
  • clear ownership;
  • safe rollback;
  • kill switches;
  • runbooks;
  • practiced communication.

Prevention matters, but recovery is part of reliability.


Observability completes testing

Tests cover known scenarios.

Production observability reveals:

  • unknown combinations;
  • real devices and networks;
  • integration drift;
  • deployment-specific failures;
  • user behavior outside assumptions.

Use production evidence to improve the next test strategy.


The production feedback loop

test hypothesis
  → release safely
  → observe real behavior
  → investigate differences
  → improve code, tests, or architecture

Production is not the end of engineering; it is a source of evidence.


Privacy and compliance

Telemetry and debugging must respect:

  • data minimization;
  • consent;
  • retention;
  • access control;
  • regional requirements;
  • deletion and subject rights;
  • sensitive fields.

Operational visibility cannot be purchased by recording everything.


Do not record form contents by default

Forms can contain:

  • passwords;
  • financial data;
  • health information;
  • personal identifiers;
  • private business content.

Mask or exclude inputs and inspect telemetry payloads before production use.


Session replay needs stronger governance

Replay can show real failures, but may capture private content and interactions.

Define:

  • masking defaults;
  • consent and regional policy;
  • sensitive route exclusions;
  • retention and access;
  • sampling and incident use.

Analytics versus observability

analytics    → product behavior and outcomes
observability → system health and failure investigation

They may share infrastructure, but have different purpose, privacy, retention, and ownership.


Product versus operational metrics

ProductOperational
conversionerror rate
task completionlatency
retentionavailability
feature adoptiondeployment health

Connect them carefully without assuming one causes the other.


Business-aware observability

A useful release dashboard can connect:

release → route → technical signal → user task → business effect

This helps teams prioritize an outage that affects a critical workflow over a noisy low-impact error.


Release dashboard example

Show:

  • current and previous release;
  • deployment status;
  • error and performance deltas;
  • critical journey success;
  • flag cohorts;
  • open incidents;
  • rollback readiness.

Keep the dashboard decision-oriented.


Deployment health window

After release, observe a defined window with:

  • expected traffic;
  • known baseline;
  • key route metrics;
  • error and support signals;
  • feature exposure;
  • rollback threshold.

Do not declare health from one successful smoke test alone.


Low-traffic products

Sparse traffic makes automatic detection harder.

Use:

  • longer observation windows;
  • synthetic checks;
  • targeted smoke tests;
  • manual review;
  • confidence-aware thresholds;
  • support and user reports.

Low volume does not mean low importance.


Maintenance and architecture decisions

Prefer architectures with:

  • understandable upgrade paths;
  • supported tooling;
  • observable boundaries;
  • reversible deployment;
  • clear ownership;
  • proportionate operational cost.

The most elegant architecture is not useful if nobody can maintain or recover it.


Boring technology has value

Mature, documented, well-understood tools can reduce:

  • operational surprise;
  • onboarding cost;
  • incident ambiguity;
  • upgrade risk;
  • dependence on one expert.

Novelty should solve a real constraint, not create one.


Upgrade path is a selection criterion

Before adopting a tool, ask:

  • how is it upgraded?
  • how are breaking changes announced?
  • can output be inspected?
  • is ownership clear?
  • can the system roll back?
  • what happens when the maintainer changes?

Current features are only one part of technology choice.


Avoid undocumented internal platforms

An internal platform without:

  • public contracts;
  • documentation;
  • release notes;
  • support ownership;
  • migration path;
  • observability;

becomes a hidden dependency and a bottleneck for every product team.


Operational readiness review

Review:

delivery / rollback / observability / performance
ownership / dependencies / privacy / support

The checklist should be proportional to user impact and system complexity.


Production readiness is proportional

A small static page and a financial workflow need different readiness depth.

Scale review according to:

  • user harm;
  • data sensitivity;
  • change frequency;
  • dependency count;
  • availability expectations;
  • recovery cost.

Proportionate does not mean careless.


Practical lab: Delivery, Observability, and Rollback Loop

Design and rehearse a safe frontend release from commit to deployment, observation, rollback, and cleanup.

The practical makes the delivery artifact, release identity, flags, telemetry, alerts, and recovery path explicit.


Practical stages 1–6: build and deploy safely

  1. Create the CI pipeline.
  2. Add parallel jobs.
  3. Produce a versioned artifact.
  4. Create a preview deployment.
  5. Protect production.
  6. Add deployment concurrency.

Verification: the tested artifact is the deployed artifact.


Practical stages 7–13: release control

  1. Expose the release version.
  2. Add a preview smoke test.
  3. Design a canary rollout.
  4. Add a release flag.
  5. Give the flag an owner.
  6. Create a kill switch.
  7. Remove a completed flag.

Feature flags do not act as authorization.


Practical stages 14–20: observe production

  1. Add error monitoring.
  2. Upload source maps privately.
  3. Add RUM.
  4. Add a custom journey metric.
  5. Create an operational dashboard.
  6. Define alerts.
  7. Add synthetic production monitoring.

Every alert needs an owner and an operational response.


Practical stages 1–3: reproducible artifact & preview verification

  1. Stage 1 (Immutable Reproducible Artifact):
    • Produce a production build with deterministic content hashes.
    • Generate release-manifest.json with commit SHA, timestamp, and metadata.
  2. Stage 2 (CI Verification Gates & Secrets Scanning):
    • Execute parallel linting, type checks, unit/integration suites, and bundle budgets.
    • Run automated secret scanning to prevent token leaks into client bundles.
  3. Stage 3 (Preview Environments & Release Identity):
    • Deploy isolated PR preview environments.
    • Inject window.__RELEASE_INFO__ for runtime telemetry attribution.

Practical stages 4–5: observability & rehearsed rollback

  1. Stage 4 (Front-End Observability & Breadcrumbs):
    • Implement zero-dependency client telemetry for unhandled errors and RUM metrics.
    • Capture user interaction breadcrumbs with strict PII masking.
  2. Stage 5 (Simulated Disaster Rehearsal & Safe Rollback):
    • Inject a deliberate production failure into v2.4 (breaking Safari form submissions).
    • Observe automated SLO breach and trigger instant CDN rollback to v2.3.
    • Verify client data compatibility (localStorage) and draft a blameless post-mortem.

Practical extension: frontend runbook

Write a short runbook covering:

detection → triage → mitigation → rollback → communication → follow-up

Include commands, owners, dashboards, thresholds, and the evidence that confirms recovery.


Try this yourself

For the next release, write:

artifact identity:
deployment strategy:
rollback trigger:
critical signal:
alert owner:
kill switch:
privacy constraint:
maintenance follow-up:

If any field is blank, the release loop has an unexamined assumption.


Troubleshooting guide (Part 1)

SymptomLikely cause
CI passes but deployed app failsTested artifact differs from deployed artifact
Rollback restores old code but breaks APICompatibility window was not designed
Alerts fire constantlyThresholds lack context or ownership
Errors cannot be debuggedRelease identity or private source maps missing
Feature flag remains foreverLifecycle and owner were never defined

Troubleshooting guide (Part 2)

SymptomLikely cause
Telemetry is expensive or unsafePayload, sampling, and privacy policy are weak
Users keep stale behaviorLong-lived clients and service-worker updates ignored
Dependency update is frighteningUpdates were deferred instead of maintained continuously
Incident response is slowRunbook and rollback were not practiced

Completion checklist

  • CI produces a reproducible, identifiable artifact;
  • the same artifact is promoted across environments;
  • production access and secrets use least privilege;
  • releases can be canaried, flagged, killed, or rolled back;
  • telemetry connects release, route, journey, and failure;
  • alerts have owners and runbooks;
  • SLOs and error budgets are meaningful and segmented;
  • dependencies, browsers, APIs, storage, and workers have maintenance paths;
  • telemetry respects privacy and data minimization;
  • rollback and recovery have been rehearsed.

Misconceptions to leave behind (Part 1)

MisconceptionBetter mental model
CI means one hosted test commandCI is clean verification and artifact production
Delivery means every commit deploysDelivery keeps a safe release ready
Passing tests makes deployment safeArtifact, environment, integration, and recovery also matter
Staging is production without usersEnvironments have different purposes and gaps
Preview replaces code reviewIt provides deployed evidence, not design judgment
Feature flags are authorizationThe server still enforces permissions
Flags can stay foreverFlags need ownership and removal

Misconceptions to leave behind (Part 2)

MisconceptionBetter mental model
Rollback is failureRollback is a normal safety mechanism
Observability means logsLogs, metrics, traces, errors, and context work together
More telemetry is always betterTelemetry has cost, privacy, and signal limits
Every error should page someoneAlerts require actionable owners and thresholds
Dependency updates should be automaticAutomation needs review, grouping, and risk triage
Users refresh after deploymentLong-lived clients require compatibility and update policy
Maintenance is separate from architectureLifecycle and recovery are architectural properties

The chapter in one sentence

Ship reproducible artifacts through a controlled release loop, observe real user impact, recover deliberately, and budget continuously for maintenance.


Next: Chapter 18

The next chapter will build on production architecture with:

  • final system integration;
  • architectural decision records;
  • capstone planning;
  • cross-cutting quality and delivery review;
  • a complete front-end platform blueprint.

Questions

If the next release harms users, how will you identify the exact artifact, detect the impact, reduce exposure, roll back safely, and learn what the system failed to tell you?

18 - Front-End Architecture & Technical Decision-Making

Chapter 18: make evidence-backed architecture decisions using quality attributes, constraints, boundaries, trade-offs, and reversible learning.

Front-End Architecture & Technical Decision-Making

Choose deliberately, learn continuously

Chapter 18

Polla Fattah


Today’s goal

Turn architecture from a collection of technologies into a transparent set of decisions.

We will connect:

  • requirements, quality attributes, constraints, and trade-offs;
  • cohesion, coupling, dependency direction, and blast radius;
  • complexity budgets, reversibility, locality, and explicit data flow;
  • progressive enhancement, platform-first design, and dependency evaluation;
  • fitness functions, ADRs, spikes, risk reduction, and failure modes;
  • performance, security, accessibility, testing, observability, and deployability;
  • team topology, governance, migration, and architecture outcomes.

By the end of today you can

  • define architecture as decisions made under constraints;
  • distinguish functional requirements from quality attributes;
  • compare alternatives without fake precision;
  • reduce coupling and blast radius through boundaries;
  • prefer reversible decisions when information is weak;
  • use ADRs to record decisions and revisit triggers;
  • turn important architecture rules into fitness functions;
  • choose a proportionate architecture for different scenarios;
  • design migration instead of assuming rewrites are cleaner;
  • defend the smallest coherent solution with evidence.

The central principle

Architecture is a set of explicit, contextual decisions about trade-offs, boundaries, and future change - not a technology stack or a prediction of everything that might happen.

Good architecture makes important change cheaper, safer, and easier to reason about.


The decision progression

flowchart TD
    Prob["1. Define Problem & Context"] --> Qual["2. Identify Quality Attributes"]
    Qual --> Const["3. Record Constraints"]
    Const --> Alt["4. Generate Alternatives"]
    Alt --> Spike["5. Reduce Unknowns via Spikes"]
    Spike --> ADR["6. Decide & Document in ADR"]
    ADR --> Guard["7. Enforce via CI Guardrails"]
    Guard --> Rev["8. Observe & Revisit with Evidence"]

    Rev -. Feeds into new decisions .-> Prob

Architecture remains a loop rather than a one-time ceremony.


Architecture is a set of decisions

Examples:

where state lives
where rendering happens
how APIs are shaped
what is shared
how failures are contained
how releases are delivered

The framework is one input to those decisions, not the decision itself.


Not everything can be optimized at once

Trade-offs may exist between:

  • autonomy and consistency;
  • speed and flexibility;
  • freshness and cacheability;
  • simplicity and isolation;
  • performance and feature richness;
  • delivery independence and runtime coupling.

Architecture makes the trade-off visible so the team can choose intentionally.


Maximum team autonomy has costs

Autonomy can increase:

  • release speed;
  • local decision quality;
  • ownership;
  • experimentation.

It can also increase:

  • duplication;
  • inconsistent behavior;
  • platform cost;
  • integration burden;
  • fragmented user experience.

The goal is useful autonomy within coherent contracts.


Quality attributes

Quality attributes describe how the system behaves:

performance / reliability / security / accessibility
maintainability / deployability / observability / scalability

They are architecture inputs because they shape boundaries and decisions.


Functional requirements versus quality attributes

functional: users can submit an inspection
quality: submission remains safe under network failure

Functional requirements describe capability.

Quality attributes describe the conditions under which that capability remains useful.


Not every quality attribute has equal priority

A government service may prioritize accessibility and reliability.

An internal analytics tool may prioritize iteration speed and data accuracy.

A marketing page may prioritize cacheability and loading performance.

Priorities should be explicit rather than assumed.


Make quality attributes measurable

fast        → p75 catalogue LCP target
reliable    → successful save ratio
accessible  → keyboard journey and audit criteria
deployable  → rollback within a defined window

An attribute becomes architectural guidance when the team can observe it.


Constraints

Constraints can come from:

  • budget;
  • deadline;
  • existing APIs;
  • browser support;
  • team skills;
  • regulation;
  • deployment platform;
  • organization;
  • data sensitivity.

Constraints are not annoyances to ignore; they define the decision space.


Requirements, constraints, and decisions

flowchart TD
    subgraph DecisionForces["The Forces Shaping Architecture"]
        Req["Functional Requirements<br/>(What the system must achieve)"]
        Qual["Quality Attributes<br/>(How well the system must perform: LCP, a11y, MTTR)"]
        Const["Constraints<br/>(Budget, team size, legacy APIs, legal compliance)"]
        
        Req --- Dec["Architecture Decision<br/>(Selected structure & accepted trade-offs)"]
        Qual --- Dec
        Const --- Dec
    end

Confusing a preference with a constraint produces unnecessary architecture.


Architecture is contextual

The same question can have different answers:

WebSocket for collaboration → plausible
WebSocket for a static page → unnecessary
global store for a workflow → perhaps useful
global store for one modal → excessive

Context is part of correctness.


Start with the problem, not the tool

Weak question:

Should we use framework X?

Stronger question:

How should a public, searchable, frequently updated catalogue render and recover?

The tool choice follows the problem model.


Technology selection is downstream

flowchart LR
    Forces["Requirements + Quality + Constraints"] --> Boundaries["Boundaries & Operating Model"]
    Boundaries --> Candidates["Evaluate 3+ Technology Candidates"]
    Candidates --> Spike["Empirical Spike & Evidence"]
    Spike --> Selection["Committed Technology Selection"]

Choosing a library before defining the need turns a decision into a justification exercise.


Architecture decisions are often about boundaries

Decide where to place:

  • state;
  • trust;
  • rendering;
  • ownership;
  • caching;
  • deployment;
  • failure;
  • team responsibility.

Boundaries determine what can change independently.


Good boundaries reduce change cost

change in one boundary
  → small, predictable set of affected consumers

The boundary should expose a stable contract and keep volatile implementation private.


Cohesion and coupling

flowchart LR
    subgraph HighCohesion["High Cohesion (Desirable)"]
        direction TB
        C1["Route Logic"] <--> C2["UI View"]
        C2 <--> C3["Local State"]
        Note1["Changes together for one feature"]
    end

    subgraph LowCoupling["Low Coupling (Desirable)"]
        direction LR
        FeatureA["Catalogue Feature"] <-- Narrow API Contract --> FeatureB["Checkout Feature"]
        Note2["Changes in A do not break B"]
    end

Good architecture seeks high cohesion inside meaningful units and intentional coupling between them.


Cohesion example

catalogue feature
  query state
  product query
  filtering
  catalogue UI
  catalogue tests

These concerns may belong together because they change around the same user capability.


Coupling example

ProductCard → entire global store → unrelated application services

The card now knows more than its responsibility requires.

Pass a product and an intent-oriented capability instead.


Dependency direction

Depend toward stability:

flowchart TD
    App["Application Entrypoint / Shell<br/>(Most Volatile)"] --> Feat["Feature Modules (Catalogue, Permits)"]
    Feat --> Domain["Domain Rules & State Reducers"]
    Domain --> Found["Shared Foundation & Design Tokens<br/>(Most Stable)"]

Dependencies should flow toward stable, reusable concepts.

When a foundation imports application behavior, the architecture becomes difficult to reuse and change.


Stable dependencies

Depend on:

  • small interfaces;
  • domain contracts;
  • platform capabilities;
  • stable shared primitives.

Avoid depending on:

  • private files;
  • temporary implementations;
  • broad containers;
  • undocumented framework internals.

Stable dependencies reduce future blast radius.


Blast radius

Blast radius asks:

If this changes or fails, what else is affected?

Reduce blast radius through:

  • narrow APIs;
  • isolated data and runtime boundaries;
  • feature flags;
  • fallback behavior;
  • independent tests and deployment;
  • clear ownership.

Centralization versus local autonomy

Centralize when consistency, security, or shared lifecycle matters.

Keep local when:

  • behavior is domain-specific;
  • requirements differ;
  • coordination cost exceeds reuse value;
  • independent evolution is important.

The right location is a trade-off, not a moral position.


The complexity budget

Every system spends complexity on:

  • code;
  • tools;
  • runtime;
  • deployment;
  • operations;
  • team coordination;
  • cognitive load.

Spend complexity only where it buys a required quality attribute.


Accidental versus essential complexity

essential → required by the product or environment
accidental → introduced by our chosen solution

Architecture work should remove accidental complexity without pretending essential complexity can disappear.


Complexity has several forms

code complexity
runtime complexity
deployment complexity
organizational complexity
conceptual complexity

A “simpler” client architecture may move complexity into the server or CI system.

Count the whole system.


Prefer reversible decisions

When information is weak, prefer choices that can be changed without rewriting the entire platform.

Examples:

  • explicit adapter around a vendor;
  • route-level boundary before full micro-frontends;
  • local state before global store;
  • package API before runtime federation.

Reversibility buys learning time.


One-way and two-way doors

Categorize decisions by reversibility:

flowchart LR
    subgraph TwoWay["Two-Way Door (Reversible)"]
        direction TB
        D1["Decision: UI styling library / Local state tool"]
        D1 --> E1["Easy to migrate or revert"]
        D1 --> A1["Rule: Decide quickly, test in production"]
    end

    subgraph OneWay["One-Way Door (Hard to Reverse)"]
        direction TB
        D2["Decision: Micro-frontends / Core database schema"]
        D2 --> E2["Costly, multi-month migration to undo"]
        D2 --> A2["Rule: Require spikes, ADRs & executive sign-off"]
    end

Do not apply heavyweight governance to every reversible choice.

Do not rush a decision that creates irreversible migration or data cost.


Delay irreversible decisions when information is weak

Use:

  • a spike;
  • a prototype;
  • a compatibility layer;
  • a small route;
  • a feature flag;
  • an explicit revisit trigger.

Learning before commitment is architecture work.


YAGNI does not mean ignore the future

Avoid building speculative systems for unknown requirements.

Do preserve:

  • clear boundaries;
  • migration paths;
  • explicit ownership;
  • replaceable dependencies;
  • observable assumptions.

The future is supported by flexibility, not by implementing every possibility now.


Premature abstraction

An abstraction created before variation is understood may encode the wrong concept.

Wait for evidence about:

  • repeated behavior;
  • stable vocabulary;
  • change direction;
  • real consumers;
  • shared accessibility and lifecycle.

The rule of three

One use may be local.

Two uses reveal similarity.

Three uses can provide enough evidence to decide whether the concept is genuinely shared.

It is a heuristic, not a command to duplicate exactly three times.


Duplication can be cheaper than coupling

Small duplicated code may be safer than:

  • a generic API with many modes;
  • a shared release dependency;
  • cross-domain assumptions;
  • a package nobody can evolve.

Compare maintenance and change cost rather than counting lines.


Locality

Locality keeps related code, data, tests, and ownership near each other.

It reduces:

  • navigation cost;
  • hidden dependencies;
  • coordination overhead;
  • accidental reuse.

Move code outward only when the new boundary creates real value.


Explicit data flow

source → transform → consumer

Explicit inputs, outputs, events, URLs, and requests are easier to test and change than invisible reads from global context.


Hidden coupling

Hidden coupling appears through:

  • global mutable state;
  • import-time side effects;
  • shared storage keys;
  • undocumented events;
  • CSS selectors across packages;
  • framework-specific assumptions;
  • environment variables used everywhere.

Make coupling visible or remove it.


Global state is an architectural decision

Use global state when:

  • lifetime is application-wide;
  • consumers are genuinely distributed;
  • synchronization is needed;
  • ownership is explicit.

Do not use it to avoid passing one value through a well-defined local boundary.


Server state is not ordinary client state

Server state has:

  • remote authority;
  • freshness;
  • errors;
  • caching;
  • invalidation;
  • concurrency.

Use a server-aware boundary rather than copying it into every client store.


Derived state should remain derived

products + query → visibleProducts

Store the inputs and calculate the result.

Duplicated derived state creates another source of truth and another synchronization path.


URL as architecture

URL state provides:

  • shareability;
  • reload persistence;
  • browser history;
  • deep links;
  • route ownership.

It should contain meaningful public view state, not secrets or every ephemeral interaction.


Progressive enhancement as a layered model

semantic HTML
  → server behavior
  → enhanced client behavior
  → optional richer capabilities

Layering can improve resilience and reduce the amount of functionality that depends on one runtime.


When progressive enhancement is valuable

It is especially useful for:

  • public content;
  • forms and navigation;
  • poor connectivity;
  • accessibility;
  • long-lived pages;
  • critical tasks.

Not every application needs full no-JavaScript operation, but every app should understand its essential failure path.


Native platform first

Before adding a dependency, ask whether the platform already provides:

  • links;
  • forms;
  • dialog;
  • details/summary;
  • URL and history;
  • storage;
  • fetch;
  • observers;
  • workers.

Platform capabilities often have strong accessibility, performance, and compatibility foundations.


Native does not automatically mean better

Evaluate:

  • browser support;
  • accessibility behavior;
  • required customization;
  • interaction consistency;
  • testing;
  • polyfills and fallbacks;
  • team familiarity.

Use the platform deliberately, not ideologically.


Dependency evaluation

Consider:

  • capability fit;
  • bundle and runtime cost;
  • security and maintenance;
  • accessibility;
  • license and ecosystem;
  • upgrade path;
  • lock-in;
  • team skill.

The smallest dependency is not always the cheapest total solution.


Build, buy, or adopt

build → unique, strategic, or tightly domain-specific capability
buy   → commodity where support is valuable
adopt  → mature open source with acceptable control and risk

Compare total lifecycle cost, not only initial implementation time.


Core versus commodity

Protect and understand capabilities that differentiate the product.

Avoid spending strategic team capacity reinventing commodity infrastructure unless the trade-off is intentional.

Also avoid outsourcing a core capability that determines security, user trust, or domain advantage without sufficient control.


Framework selection is contextual

Evaluate frameworks through:

  • rendering and data needs;
  • team familiarity;
  • ecosystem and support;
  • accessibility;
  • testing;
  • deployment;
  • upgrade path;
  • performance on target users.

Benchmark claims without context are weak evidence.


Framework benchmarks are contextual

Results vary by:

  • application shape;
  • route size;
  • data and interaction;
  • device and network;
  • build configuration;
  • developer usage.

Run a spike with representative behavior instead of selecting a framework from a leaderboard.


React versus Vue is often not the main decision

The larger questions may be:

  • where server and client boundaries lie;
  • how state is owned;
  • how routes are delivered;
  • how APIs are validated;
  • how teams release;
  • how failures recover.

Framework syntax rarely determines all architecture.


Framework lock-in

Lock-in can be acceptable when the framework provides strong value and the team accepts its lifecycle.

Reduce unnecessary lock-in at boundaries with:

  • domain models;
  • adapters;
  • platform APIs;
  • stable package contracts;
  • route and API boundaries.

Avoid turning lock-in avoidance into architecture by itself.


Architecture fitness

Fitness asks whether the system continues to support important properties as it evolves.

Examples:

  • no dependency cycles;
  • route budget stays within threshold;
  • public package imports only;
  • all remote responses are validated;
  • critical journey remains keyboard-usable.

Evolutionary architecture

Instead of assuming the initial architecture is perfect:

decision → guardrail → measure → learn → adjust

Architecture can evolve safely when important properties are observable and protected.


Fitness functions

Automate checks that preserve properties continuously:

flowchart LR
    Code["Pull Request Code Push"] --> Linter["ESLint Boundary Rules<br/>(No feature-to-feature imports)"]
    Code --> Budget["Bundle Budget Checker<br/>(Initial route < 180 KB)"]
    Code --> Schema["Contract Validator<br/>(Validates Zod schemas)"]
    Linter --> Gate{"CI Gate"}
    Budget --> Gate
    Schema --> Gate
    Gate -->|All Pass| Merge["Merge Approved"]
    Gate -->|Any Fail| Block["Block PR Deployment"]

Automate rules that matter enough to protect continuously.


Architecture rule as code

Rules can be enforced through:

  • lint boundaries;
  • package exports;
  • dependency graph checks;
  • CI budgets;
  • contract tests;
  • runtime monitoring.

Written policy that cannot be observed often becomes forgotten policy.


Fitness functions should protect important properties

Do not create checks for every preference.

Protect properties whose regression would cause meaningful harm:

  • security;
  • accessibility;
  • performance;
  • deployability;
  • dependency direction;
  • public contract compatibility.

Architecture Decision Records

An ADR records:

  • context;
  • decision;
  • alternatives;
  • consequences;
  • validation;
  • status;
  • revisit trigger.

It preserves reasoning after the original meeting is forgotten.


ADRs are not meeting minutes

An ADR should not capture every discussion detail.

It should explain:

what problem existed
what was chosen
why it was chosen
what trade-offs were accepted
when to revisit

Example ADR: URL state for catalogue filters

flowchart LR
    URL["URL Query Params<br/>(?category=permits&page=2)"] --> Router["Router State Hook"]
    Router --> Search["Search Input Component"]
    Search --> API["Fetch API Client"]
    API --> Results["Display Results Table"]

Worked decision: Civic Platform Architecture

Context & Forces (Erbil Citizen Portal):

  • 4 autonomous product squads (Health, Transport, Commerce, Education).
  • Target users: 70% mobile browsers on congested 3G/4G networks; WCAG 2.1 AA legal mandate.
  • High SEO requirement for public municipal announcements and legal circulars.
flowchart TD
    Req["Citizen Portal Forces"] --> F1["Autonomy for 4 Squads"]
    Req --> F2["Fast Mobile LCP < 2.0s over 3G"]
    Req --> F3["Zero Accessible Regressions"]
    Req --> F4["SEO for Legal Circulars"]

Candidate architectures: Rejected options

Candidate OptionArchitecture ModelReason for Rejection
Option A: Micro-FrontendsWebpack Module Federation; multi-repo independent deploys4.2s mobile LCP on 3G: duplicate React vendor runtimes violate performance budget
Option B: Client-Side SPAPure CSR single-page app; static CDN hostingSEO & blank screen: fails public legal circular indexing and initial 3G render

Candidate architectures: Accepted decision

Option C: Modular Monolith with Edge SSR (ACCEPTED)

Architectural DimensionStrategy & Evaluation
Mobile LCP1.4s (p75): cached semantic HTML at CDN edge
Search Engine Indexing100% crawlable: complete server-rendered document
Squad Autonomypnpm monorepo with strict package.json "exports"
Operational OverheadSingle container pipeline; zero distributed federation complexity

Worked decision: Evidence that would reverse it

The decision to adopt a Modular Monolith with Edge SSR will be formally revisited if:

flowchart TD
    Rev["Reversal Triggers (ADR-018)"]
    Rev --> T1["Team Scale: Engineering squads grow from 4 to >12 squads<br/>and CI queue times exceed 30 minutes"]
    Rev --> T2["Regulatory Mandate: Ministry of Education mandates<br/>hosting in an independent sovereign data center"]
    Rev --> T3["Traffic Spikes: SSR compute costs exceed budget<br/>requiring static HTML export for catalog routes"]

The ADR makes ownership and trade-offs explicit.


ADR status

stateDiagram-v2
    [*] --> Proposed: Drafted by Engineer
    Proposed --> Accepted: Team Architectural Consensus
    Proposed --> Rejected: Fails Constraints / Trade-offs
    Accepted --> Superseded: Replaced by Newer ADR
    Accepted --> Deprecated: Capability Retired
    Superseded --> [*]
    Deprecated --> [*]
    Rejected --> [*]

Status tells readers whether the decision is active and whether another document replaces it.


Decision scope

Record what the decision does and does not cover.

scope: catalogue filter state
not scope: all application state or authentication

Clear scope prevents a local decision from becoming accidental universal policy.


Decision matrices

A matrix can expose trade-offs:

OptionFreshnessComplexityReversibilityTeam fit
Ahighmediumhighhigh
Bmediumhighlowmedium

Use it to structure reasoning, not to pretend qualitative judgment is exact arithmetic.


Avoid weighted-score theater

Numbers can create false certainty when:

  • criteria are subjective;
  • weights are arbitrary;
  • unknowns are hidden;
  • important risks are averaged away.

Show the assumptions and discuss the decisive trade-offs directly.


Proof of concept and spike

flowchart LR
    subgraph Spike["Technical Spike"]
        Q1["Question: Does Rolldown bundle this in <500ms?"] --> Code1["Write 30-line test script"]
        Code1 --> Ans1["Answer: Yes (420ms). Spike discarded."]
    end

    subgraph PoC["Proof of Concept"]
        Q2["Question: Can Edge SSR integrate with legacy auth?"] --> Code2["Build skeleton prototype with 2 routes"]
        Code2 --> Ans2["Validate feasibility across systems"]
    end

Keep the purpose narrow and record what was learned.

Do not mistake experimental code for a production architecture.


Architecture review

Review asks:

  • does the decision fit the requirements?
  • are quality attributes explicit?
  • are failure modes understood?
  • is the boundary appropriate?
  • can it be operated and changed?
  • what evidence remains weak?

It should improve the decision, not only approve a document.


Decision confidence

State confidence and uncertainty:

high confidence → well-understood constraint and evidence
low confidence  → important unknown remains

Low confidence should trigger a spike, guardrail, or revisit date.


Assumption register

Track assumptions such as:

  • expected traffic;
  • team capacity;
  • API stability;
  • browser support;
  • deployment independence;
  • data sensitivity;
  • user behavior.

An assumption that changes can invalidate the decision.


Architecture risk and risk reduction

risk → assumption → experiment / guardrail → evidence → decision update

Reduce uncertainty early when it could make the chosen architecture expensive to reverse.


Failure mode analysis

For a feature, ask:

  • what can fail?
  • how will the user see it?
  • what remains usable?
  • can the operation retry?
  • can the release roll back?
  • who is alerted?

Failure behavior is part of the architecture, not a later polish step.


Graceful degradation

full capability → reduced capability → understandable failure

Examples:

  • cached content when the network fails;
  • static navigation when enhancement fails;
  • panel fallback when one remote fails;
  • queued draft when submission is offline.

Critical path

Identify the work required for the user’s first valuable outcome.

Keep optional features, telemetry, personalization, and secondary data off the critical path when possible.

Critical-path design affects rendering, performance, reliability, and architecture together.


Availability architecture

Availability includes:

  • resilient requests;
  • cache and fallback;
  • independent failure boundaries;
  • deployment and rollback;
  • monitoring;
  • recovery ownership.

“The server is up” does not prove that the user can complete the task.


Security architecture

Define:

  • trust boundaries;
  • identity and authorization;
  • secret location;
  • content and request validation;
  • browser policies;
  • third-party trust;
  • logging and privacy.

Security belongs in the decision, not in a final header checklist.


Performance architecture

Make choices about:

  • rendering topology;
  • JavaScript responsibility;
  • cacheability;
  • code splitting;
  • data waterfalls;
  • device cost;
  • route and journey budgets.

Performance constraints should influence boundaries before slow code accumulates.


Accessibility architecture

Accessibility is affected by:

  • semantic HTML;
  • component contracts;
  • focus ownership;
  • routing and navigation;
  • progressive enhancement;
  • design-system governance;
  • testing and release policy.

Treat it as a system property, not only a component task.


Internationalization architecture

Plan for:

  • locale and direction;
  • text expansion;
  • date and number formatting;
  • routing and content negotiation;
  • font coverage;
  • translation loading;
  • persistence and URL state.

Adding i18n late often reveals hidden assumptions across every layer.


Testability as an architecture attribute

Testability improves when the system has:

  • explicit inputs and outputs;
  • replaceable boundaries;
  • deterministic transitions;
  • isolated ownership;
  • observable failures;
  • stable contracts.

If architecture makes the important behavior impossible to isolate, testing cost is design feedback.


Observability as an architecture attribute

Make it possible to answer:

  • which release failed?
  • which route and user journey?
  • which dependency or boundary?
  • which users are affected?
  • can the issue be mitigated or rolled back?

Observability should be designed into boundaries and identifiers.


Deployability as an architecture attribute

A deployable system has:

  • build reproducibility;
  • artifact identity;
  • safe configuration;
  • progressive delivery;
  • compatibility strategy;
  • rollback;
  • operational ownership.

An architecture that cannot be released safely limits product change.


Maintainability, reliability, scalability

maintainability → can we change it?
reliability     → does it continue to work?
scalability     → does it work as demand or scope grows?

These properties interact and should be evaluated together.


Different kinds of scale

user scale
data scale
team scale
feature scale

A system may scale users while failing teams, or scale features while becoming impossible to reason about.

Identify which scale is actually driving the decision.


Cognitive load

Architecture should reduce the amount a developer must understand at once.

Use:

  • local reasoning;
  • clear names;
  • stable boundaries;
  • focused modules;
  • useful documentation;
  • predictable workflows.

The fastest system to change is often the one with the clearest mental model.


Architecture documentation

Useful documentation includes:

  • context and system diagrams;
  • ADRs;
  • package and dependency maps;
  • ownership;
  • runbooks;
  • contracts;
  • failure and recovery behavior.

Document decisions that future maintainers need to preserve or revisit.


C4-style thinking

Move through levels of architectural abstraction:

flowchart TD
    L1["1. System Context<br/>Citizens, Portal, External Auth, Ministry APIs"] --> L2["2. Containers<br/>Web App SPA, Edge SSR Gateway, Redis Cache"]
    L2 --> L3["3. Components<br/>Permit Catalogue, Routing Shell, Auth Provider"]
    L3 --> L4["4. Code<br/>TypeScript Classes, Functions, React/Vue Components"]

Use the level that answers the current question.

Do not create diagrams so detailed that nobody can use them.


Architecture diagram purpose

A diagram should help someone:

  • understand ownership;
  • find a trust boundary;
  • trace data flow;
  • identify failure containment;
  • evaluate a change;
  • operate the system.

If it only decorates a presentation, it is not architecture documentation.


Avoid architecture astronautics

Warning signs:

  • abstractions with no current consumer;
  • patterns chosen for prestige;
  • distributed systems without distributed requirements;
  • infrastructure whose operation exceeds its value;
  • diagrams detached from code and ownership.

Complexity is not evidence of maturity.


Simplicity is a feature

Simple architecture can provide:

  • faster onboarding;
  • easier debugging;
  • fewer failure modes;
  • lower operating cost;
  • more reversible change.

Simple does not mean primitive or careless.


Standards before custom infrastructure

Prefer platform and established protocols when they meet the requirement:

  • HTTP;
  • URL and history;
  • HTML forms;
  • browser storage;
  • standard security headers;
  • common observability formats.

Custom infrastructure should solve a demonstrated gap.


Libraries should amplify the platform

A library is valuable when it:

  • reduces repeated risk;
  • provides missing capability;
  • preserves accessibility;
  • improves productivity;
  • remains replaceable enough.

It is risky when it hides basic browser behavior the team needs to understand.


Architecture and team skill

Choose a design the team can:

  • explain;
  • debug;
  • test;
  • deploy;
  • operate;
  • migrate.

An architecture that requires one specialist for every incident has a bus-factor problem.


Hiring and onboarding

Architecture affects:

  • how quickly new engineers understand the system;
  • how safely they make changes;
  • how much tribal knowledge is required;
  • how ownership is transferred.

Good boundaries and documentation are productivity tools.


Bus factor and ownership

Reduce dependence on one person through:

  • shared runbooks;
  • pair or review practice;
  • explicit package ownership;
  • reproducible workflows;
  • decision records;
  • operational rehearsal.

Ownership should create accountability, not private territory.


Guardrails versus gates

guardrail → guides or prevents common unsafe behavior
gate      → blocks progress until a condition is met

Use gates for security, compatibility, or release-critical requirements.

Use guardrails for recommended patterns where local judgment remains useful.


Paved roads

A paved road provides:

  • safe defaults;
  • templates;
  • supported tooling;
  • examples;
  • deployment and observability;
  • extension points.

It should reduce cognitive load without pretending every product has the same needs.


Architecture debt

Architecture debt is accumulated cost from decisions that no longer fit:

  • obsolete boundaries;
  • excessive coupling;
  • stale platform assumptions;
  • missing migrations;
  • unsupported dependencies;
  • unclear ownership.

Track it and prioritize by change cost and user impact.


Refactoring architecture

Refactor toward boundaries through small steps:

observe coupling → define contract → move one consumer → verify → repeat

Preserve behavior while changing structure.


Strangler strategy revisited

Replace one capability at a time while the old system continues to serve the rest.

The migration needs:

  • routing or proxy boundary;
  • data compatibility;
  • shared authentication;
  • observability;
  • rollback;
  • ownership during coexistence.

Big-bang rewrite risk

A rewrite can:

  • delay user value;
  • lose hidden behavior;
  • recreate old mistakes;
  • require long coexistence anyway;
  • concentrate technical and organizational risk.

Rewrite only when the expected value and migration control justify the discontinuity.


When a rewrite is justified

Possible evidence:

  • the current platform cannot meet a required constraint;
  • incremental migration is more expensive than replacement;
  • ownership and product scope are stable;
  • behavior is well understood;
  • rollout and rollback are credible;
  • the organization can sustain the work.

“The code feels old” is not enough evidence by itself.


Technical-debt prioritization

Prioritize by:

  • frequency of change;
  • blast radius;
  • user harm;
  • security or compliance;
  • delivery friction;
  • probability of failure;
  • reversibility.

Debt that is stable and harmless may be lower priority than a small coupling point that blocks every release.


Cost of change

Measure:

  • how many files or packages change;
  • how many teams coordinate;
  • how many environments verify;
  • how many tests break;
  • how much release work is needed;
  • how hard rollback is.

Change amplification is architecture evidence.


Scenario: public documentation

Likely priorities:

  • cacheability;
  • discoverability;
  • accessibility;
  • content publishing;
  • low client cost.

Static or revalidated output with focused enhancement may be a strong fit.


Scenario: internal admin

Likely priorities:

  • authentication;
  • dense interaction;
  • productivity;
  • resilient forms;
  • domain workflows;
  • observability.

A modular client application or hybrid route may be simpler than a public-content topology.


Scenario: e-commerce

Different routes can need different strategies:

marketing → static
catalogue  → cacheable hybrid
checkout   → secure interactive flow
account    → personalized server/client boundary

One application does not require one rendering strategy.


Scenario: collaborative editor

Priorities may include:

  • real-time updates;
  • conflict handling;
  • offline drafts;
  • presence;
  • low interaction latency;
  • strong domain state.

The architecture should begin with synchronization and failure requirements, not a component library choice.


Scenario: government service

Likely priorities:

  • accessibility;
  • reliability;
  • security;
  • auditability;
  • long support life;
  • progressive enhancement;
  • clear recovery.

Operational simplicity and standards may matter more than fashionable distribution.


Scenario: large enterprise platform

Possible pressures include:

  • team autonomy;
  • shared foundations;
  • independent release cadence;
  • multiple domains;
  • legacy migration;
  • governance and security.

Use explicit contracts and incremental separation before reaching for runtime federation.


No architecture is context-free

A decision without context is a slogan.

Record:

  • product;
  • users;
  • team;
  • constraints;
  • quality priorities;
  • deployment model;
  • expected change;
  • evidence and unknowns.

Resume-driven architecture

This occurs when a technology is chosen mainly because it is interesting or marketable.

Counter it with:

  • explicit requirements;
  • measurable constraints;
  • a representative spike;
  • total lifecycle cost;
  • a revisit trigger.

Technology can be valuable without being the reason for the decision.


Cargo-cult architecture

Copying a pattern from another company without its context can import:

  • unnecessary services;
  • incompatible team assumptions;
  • different scale costs;
  • hidden operational requirements.

Learn the principle, then re-evaluate it against your constraints.


Framework as architecture

framework → routing, rendering, components, build mechanisms
architecture → ownership, boundaries, data, quality, operations

The framework participates in architecture but does not replace system decisions.


Pattern collection

Using every pattern creates:

  • duplicated concepts;
  • conflicting state paths;
  • more documentation;
  • harder onboarding;
  • ambiguous ownership.

Patterns should solve a problem that exists in the system.


Hidden architecture

Architecture is hidden when behavior depends on:

  • undocumented globals;
  • convention nobody can find;
  • private package imports;
  • manual deployment steps;
  • tribal knowledge;
  • unrecorded exceptions.

Make important decisions visible in code, checks, docs, and ownership.


Architecture freeze

Architecture should be stable enough to guide work but revisable when evidence changes.

Use:

  • ADR status;
  • review triggers;
  • fitness functions;
  • migration paths;
  • measured outcomes.

“We decided once” is not a reason to ignore new constraints.


Architecture review cadence

Review architecture when:

  • product scope changes;
  • team topology changes;
  • a quality attribute regresses;
  • an integration becomes a bottleneck;
  • a dependency reaches end of life;
  • a migration creates new boundaries.

Review should be triggered by change, not only by calendar ceremony.


Measure architecture outcomes

Possible outcomes include:

  • change lead time;
  • deployment failure and recovery;
  • accessibility defects;
  • performance budgets;
  • dependency cycle count;
  • onboarding time;
  • support burden;
  • team autonomy in practice.

Architecture quality is visible through system behavior.


Architecture is sociotechnical

Technology and organization interact through:

  • ownership;
  • incentives;
  • communication;
  • release authority;
  • expertise;
  • support expectations.

A design that ignores people will be changed by people in ways the diagram did not predict.


Architecture and product lifecycle

Prototype, growth, maturity, and retirement may need different priorities.

prototype → learning speed
growth    → boundaries and scale
maturity  → reliability and maintenance
retirement → migration and safe shutdown

Do not optimize a prototype for the operational needs of a global platform without evidence.


Architecture and deadlines

Deadlines do not eliminate architecture decisions.

Make the trade-off explicit:

  • what is deferred;
  • what risk is accepted;
  • what boundary remains safe;
  • what follow-up is required;
  • what cannot be compromised.

Shortcuts are safer when named and bounded.


Architecture and future change

Do not attempt to predict every future feature.

Instead, preserve:

  • clear ownership;
  • replaceable boundaries;
  • stable contracts;
  • observable assumptions;
  • reversible choices where possible.

Good architecture makes learning and change affordable.


The decision framework

1 define problem
2 identify quality attributes
3 record constraints
4 generate alternatives
5 analyze trade-offs
6 test unknowns
7 decide
8 record ADR
9 add guardrails
10 observe
11 revisit

This turns architecture into a repeatable practice.


Full ADR template

Status
Context
Decision
Alternatives considered
Consequences
Validation
Revisit trigger

Keep the document concise enough to read and specific enough to preserve reasoning.


Revisit triggers

Examples:

  • traffic exceeds the assumed range;
  • team ownership changes;
  • deployment becomes a bottleneck;
  • performance budget fails;
  • a dependency is retired;
  • a security requirement changes;
  • evidence from a spike contradicts an assumption.

Triggers turn an ADR into a living decision rather than a permanent verdict.


Example: catalogue state decision

URL owns committed filters.
Local state owns transient input.
Server cache owns remote results.
Derived state owns visible rows.

The decision improves reload, sharing, caching, and testability without requiring a global store.


Example: rendering decision

public home → revalidated static
catalogue   → cacheable server/hybrid route
account     → personalized boundary
admin       → interaction-focused client route

The architecture is route-specific because requirements differ.


Example: design-system decision

Share:

  • tokens;
  • accessible primitives;
  • stable interaction patterns.

Keep local:

  • domain workflows;
  • product-specific orchestration;
  • experimental layouts.

Promote concepts only when stability and ownership justify it.


Example: micro-frontend decision

Choose runtime separation only when:

  • independent deployment is a real bottleneck;
  • capability boundaries are stable;
  • contracts and observability exist;
  • failure containment matters;
  • the organization can operate the integration.

Otherwise start with packages or a modular monolith.


An architectural “no”

Saying no can protect the system:

No runtime federation yet.
No global store for this local state.
No public package for a one-off component.
No client exposure of this secret.

An architectural no should include the reason and the condition that could change it.


An architectural “yes”

A good yes is bounded:

Yes, use a shared package for stable accessibility primitives.
Scope: three products, public entry points, versioned releases.
Review: after two migration cycles or a major API change.

Boundaries make adoption safer.


The browser is still the foundation

Even sophisticated systems eventually interact with:

  • HTML;
  • CSS;
  • JavaScript;
  • URLs;
  • HTTP;
  • browser security;
  • input, focus, and rendering.

Architecture should amplify platform capabilities rather than hide them completely.


HTML, CSS, and JavaScript are architecture

HTML → document, semantics, navigation, forms
CSS  → layout, themes, responsiveness, visual stability
JS   → interaction, state, synchronization, runtime behavior

Choices at the platform layer shape accessibility, performance, testing, and resilience.


Testing and observability are architecture feedback

Tests reveal whether boundaries are behaviorally useful.

Observability reveals whether boundaries work under real users, networks, devices, and releases.

Use both to decide whether architecture is serving its purpose.


Architecture is continuous

decide → implement → observe → learn → refactor → decide again

The goal is not a final perfect diagram.

The goal is a system that can change without losing control.


Practical lab: Make and Defend an Architecture Decision

Create an evidence-backed architecture decision for a product platform without ranking frameworks or adopting complexity by default.

The practical ends with an ADR, fitness functions, failure modes, migration path, and decision report.


Practical stages 1–5: define the problem

  1. Define the product.
  2. List functional requirements.
  3. Rank quality attributes.
  4. Record constraints.
  5. Define state ownership.

State the user, team, deployment, security, performance, and organizational context.


Practical stages 6–12: design the boundaries

  1. Choose rendering topology.
  2. Choose component boundaries.
  3. Decide on shared state.
  4. Choose API boundary strategy.
  5. Define security boundaries.
  6. Define performance constraints.
  7. Define accessibility requirements.

Treat each as a decision with trade-offs, not a framework checkbox.


Practical stages 13–19: evaluate the platform

  1. Evaluate dependencies.
  2. Evaluate framework fit.
  3. Decide repository structure.
  4. Evaluate micro-frontends.
  5. Define testing layers.
  6. Define delivery.
  7. Define observability.

Prefer the smallest coherent solution that satisfies the constraints.


Practical stages 20–24: reduce uncertainty

  1. Create two alternatives.
  2. Run a technical spike.
  3. Write the ADR.
  4. Add fitness functions.
  5. Define failure modes.

The spike should answer the highest-risk unknown, not build the entire future system.


Practical stages 25–27: migration and defense

Practical stages 1–3: constraints, options & empirical spike

  1. Stage 1 (Explicit Problem & Constraint Mapping):
    • Document functional goals and prioritize non-negotiable quality attributes.
  2. Stage 2 (Formulating Three Viable Candidate Architectures):
    • Candidate A (Micro-Frontends), Candidate B (Client-Side SPA), Candidate C (Modular Monolith with Edge SSR).
  3. Stage 3 (The Investigative Technical Spike):
    • Run a benchmark spike measuring bundle size and p75 LCP under 4x CPU throttling.

Practical stages 4–5: ADR & reversal plan

  1. Stage 4 (Drafting the Formal ADR):
    • Record Title, Status, Context, Decision, Consequences, and Automated Fitness Functions.
  2. Stage 5 (Reversal Plan & Review Trigger Conditions):
    • Define exact quantitative thresholds (team size >12 squads, CI queue >30m) that trigger architectural review.

Verification: Decisions reflect empirical evidence and trade-offs rather than technology fashion.


Try this yourself

Choose one architecture decision and write:

problem:
quality priorities:
constraints:
alternatives:
unknown:
spike:
decision:
consequences:
revisit trigger:

If the decision cannot name a problem, it may be technology fashion rather than architecture.


Troubleshooting guide (Part 1)

SymptomLikely cause
Architecture debate never endsRequirements and decision criteria are unclear
Every solution is globalOwnership and locality were not evaluated
Micro-frontends are proposed immediatelyOrganizational pressure was mistaken for runtime need
ADRs are long and unreadThey record meetings instead of decisions
Fitness checks are ignoredThey protect preferences rather than important properties

Troubleshooting guide (Part 2)

SymptomLikely cause
Rewrite feels safer than migrationHidden behavior and compatibility cost are underestimated
Shared package changes break everyonePublic API and versioning are weak
Architecture depends on one expertOwnership, documentation, and runbooks are insufficient
“Simple” system fails at scaleThe relevant quality attribute was not measured

Completion checklist

  • the problem and context are explicit;
  • quality attributes are prioritized and measurable;
  • constraints are distinguished from preferences;
  • alternatives and trade-offs are documented;
  • unknowns are tested with focused spikes;
  • boundaries reduce change cost and blast radius;
  • important architecture properties have guardrails;
  • failure, security, performance, and accessibility are included;
  • migration and rollback are credible;
  • the decision has an owner and revisit trigger.

Misconceptions to leave behind (Part 1)

MisconceptionBetter mental model
Architecture is the technology stackIt is contextual decisions and boundaries
A good architecture optimizes everythingIt makes explicit trade-offs
More abstraction is betterAbstraction has coupling and cognitive cost
Duplication is always badLocal duplication can preserve autonomy
Global state is required for scaleOwnership and lifetime determine scope
SSR is more architectural than CSRRendering is one contextual decision
Micro-frontends are the natural futureDistribution is justified by real independence needs

Misconceptions to leave behind (Part 2)

MisconceptionBetter mental model
Every shared component belongs centrallyStable shared concepts deserve promotion
ADRs are bureaucracyThey preserve reasoning and revisit conditions
A decision cannot changeArchitecture should learn from evidence
A rewrite is cleanerIncremental migration often reduces risk
Technical debt must always be removedPrioritize by impact and change cost
Simplicity means underengineeringSimplicity can be deliberate, bounded architecture

The chapter in one sentence

Make architecture decisions from requirements, quality attributes, constraints, evidence, and reversible boundaries - and keep revisiting them as the system and organization learn.


Course completion: Capstone architecture

Congratulations on completing all 18 chapters of modern web application engineering!

Next steps to master front-end architecture:

  • Capstone Architecture Project: Build and defend an end-to-end civic application platform;
  • Appendix A: Review the Architectural Rosetta Stone (React vs. Vue mechanics);
  • Appendix B: Consult the Modern Browser APIs Reference;
  • Appendix C: Run the Production Deployment Checklist before every release.

Questions

Which architecture decision in your system is treated as permanent even though its assumptions, evidence, or constraints have already changed?