This is the multi-page printable view of this section. .
Lecture slides
- 1: The Modern Web Platform & Browser Internals
- 2: Semantic HTML, Accessibility, Internationalization & the DOM
- 3: Modern CSS Architecture & Layout Systems
- 4: Modern JavaScript & Asynchronous Programming
- 5: TypeScript, Runtime Contracts & Safe Data Boundaries
- 6: Component-Driven Architecture & Design Patterns
- 7: Reactivity & Rendering Mechanics
- 8: State Management, Routing & Form Architecture
- 9: Client–Server Communication, APIs & Cache Management
- 10: Real-Time Communication, Offline Systems & Client Persistence
- 11: Rendering Topologies: CSR, SSR, SSG & Beyond
- 12: Modern Build Systems, Development Tooling & Team Workflows
- 13: Front-End Security, Authentication & Browser Isolation
- 14: Scaling Front-End Architecture: Design Systems, Monorepos & Micro-Frontends
- 15: Core Web Vitals & Performance Engineering
- 16: Testing Strategies for Resilient Interfaces
- 17: Continuous Delivery, Observability & Maintenance
- 18: Front-End Architecture & Technical Decision-Making
Lecture decks for the book. Open a deck to present it, use the arrow keys or space bar to move between slides, and press F for fullscreen. Each deck also prints one slide per landscape page for handouts or PDF export.
1 - The Modern Web Platform & Browser Internals
The Modern Web Platform & Browser Internals
From the first HTML bytes to the next interaction
Chapter 1 · Modern Front-End Engineering
Polla Fattah
The files loaded. Why does the page still feel slow?
A catalogue heading appears before its image.
Clicking Load products appears to do nothing - then everything changes together.
What would you inspect first: the requests, the handler, or the rendering work?
What we will explain
- How the browser discovers resources.
- Why downloading and executing a script are different events.
- How document state becomes a frame.
- Why synchronous code, microtasks, and timers behave differently.
- What evidence DevTools can - and cannot - provide.
One small page
Use the full document in the chapter for the practical.
A browser provides more than a JavaScript engine
The engine executes language code.
The browser also provides networking, the DOM, events, timers, storage, and rendering.
Work can happen concurrently, but ordinary page JavaScript and much rendering work compete for main-thread time.
Navigation has setup costs
For https://shop.example.com/products?page=2:
- Resolve the destination if necessary.
- Establish or reuse a connection.
- Request the document.
- Begin processing available HTML.
Caches and restored pages can change this path. A small resource on a new origin may still incur setup cost.
Parsing and downloading overlap
The browser does not generally wait for the complete HTML before discovering other resources.
The arrows show dependencies, not measured durations.
Discovery can be the delay
versus a URL assigned only when JavaScript runs:
Unless another reference reveals it earlier, the second request depends on script execution. Assigning src can start the request before insertion into the DOM.
A stylesheet can hide another dependency
The browser needs the CSS before discovering the image through that rule.
Discovery, resource use, and request scheduling are different steps. A declared font that is never used need not be downloaded.
Speculative discovery can look ahead
A preload scanner may inspect available markup while the main parser waits for a script.
It can find a later image or script reference.
It cannot execute arbitrary JavaScript to discover its future results or inspect bytes that have not arrived.
Do not assume a universal browser thread arrangement.
Resource hints have different jobs
| Hint | Purpose |
|---|---|
| Preload | Fetch a current-page resource earlier |
| Prefetch | Speculate about likely future use |
| Preconnect | Begin connection setup |
A hint does not make bandwidth unlimited. Measure whether it addresses an actual delay.
Source HTML is not the live DOM
The DOM changes; the original response does not.
Demonstration: compare the document response in Network with the status paragraph in Elements/Inspector.
Classic scripts can stop the parser
Place this in legacy.js in the head:
The heading has not been parsed: false.
Move the script after the heading: it can find the node.
Node availability does not prove that it has painted.
Choose script timing
For external scripts declared in the initial HTML:
| Mode | Main execution condition |
|---|---|
| Classic | Parser waits at the script |
| Defer | After parsing; ordered with deferred classic scripts |
| Async | When ready; document order not guaranteed |
| Module | Deferred by default; dependency graph applies |
async does not mean execution on a background thread.
CSS can delay a script too
A previously encountered stylesheet can block this script.
The parser is waiting for the script, so CSS can indirectly prolong the pause.
CSS blocking presentation and a script blocking parsing are different dependencies.
Readiness is not responsiveness
DOMContentLoaded: parsing and relevant deferred script processing have progressed to this milestone.- Window
load: additional load-delaying resources have completed. - Neither guarantees a particular paint time or fast future interactions.
Module experiments here omit imports and top-level await to keep the comparison focused.
From state to a frame
An update may reuse work or skip stages. This is a conceptual model, not an engine’s thread diagram.
Width and transform do different things
Width can change wrapping and neighboring geometry.
A transform changes visual placement without reallocating normal-flow space.
Some transforms can reuse painted content. That is not a universal performance guarantee.
A read can force layout
Compare with writing all widths first, then reading them.
Reset the starting state. Batching may reduce repeated layout; the first read can still require it.
Predict the output
Write your prediction before running it.
Explain the output
Synchronous work finishes first. The already-fulfilled Promise’s reaction runs at a microtask checkpoint before the timer task.
This is not a claim that all asynchronous operations finish before timers.
A task is not followed by a guaranteed paint
requestAnimationFrame() schedules a callback for a rendering update, not an after-paint notification.
An expensive callback remains expensive.
Why might “Working…” never appear?
Inside the click handler:
Both DOM changes occur. The intermediate state may never reach the screen.
Microtasks are not background work
A microtask can queue another microtask.
The checkpoint keeps processing them, delaying later tasks if the chain grows.
Use the bounded chain in the chapter. Do not freeze the page with an endless loop.
Moving expensive work into a Promise callback does not make it free.
DevTools demonstration: discovery
- Serve the chapter’s small page over local HTTP.
- Record browser, cache, and throttling settings.
- Capture the image request from initial markup.
- Remove that reference; assign the same URL after a timer.
- Compare initiators and request start times in separate runs.
Explain differences; do not expect identical millisecond values.
DevTools demonstration: execution and rendering
- Record a click on the bounded slow handler.
- Locate its execution in a performance trace.
- Compare the DOM update with the visible result.
- Reset the page before comparing layout experiments.
A breakpoint changes timing. A console log proves execution, not paint.
Evidence before optimization
| Observation | Next question |
|---|---|
| Image request starts late | What revealed its URL? |
| Click response is delayed | What occupied the main thread? |
| Layout repeats | Are writes and geometry reads interleaved? |
| Trace looks different | Did cache, viewport, or throttling change? |
An unexplained trace is a reason to investigate, not to invent a guaranteed sequence.
Practical 01
Produce predictions, comparable observations, and an explanation of one surprise.
The core exercise covers discovery, script timing, scheduling, and rendering work.
Fixed-height virtualization is optional; deeper profiling belongs with Chapter 15.
Check your understanding
- Why can a downloaded script still be waiting?
- What does finding a DOM node tell you about paint?
- Why does a zero-delay timer run later?
- Which evidence would distinguish loading delay from execution cost?
Next: what should the document mean?
Browser behavior explains how a page runs.
Chapter 2 examines how semantic structure, accessible controls, language, and direction describe the interface people use.
2 - Semantic HTML, Accessibility, Internationalization & the DOM
Semantic HTML, Accessibility, Internationalization & the DOM
Meaningful structure, inclusive interaction, and the browser’s live document tree
Chapter 2
Polla Fattah
Today’s goal
Build a mental model in which one document is shared by:
- the browser;
- keyboard users;
- assistive technologies;
- JavaScript;
- search and other software.
The goal is not to memorise tags. It is to choose structure, behaviour, language, and DOM operations deliberately.
By the end of today you can
- choose HTML for meaning and behaviour, not default appearance;
- structure a page with headings and landmarks;
- explain accessible names, focus order, labels, and errors;
- use native controls before inventing ARIA widgets;
- distinguish language from writing direction;
- query, traverse, create, and update DOM nodes safely;
- explain event propagation and delegation;
- describe where Web Components fit without replacing semantic HTML.
The chapter’s single model
flowchart TD
A[Semantic HTML] --> B[Native browser behaviour]
B --> C[Accessibility representation]
C --> D[Language and direction]
D --> E[DOM tree and events]
E --> F[Web Components]These are not separate tricks. They are different views of the same interface model.
A running example: a public-service portal
We will imagine a multilingual service interface with:
- a site header and primary navigation;
- a page explaining one service;
- a request form;
- a table of existing requests;
- a status filter;
- Arabic and English content;
- a small custom element where a reusable boundary is useful.
The interface will remain understandable before JavaScript enhances it.
HTML describes meaning, not appearance
These fragments may look similar after CSS:
But the browser cannot infer the intended roles reliably.
Prefer the meaningful structure
The second version identifies headings, navigation, links, and main content without requiring a stylesheet or a JavaScript guess.
One durable rule
Choose HTML according to meaning and behaviour before styling it according to appearance.
HTML answers:
CSS answers:
JavaScript answers only the behaviour that needs application logic.
Build a document hierarchy
The hierarchy should make sense if all styling disappears.
Headings are content structure
Use heading levels to represent relationships:
Do not choose <h3> because its default font size looks convenient.
There is no automatic outline rescue
Do not assume that nested sections automatically create a perfect heading outline for every tool.
Make the heading hierarchy explicit:
- one clear page-level
<h1>; - meaningful
<h2>regions; <h3>only when a subsection belongs to the preceding<h2>;- no skipped levels just for visual size.
Use CSS for size. Use headings for structure.
Landmarks give the page a map
| Element | Meaning |
|---|---|
<header> | Introductory content for the page or section |
<nav> | A group of navigation links |
<main> | The page’s primary content |
<footer> | Closing information for the page or section |
<aside> | Related content, not the main flow |
Landmarks help users and tools move through a large interface.
<main> and <nav>
There should normally be one primary <main> for the page.
Give multiple navigation regions useful names:
The label distinguishes regions that would otherwise all be announced as navigation.
<header>, <footer>, and <aside>
These elements are contextual:
They are meaningful when their content has the corresponding relationship; they are not decorative wrappers to use everywhere.
<section> and <article>
Use <section> for a thematic region that normally has a heading.
Use <article> for content that could stand on its own or be reused:
If a generic container has no semantic purpose, <div> is the honest choice.
Lists carry relationships
Use list elements when the content is a list:
Do not build a list from paragraphs and line breaks merely because CSS will make it look like one.
Tables carry data relationships
The caption and header relationships make the data understandable outside the visual grid.
Figures, captions, and media
An image’s alternative text describes its purpose in context. Decorative
images may use an empty alt=""; omitting alt is not the same decision.
Native controls are behaviour with a contract
Native elements bring established behaviour:
The browser, keyboard, accessibility APIs, and form system already understand these controls.
Buttons perform actions; links navigate
Do not use a clickable <div> when a button or link expresses the real job.
The clickable <div> problem
This creates a visual imitation, not a complete control:
It lacks, unless you recreate them correctly:
- keyboard activation;
- focusability;
- the button role;
- disabled semantics;
- expected browser behaviour;
- a reliable accessible name.
Start with <button> and style it.
The accessibility tree is not the DOM tree
The browser uses the DOM and other information to expose an accessibility representation.
Assistive technology may receive:
Good structure gives the browser useful information to expose.
Accessible names answer “what is this?”
Controls need names users can perceive through their chosen interface:
Visible labels are usually the strongest starting point. Placeholder text is a hint and should not replace a label.
Name icon-only controls deliberately
The icon is visual decoration. The accessible name communicates the action.
Prefer visible text when space allows; it helps every user, not only screen reader users.
Keyboard interaction is part of the interface
Ask these questions for every interactive feature:
- Can the user reach it with the keyboard?
- Is the focus order logical?
- Is the focus indicator visible?
- Can the user operate it without a pointer?
- Does focus move somewhere sensible after a state change?
Accessibility is interaction design, not a final inspection step.
Focus order follows the interface’s meaning
Avoid using positive tabindex values to force an artificial order.
Prefer:
- meaningful document order;
- native controls;
- a visible focus style;
- a small, deliberate focus-management rule only when a component requires it.
If the DOM order is wrong, CSS reordering does not repair the reading or keyboard model.
Forms are an accessibility system
The form needs a relationship between label, control, name, value, validation, and error feedback.
Native validation is useful information
Native constraints can provide a baseline. If JavaScript adds custom errors, it must preserve a clear message, association, and recovery path.
Do not make invalid state visible only through colour.
Group related controls
fieldset and legend communicate the group relationship that a visual box
alone cannot provide.
Associate errors with the control
The error should explain what is wrong and how to fix it. Do not only add a red border and expect the user to infer the problem.
ARIA adds semantics only when HTML cannot
ARIA can describe a custom interaction, but it does not automatically create the keyboard behaviour, focus management, state changes, or event handling.
flowchart TD
A["First choice: Native HTML"] --> B["Second choice: Native HTML + enhancement"]
B --> C["Last choice: Custom ARIA widget with full contract"]ARIA can make good HTML worse
Do not add roles that contradict the element’s real meaning:
The role does not repair a confused design. Use the element that expresses the job, then add only the missing semantic information.
Language begins in the document
For a passage in another language:
Language metadata supports pronunciation, hyphenation, spell checking, translation, search, and assistive technology decisions.
Language and direction are different
lang identifies the language. dir identifies the writing direction.
Arabic does not mean every surrounding interface region must be permanently right-aligned. Direction belongs to the text and layout context.
Direction can be automatic
dir="auto" lets the browser infer direction from the first strong character.
It is useful for unknown user content, but explicit direction is better when
the application knows the language and layout contract.
Bidirectional text needs isolation
User-generated text can contain characters with different directionality.
Use <bdi> when embedded user text should not disturb the surrounding order.
Use <bdo dir="rtl">...</bdo> only when deliberately overriding direction.
RTL is not “mirror everything”
RTL-aware design considers:
- text direction;
- logical properties such as
margin-inline-start; - icon meaning;
- navigation order;
- tables and numbers;
- mixed-language content;
- focus and reading order.
Do not replace every left value with right and call the interface localized.
The DOM is a live tree
HTML is the source representation. The DOM is the browser’s live object tree.
flowchart TD
Doc[Document] --> HTML[html]
HTML --> Head[head]
HTML --> Body[body]
Body --> Header[header]
Body --> Main[main]JavaScript reads and changes this live tree. It does not edit the original Markdown or HTML source file.
Nodes and elements are related
The DOM contains different node types:
- document nodes;
- element nodes;
- text nodes;
- comment nodes.
An element is a node, but not every node is an element. APIs such as
children, childNodes, parentElement, and textContent expose different
views of that tree.
Query the DOM precisely
Use stable, meaningful selectors. Check whether a query can return null, and
avoid assuming that a selector matched the element you intended.
querySelector() returns the first match. querySelectorAll() returns a
static NodeList of all matches.
Traverse relationships deliberately
Useful relationships include:
parentElement;children;nextElementSibling;closest();matches().
Traversal should express the component boundary, not depend on fragile visual nesting.
Create and insert elements safely
Prefer textContent for untrusted text. Do not use innerHTML as a convenient
string formatter when the value may contain user or remote data.
Attributes and properties are connected, not identical
Attributes are markup-facing values. Properties are live JavaScript-facing state. Boolean controls make the distinction especially visible:
Events connect users to the DOM
Events carry an interaction through a tree. Understanding that path is more reliable than attaching random handlers until the interface appears to work.
Event propagation has phases
flowchart LR
subgraph Capture["1. Capture Phase"]
W1[window] --> D1[document] --> A1[ancestors]
end
subgraph Target["2. Target Phase"]
A1 --> T[target element]
end
subgraph Bubble["3. Bubble Phase"]
T --> A2[ancestors] --> D2[document] --> W2[window]
endThe event can be observed at different points in the path. This is why a parent can respond to an interaction that occurred on a child.
target and currentTarget differ
Confusing these values is a common source of bugs in delegated interfaces.
Default actions and propagation are different
stopPropagation() does not cancel the browser’s default action. Use each
operation only for the problem it solves.
Event delegation scales repeated UI
One listener can handle current and future matching descendants. The handler must still validate the target and preserve keyboard-accessible controls.
Put the public-service interface together
Now add language metadata, labels, native controls, meaningful table headers, and safe DOM updates before adding custom interaction.
Web Components extend HTML
Web Components provide browser-native mechanisms for reusable boundaries:
- Custom elements define a new element name;
- Shadow DOM can isolate internal markup and styles;
- Slots allow controlled content insertion.
They are tools for encapsulation, not a reason to discard document meaning.
A custom element still needs a contract
The element needs clear inputs, output meaning, lifecycle behaviour, and an accessibility story. Encapsulation does not make an interface accessible by itself.
Shadow DOM is not a security boundary
Shadow DOM can hide implementation details and scope styles.
It does not automatically provide:
- security isolation;
- an accessible name;
- keyboard behaviour;
- correct focus management;
- a good component API.
Treat it as an encapsulation mechanism, not a trust boundary.
Slots preserve controlled composition
Slots allow consumers to provide content to named insertion points. The component still owns the surrounding semantics and must define how slotted content participates in the accessible interface.
The component should not erase the document
Good component design preserves:
- meaningful HTML where native elements already work;
- accessible names and labels;
- predictable keyboard behaviour;
- language and direction metadata;
- a clear DOM and event contract.
Web Components are an extension point, not a replacement for semantic HTML.
Misconceptions to leave behind (Part 1)
| Misconception | Better model |
|---|---|
| “If it looks like a heading, it is a heading.” | Meaning comes from the element and structure. |
“A <div> can replace any element.” | Native elements bring behaviour and semantics. |
| “ARIA makes custom controls accessible.” | ARIA describes; code must implement interaction. |
| “Placeholder text is a label.” | Use a real label and associate it with the control. |
Misconceptions to leave behind (Part 2)
| Misconception | Better model |
|---|---|
| “RTL means align everything right.” | Direction, layout, icons, and mixed text need separate decisions. |
| “DOM means HTML.” | HTML is source; the DOM is the live runtime tree. |
“stopPropagation() prevents navigation.” | Default actions and propagation are different. |
| “Shadow DOM is security.” | Shadow DOM is encapsulation, not isolation. |
The practical lab
Build a semantic multilingual service interface.
The runnable code will be in the separate Playground repository. This deck provides the learning sequence and the current repository provides the instructions.
Practical stages 1–3: structure and meaning
- Create the document hierarchy with
header,nav,main, andfooter. - Add meaningful headings, sections, articles, lists, a figure, and a table.
- Build a request form with labels, native controls, a fieldset, and a legend.
Check that the page still communicates its structure without CSS.
Practical stages 4–6: interaction and language
- Test the interface with keyboard-only navigation.
- Add English and Arabic content with correct
langanddirvalues. - Inspect the accessibility representation and repair missing names or relationships.
Do not treat a passing visual inspection as an accessibility test.
Practical stages 7–10: DOM and components
- Query and update the DOM with safe text insertion.
- Add event delegation for repeated request rows.
- Create a small Web Component with an explicit contract.
- Draw the final architecture and identify what remains native HTML.
The exercise is complete when the component boundary is explainable, not merely when the screen looks correct.
Try it yourself
Choose one interface feature and review it in four views:
- What does the raw HTML say?
- What does the keyboard user experience?
- What name and role does assistive technology receive?
- What does the DOM event path do when the user interacts?
Write one improvement that helps at least two of these views at once.
Troubleshooting questions
| Symptom | First question |
|---|---|
| The control cannot be reached | Is it a native interactive element? |
| The label is not announced | Is the label associated with the control? |
| The Arabic text changes layout unexpectedly | Are language and direction declared at the right boundary? |
| The delegated handler does nothing | What are target and currentTarget? |
| Text displays unexpectedly as markup | Did code use innerHTML where textContent was needed? |
| The custom element is confusing | What is its semantic and accessibility contract? |
Change one thing at a time and inspect the actual DOM after each change.
Completion check
- I can choose semantic elements without relying on their default styles.
- I can describe the page using landmarks and a heading hierarchy.
- Every form control has a meaningful accessible name.
- I can use the keyboard through the interface and see focus.
- I can distinguish
langfromdirand handle mixed-direction text. - I can query and update the DOM without unsafe string insertion.
- I can explain event propagation, default actions, and delegation.
- I can describe what Web Components add and what they do not guarantee.
- I completed the Chapter 2 practical in the companion Playground.
Next: Modern CSS Architecture & Layout Systems
Chapter 3: the cascade, layout systems, responsive design, container queries, layers, design tokens, and the difference between visual flexibility and structural chaos.
Modern Front-End Engineering
3 - Modern CSS Architecture & Layout Systems
Modern CSS Architecture & Layout Systems
From cascade decisions to responsive, internationalized component systems
Chapter 3
Polla Fattah
Today’s goal
Stop treating CSS as a collection of visual fixes.
Today we will use CSS as:
- a precedence system;
- a value and token system;
- a layout system;
- a responsive system;
- an internationalization system;
- and an architecture for change.
By the end of today you can
- explain why a CSS declaration wins or loses;
- use cascade layers to control ownership;
- distinguish raw tokens from semantic tokens;
- choose normal flow, Flexbox, Grid, or positioning intentionally;
- use intrinsic sizing,
minmax(),clamp(), andsubgrid; - distinguish viewport responsiveness from container responsiveness;
- build direction-independent layouts with logical properties;
- use modern selectors without creating specificity traps;
- compare plain CSS, CSS Modules, utility CSS, and CSS-in-JS by problem;
- design a responsive multilingual catalogue without JavaScript measurement.
CSS is a system, not decoration
A dashboard must answer architectural questions:
- What happens when the page narrows?
- What happens when a card moves into a narrow sidebar?
- How do cards align when descriptions have different lengths?
- How do long translations affect the layout?
- Which rule wins when a library and the application disagree?
- How can one design decision update the whole interface?
These are system-design questions expressed through CSS.
The chapter’s progression
flowchart TD
A[Cascade & Precedence] --> B[Values & Tokens]
B --> C[Layout Systems]
C --> D[Responsive Design]
D --> E[Internationalized Layout]
E --> F[Modern Selectors]
F --> G[Styling Architecture]The goal is not to memorise properties. It is to make CSS behaviour predictable enough to change safely.
The cascade is the foundation
The cascade decides which declaration supplies the final value.
The answer is not always “the last rule”.
Why declarations compete
Candidate declarations may differ by:
- origin and importance;
- cascade layer;
- specificity;
- scope or proximity where relevant;
- source order.
The useful question is:
Why did this declaration win?
That question is more valuable than memorising a score.
Browser, user, and author styles
The page starts with more than our stylesheet:
- the browser supplies user-agent styles;
- users may supply preferences or styles;
- the author supplies application styles.
An unstyled <h1> is already bold, large, and separated from nearby text.
Author CSS is not the only styling authority. This matters especially for
accessibility preferences and !important declarations.
Inheritance is an architectural tool
Descendant text normally receives these values without repeating them on every element.
Properties such as color and font-family commonly inherit. Properties such
as margin, padding, border, and width normally do not.
Use inheritance for broad design decisions, then override at meaningful boundaries.
Specificity is not a quality ranking
The second selector is harder to override, but that does not make it better.
High specificity increases the cost of future changes.
Prefer a selector that is easy to reason about and easy to replace.
Do not reduce specificity to arithmetic alone
Specificity is one stage in the cascade, not the entire cascade.
Before comparing selectors, ask:
- Are both declarations in the same origin?
- Are they in the same layer?
- Is one declaration important?
- Is a scoped rule involved?
- Only then: which selector is more specific?
The browser is resolving a precedence system, not grading CSS style.
Source order is the final tie-breaker
When the earlier cascade stages tie, the later declaration wins.
Source order is useful for intentional overrides. It is dangerous when a stylesheet becomes a sequence of unexplained patches.
Cascade layers make ownership explicit
The layer order is declared once. A later layer can win without requiring an increasingly specific selector.
Layer precedence comes before specificity
The simple selector in the later components layer can win over the highly
specific selector in the earlier base layer.
This lets architecture control precedence before selector complexity grows.
A practical layer architecture
Possible ownership:
| Layer | Responsibility |
|---|---|
reset | predictable browser baseline |
tokens | custom properties and theme values |
base | elements, typography, document defaults |
components | reusable UI boundaries |
utilities | small intentional helpers |
overrides | explicit application exceptions |
The names matter less than the stable ownership contract.
Custom properties participate in the cascade
Custom properties are live CSS values. They inherit, cascade, and can be overridden by component or theme boundaries.
They are more powerful than text replacement.
Custom properties can represent decisions
The component keeps one layout rule. A meaningful boundary changes the value.
This is different from searching and replacing every 1rem in a project.
Raw tokens and semantic tokens
Raw token:
Semantic token:
Components should usually consume semantic meaning. If the brand palette changes, the component should not need to know the replacement colour.
Token ownership has direction
flowchart TD
A[Raw palette] --> B[Semantic meaning]
B --> C[Component usage]
C --> D[Page composition]Avoid letting a component reach backward into arbitrary raw palette values.
Stable semantic names reduce the blast radius of visual change.
Theming with custom properties
The component stays attached to meaning. The theme changes the values.
Start with normal flow
Normal flow is the layout system already provided by the browser.
Let content determine height and let blocks participate in the document before reaching for positioning.
Positioning changes the relationship
Absolute positioning can be correct for a badge anchored to a card. It is a poor replacement for a page layout system when content can grow or translate.
Use fixed and sticky positioning with equally explicit viewport and scrolling assumptions.
Flexbox is one-dimensional
Use Flexbox when items are primarily arranged along one main axis:
The main axis follows flex-direction. The cross axis is perpendicular to it.
Main axis and cross axis
For a row:
- the main axis is inline/horizontal in the usual LTR case;
- the cross axis is block/vertical.
Use justify-content for main-axis distribution and align-items for
cross-axis alignment. Always consider writing direction before describing an
axis as “left” or “right”.
Flex sizing is not just alignment
Flex sizing considers:
- the flex basis;
- grow and shrink factors;
- available free space;
- minimum sizes;
- intrinsic content size.
min-inline-size: 0 can be necessary when a flexible child must be allowed to
shrink rather than force overflow.
Grid is two-dimensional
Grid describes rows and columns together. It is a strong choice when the relationship between both axes matters.
Grid tracks express constraints
This says:
- create as many tracks as fit;
- do not make a card narrower than its useful minimum;
- share remaining space between tracks.
The layout responds to available space without a device list.
Intrinsic sizing asks the content
Useful intrinsic concepts include:
min-content: the smallest size without avoidable breaking;max-content: the size needed by the unwrapped content;fit-content(): a bounded content-aware size.
Intrinsic sizing is especially important when translations and user content change the amount of text.
Alignment is a separate decision
Do not mix up:
- track sizing;
- item alignment inside a track;
- content distribution across tracks.
The right property depends on which relationship you are trying to control.
Grid or Flexbox?
| Question | Prefer |
|---|---|
| Is the layout mainly one axis? | Flexbox |
| Are rows and columns both meaningful? | Grid |
| Does content define a small toolbar? | Flexbox |
| Does the page define a shared card matrix? | Grid |
| Does one element need to distribute remaining space? | Flexbox |
They are complementary systems, not competing religions.
Nested Grid has a boundary
A nested grid creates a new grid context:
The inner card does not automatically inherit the outer grid’s tracks. This is often correct, but it means related content may fail to align across siblings.
subgrid shares tracks intentionally
subgrid allows descendants to participate in the parent grid’s track sizing.
It is useful when card headings, descriptions, and actions should align across
cards with different content lengths.
Do not use subgrid automatically
Use it when shared track alignment is a real requirement.
Avoid it when:
- cards have independent internal structure;
- the extra coupling makes the component harder to reuse;
- a simpler normal flow is sufficient;
- the alignment is only decorative.
Layout power should follow a demonstrated relationship.
Responsive design is not a device list
Avoid designing only for:
flowchart LR
A[Phone] --> B[Tablet] --> C[Desktop]Real interfaces encounter:
- split-screen windows;
- zoom;
- long translations;
- embedded cards;
- sidebars;
- accessibility text size changes;
- unusual aspect ratios.
Responsive design responds to constraints, not product labels for devices.
Begin with a fluid layout
Let the interface adapt continuously before adding a breakpoint.
Media queries answer environment questions
Media queries respond to the viewport or other environment features:
- viewport width;
- colour scheme;
- pointer precision;
- reduced motion;
- contrast preferences;
- print.
They are not the only responsive tool.
Container queries answer component-space questions
The component responds to the space it actually receives, whether that space comes from the page, a sidebar, or a nested layout.
Media queries and container queries differ
| Question | Tool |
|---|---|
| Is the viewport narrow? | Media query |
| Is this component’s parent narrow? | Container query |
| Does the user prefer reduced motion? | Media query |
| Should a card switch from horizontal to vertical? | Container query |
Container queries do not replace media queries. They solve a different scope of responsive reasoning.
Name containers when the relationship matters
Names make the component’s dependency explicit and avoid accidental reliance on an unrelated ancestor container.
Container query units
Container-relative units include:
cqwandcqh;cqiandcqbfor logical axes;cqminandcqmax.
Use them when the component’s typography or spacing should track its container.
Responsive typography needs limits
clamp(min, preferred, max) prevents a fluid value from becoming unusably
small or excessively large.
Readable text is a layout constraint, not a decorative afterthought.
Modern viewport units
Some mobile viewport units distinguish the small, dynamic, and large viewport:
Choose viewport units according to whether browser UI changes should affect the
layout. Do not assume every mobile viewport is a stable 100vh rectangle.
Responsive images belong to layout
Image dimensions, aspect ratios, loading behaviour, and object fitting affect layout stability and performance together.
CSS has logical axes
Prefer logical properties when the interface can change direction:
This expresses the relationship instead of hard-coding physical left and right values.
Inline and block are relationships
flowchart TD
subgraph Inline["Inline Axis"]
I[Direction text progresses: horizontal LTR/RTL or vertical]
end
subgraph Block["Block Axis"]
B[Direction blocks stack: perpendicular to inline axis]
endIn horizontal English text these often resemble horizontal and vertical. In RTL or vertical writing modes, the physical interpretation changes.
Logical CSS keeps the component contract stable.
Direction-independent components
Test the component with different dir values rather than copying the whole
stylesheet and swapping every physical property.
Some things should not mirror
Direction-aware layout does not mean every visual symbol flips.
Consider separately:
- text and reading order;
- navigation arrows;
- media-play controls;
- brand marks;
- charts and geographic maps;
- numbers and code.
The correct decision follows meaning, not a blanket mirror operation.
Modern selectors express relationships
CSS can now describe relationships more directly:
Use these features to make intent clearer, not to hide an unmaintainable selector strategy.
:has() selects based on a relationship
The parent can respond to the state of a descendant without JavaScript merely to add a class.
Still check whether the selector expresses a stable relationship and whether a class would make the ownership clearer.
CSS nesting needs readable boundaries
Nesting can keep a component’s local rules together. Deep nesting still creates coupling and should not replace a clear component contract.
User preferences are part of responsiveness
Other preferences can affect colour, contrast, and interaction assumptions. Responsive design includes the user’s environment, not only its dimensions.
Motion should communicate something
Use motion to communicate:
- a change of state;
- a relationship between locations;
- progress or feedback;
- continuity during an interaction.
Avoid motion that delays reading, causes discomfort, or exists only because a component has an animation library available.
Styling architecture is a dependency system
Every approach answers questions about:
- where styles live;
- how names are scoped;
- how styles are composed;
- how themes are represented;
- how overrides work;
- how unused styles are removed;
- how runtime and build-time costs are distributed.
The right choice depends on the project’s constraints and ownership boundaries.
Plain CSS
Strengths:
- native browser model;
- no framework requirement;
- easy to inspect in DevTools;
- supports layers, custom properties, and modern selectors.
Risks:
- unclear naming can create collisions;
- global ownership can become ambiguous;
- unused styles need a deliberate strategy.
Architecture is still required even when no tool generates the CSS.
Component-oriented CSS and CSS Modules
Component-oriented CSS groups styles around an interface boundary.
CSS Modules add build-time scoping:
flowchart LR
A[Card.module.css] --> B[Build Step] --> C[Generated Local Scoped Class Names]They reduce accidental collisions, but they do not decide:
- token ownership;
- layout responsibility;
- accessibility;
- responsive strategy;
- whether a component boundary is well designed.
Utility-first CSS
Utility classes make small decisions explicit in markup:
Benefits can include consistency and fast composition. Costs can include noisy markup, difficult component-level meaning, or a false belief that utilities remove architecture.
Utilities are a vocabulary, not a complete design system.
CSS-in-JS is not one technology
CSS-in-JS can mean different things:
- runtime style generation;
- build-time extraction;
- component-local syntax;
- theme-aware value functions;
- atomic style generation.
Compare the actual tool’s runtime cost, debugging model, SSR behaviour, accessibility support, and ownership boundaries - not the category label.
Compare problems, not fashion
Ask:
- How many teams own the styles?
- Do styles need to work without a framework runtime?
- How important is runtime theming?
- What is the build and delivery environment?
- How will developers debug the final CSS?
- How will shared components evolve?
Choose the smallest architecture that meets the real constraints.
The running catalogue interface
We will build a responsive multilingual catalogue with:
- semantic page structure;
- a toolbar;
- summary cards;
- a product grid;
- theme tokens;
- container-aware cards;
- RTL-safe spacing;
- reduced-motion support.
The interface is a test of relationships, not a gallery of CSS tricks.
Establish the cascade and tokens
Start by deciding who owns values and who owns component rules.
Build the page layout
The page breakpoint responds to the viewport. The components inside it should still respond to the space each one receives.
Build the toolbar with Flexbox
Wrapping is part of the design. It is better than forcing every control into a single row that cannot contain translated labels.
Build the product grid with Grid
The grid expresses a useful minimum and lets the available container determine how many columns fit.
Adapt cards to their container
The card does not need to know whether the product grid is on the full page or inside a sidebar.
Use subgrid only for a shared alignment
If product names, descriptions, and actions should line up across a row,
subgrid can share the parent tracks.
If cards do not need shared rows, keep their internal layout independent.
Make the catalogue RTL-safe
Test English, Arabic, and Sorani Kurdish content with long labels. Do not
assume that replacing left with right is localization.
Respect reduced motion
The user preference is a real requirement. It belongs in the component’s architecture and verification, not only in a final accessibility audit.
CSS architecture has dependencies
flowchart TD
Tokens[Tokens] --> Components[Components] --> Composition[Layout Composition]
Tokens -.-> Themes[Themes]
Components -.-> Responsive[Responsive Contexts]
Composition -.-> Pages[Pages]When a component reaches into page-specific selectors, or a page owns a component’s internal spacing, the dependency direction becomes unclear.
Good CSS makes change flow through deliberate boundaries.
Misconceptions to leave behind (Part 1)
| Misconception | Better model |
|---|---|
| Specificity decides every conflict. | Origin, layers, specificity, and order all matter. |
| The most specific selector is best. | The most maintainable selector is usually better. |
| Flexbox replaced Grid. | Flexbox and Grid solve different dimensional problems. |
| Grid replaced Flexbox. | Tool choice follows the relationship being laid out. |
| Responsive means phone/tablet/desktop. | Respond to constraints and user environment. |
Misconceptions to leave behind (Part 2)
| Misconception | Better model |
|---|---|
| Container queries replace media queries. | They answer different scope questions. |
| RTL means swap every left and right. | Use logical properties and inspect meaning. |
| CSS variables are text substitution. | They cascade, inherit, and change at runtime. |
| CSS Modules solve architecture. | Scoping helps; ownership and layout still need design. |
| Utility CSS removes architecture. | A utility vocabulary still needs rules and boundaries. |
The practical lab
Build a responsive multilingual catalogue.
The runnable implementation belongs in the separate Playground repository. The practical instructions define the evidence that the layout works.
Practical stages 1–3
- Establish semantic HTML and a small token layer.
- Add
reset,base,components, andutilitiescascade layers. - Convert raw values into semantic tokens and add light/dark theme values.
At each stage, inspect which layer and token supplied the final style.
Practical stages 4–6
- Build the page layout with normal flow and Grid.
- Build the toolbar with wrapping Flexbox.
- Build the product grid with
minmax()and intrinsic sizing.
Test narrow widths and long translated labels before adding more breakpoints.
Practical stages 7–9
- Demonstrate
subgridwhere card rows must align. - Add container queries for cards and the toolbar.
- Make spacing, borders, and alignment direction-independent.
Use DevTools to test the same component in more than one container.
Practical stages 10–12
- Respect reduced motion.
- Use
:where(),:is(),:has(), or nesting where they make relationships clearer. - Draw the CSS architecture and document ownership boundaries.
The final deliverable is an explainable system, not merely a screenshot.
Try it yourself
Take one product card and test it in four environments:
- wide page content;
- narrow sidebar content;
- English text;
- Arabic or Sorani Kurdish text.
Record one layout decision that remains valid in all four environments and one decision that must adapt.
Troubleshooting questions (Part 1)
| Symptom | First question |
|---|---|
| A rule will not win | Which layer, specificity, origin, and order are involved? |
| A flexible item overflows | Is its minimum size preventing shrinkage? |
| Cards do not align | Is shared track alignment actually required? |
| A component breaks in a sidebar | Is it using a viewport query instead of a container query? |
Troubleshooting questions (Part 2)
| Symptom | First question |
|---|---|
| Arabic text collides with controls | Are logical properties and direction boundaries correct? |
| A theme change needs many edits | Are components consuming semantic tokens? |
| Motion is uncomfortable | Is reduced motion handled at the component boundary? |
Completion check
- I can explain why a declaration wins.
- I can use cascade layers to represent ownership.
- I can distinguish raw tokens from semantic tokens.
- I can choose Flexbox, Grid, or normal flow for a reason.
- I can use intrinsic sizing and
minmax()without a device list. - I can explain when
subgridis useful. - I can distinguish media queries from container queries.
- I can use logical properties for direction-independent layout.
- I can respect reduced motion and long translated content.
- I can compare styling architectures by project constraints.
- I completed the responsive multilingual catalogue practical.
Next: Modern JavaScript & Asynchronous Programming
Chapter 4: lexical scope, closures, asynchronous work, promises, cancellation, and the event loop as an application runtime.
Modern Front-End Engineering
4 - Modern JavaScript & Asynchronous Programming
Modern JavaScript & Asynchronous Programming
Control state, boundaries, timing, cancellation, and data flow
Chapter 4
Polla Fattah
Today’s goal
Move beyond syntax toward the language model needed for modern front-end work.
We will connect:
- scope and closures;
- objects and composition;
- data transformation and immutability;
- modules and dynamic loading;
- promises and asynchronous execution;
- concurrency, races, and cancellation;
- locale-aware output.
By the end of today you can
- explain where a JavaScript binding lives;
- use closures deliberately without confusing them with copied values;
- describe prototype delegation and why classes do not remove it;
- choose data transformation methods for clarity;
- update nested state without accidental mutation;
- design module boundaries and dynamic imports;
- distinguish sequential work from concurrent work;
- handle rejected promises and cancellation explicitly;
- prevent stale results from winning a race;
- format numbers, dates, and plurals for a locale.
The chapter’s progression
flowchart TD
A[Scope & Closures] --> B[Objects & Composition]
B --> C[Data Operations & Immutability]
C --> D[Modules & Iteration]
D --> E[Promises & async/await]
E --> F[Concurrency & Cancellation]
F --> G[Locale-Aware Output]The core idea is:
Modern JavaScript controls state, boundaries, timing, and data flow.
Lexical scope: where a name can be used
The function can look outward to the scope that surrounds it.
An outer scope cannot look inward into a function’s private bindings.
Scope follows the source-code structure. It is lexical, not based on which function happened to call another function at runtime.
Block scope is a lifetime boundary
Practical rule:
- prefer
constby default; - use
letwhen reassignment is required; - understand that both respect block scope;
- treat a block as a useful lifetime boundary.
Closures retain bindings
createCounter() finished, but increment retains access to its lexical
environment.
The closure does not copy count. It retains access to the binding.
Closures appear everywhere in front-end code
The callback remembers input and applyFilter from the surrounding scope.
Common uses include:
- event handlers;
- callback configuration;
- private state factories;
- debounced functions;
- memoized computations;
- request controllers.
Closures can keep data alive
As long as the returned logger is reachable, records remains reachable too.
Closures provide privacy and state, but they can also retain large objects, DOM nodes, or subscriptions longer than intended.
Release listeners, timers, and references when the feature is destroyed.
Objects have a prototype model
An object can delegate property lookup to another object through its prototype.
The object does not need to own every method directly.
Prototype delegation is not copying
Property lookup conceptually follows:
flowchart TD
A[Request object itself] -->|if absent| B[Service prototype]
B -->|if absent| C[Object.prototype]
C -->|if absent| D[undefined]This affects identity, mutation, method lookup, and debugging. Understand the delegation even when your code uses class syntax.
Classes are syntax over prototypes
The method is associated with the class prototype. Class syntax improves readability for some designs; it does not replace JavaScript’s prototype model.
Composition is often simpler than inheritance
Instead of a deep hierarchy:
flowchart LR
A[BaseRecord] --> B[ServiceRecord] --> C[UrgentServiceRecord] --> D[LocalisedUrgentRecord]compose focused capabilities:
Composition can reduce coupling and make a capability easier to reuse or replace.
Destructuring names the data you need
Destructuring can make data flow obvious, but avoid extracting so much that the source of a value becomes difficult to follow.
Rest and spread have different jobs
- rest collects the remaining values;
- spread expands values into a new object or array.
Both are shallow operations.
Spread is not a deep copy
The outer object is new. The nested object is still shared.
Know which references remain shared before calling something an immutable update.
Optional chaining is a guarded lookup
It is useful when a value is legitimately absent.
It does not:
- validate the data shape;
- supply a meaningful fallback;
- prove that a missing value is acceptable;
- repair a broken API contract.
Use it where absence is part of the model, not everywhere as a blanket guard.
Nullish coalescing preserves meaningful zeroes
?? falls back only for null or undefined.
That differs from ||:
Choose the operator according to the domain meaning of empty strings, zero, false, null, and undefined.
Transform collections by intent
Each method communicates a different question. Prefer the method whose name matches the operation.
Use reduce() when it clarifies
reduce() can express aggregation, but it can also hide a complicated
algorithm inside one callback.
Use a loop or named helper when the steps matter more than compactness.
Immutable update patterns
The update creates new references along the changed path and preserves unrelated references.
This helps state systems compare identity and makes change boundaries visible.
Identity and change detection
Identity can communicate which part of a state tree changed.
That does not mean every object must be recreated on every update. Preserve references when the value did not change.
Do not turn immutability into dogma
Mutation can be reasonable when:
- the object is local to one operation;
- no other consumer observes it;
- the mutation improves clarity or performance;
- ownership is explicit.
The architectural question is not “mutation or no mutation?”
It is:
Who can observe this value, and what change signal do they rely on?
ES Modules create boundaries
Modules provide:
- explicit dependencies;
- module scope;
- reusable exports;
- a graph the build tool can inspect.
Modules are architecture, not only file organization.
Named and default exports
Named exports make the public vocabulary explicit:
Default exports can represent one primary value:
Choose a consistent convention. Avoid making consumers guess whether a module has one canonical thing or a set of named capabilities.
Module scope is private by default
cache is not a global just because the module can access it. Consumers see
only what the module exports.
This is a useful boundary for implementation details and state ownership.
The module graph is a dependency graph
flowchart TD
Screen[screen.js] --> ReqClient[request-client.js]
Screen --> FormatStatus[format-status.js]
ReqClient --> HTTP[http.js]The graph affects:
- evaluation order;
- bundling;
- caching;
- code splitting;
- circular-dependency risk;
- ownership and testing boundaries.
Design imports as deliberately as public API calls.
Dynamic imports load on demand
Dynamic imports can reduce initial work and defer rarely used features.
They add asynchronous failure, loading state, and chunk-boundary concerns. A dynamic import is not free merely because it is written in one line.
Iteration is more than for loops
An iterable provides a protocol for producing values over time.
Arrays, strings, Maps, Sets, DOM collections, and custom objects can participate in iteration.
The protocol separates “how values are produced” from “how they are consumed.”
Iterators produce one value at a time
The done flag is part of the contract. Iterators can represent sequences that
are calculated lazily rather than stored as a complete array.
Generators pause and resume
Generators provide a convenient way to define an iterator. They can model progressive work, but they do not automatically make expensive work asynchronous or cancellable.
Async iteration represents arriving values
Async iteration is useful when values arrive over time. The consumer still needs to understand completion, errors, backpressure, and cancellation.
From callbacks to promises
Callbacks can express completion and failure, but nested workflows become hard to compose:
Promises represent a future result and provide composable success and failure paths.
A promise represents a future result
A promise is not the result itself. It represents a pending, fulfilled, or rejected computation.
Creating promises is an ownership decision
Create a promise when adapting a callback or exposing an asynchronous contract.
Do not wrap an API in a new promise without a reason; unnecessary wrappers can hide errors and complicate cancellation.
Promise chaining carries values forward
Each handler returns the value or promise for the next step.
Returning nothing intentionally passes undefined; forgetting to return is a
common source of a broken chain.
Promise errors propagate through the chain
An exception or rejection skips forward to the next compatible rejection handler. Place recovery at the boundary that can actually decide what to do.
async and await improve expression
await makes a promise result available inside the function. It does not turn
the browser into a blocking environment.
Async functions always return promises
Even a plain return value is wrapped in a fulfilled promise. Callers must still decide how to handle rejection.
Handle errors at the right boundary
Do not catch an error merely to log it and then pretend the operation succeeded.
Not every failure is a network error
An async operation can fail because:
- the request was blocked or offline;
- the server returned a non-2xx response;
- the body was invalid JSON;
- validation rejected the data;
- rendering code threw;
- the user cancelled the work;
- a stale result was deliberately ignored.
Your error model should preserve the difference when the UI needs it.
Sequential work is sometimes correct
The second operation depends on the first. Sequential execution expresses that dependency.
The cost is that total time includes both waits.
Accidental sequential work is a performance bug
If the requests are independent, this waits unnecessarily.
Start independent work together, then await the combined result.
Promise.all() expresses all-or-nothing coordination
The operations start without waiting for one another. The combined promise fulfils when all fulfil and rejects when one rejects.
Use it when one failure should prevent the combined result from being used.
Promise.allSettled() preserves every outcome
Use it when partial success is meaningful and each result needs its own status.
The choice communicates product behaviour, not only JavaScript preference.
Concurrency is not parallelism
Several asynchronous operations can be in flight at once while JavaScript continues on its event loop.
That does not mean JavaScript is running those callbacks on several CPU cores.
Concurrency is about overlapping waiting and coordinating completion.
Parallelism is about simultaneous execution.
Race conditions can happen in one thread
sequenceDiagram
participant UI as User Interface
participant Net as Network
Note over UI: Types "ca" (Req A)
UI->>Net: Request A ("ca")
Note over UI: Types "car" (Req B)
UI->>Net: Request B ("car")
Net-->>UI: Response B arrives (150ms)
Note over UI: UI shows "car" results
Net-->>UI: Response A arrives late (800ms)
Note over UI: Race! Stale "ca" overwrites "car"If every response renders immediately, request A can overwrite the newer
result for care.
Single-threaded JavaScript does not remove ordering races between asynchronous operations.
Race conditions are correctness problems
The question is not only:
Which request finished last?
It is:
Which result is still valid for the current user intent?
Freshness is an application rule. The network does not know which response the user currently cares about.
Strategy 1: ignore stale results
This is simple and safe when the old work is harmless. It still spends the resources needed to finish the old request.
Strategy 2: cancel outdated work
Cancellation can save work and communicate that the previous intent is no longer relevant.
The UI must distinguish cancellation from an unexpected failure.
AbortController is a general signal
The signal is a shared cancellation contract. The operation decides how to listen and clean up when the signal aborts.
Cancellation is an application design problem
Ask:
- What work belongs to the cancelled intent?
- Who owns the controller?
- What resources need cleanup?
- Should cancellation be silent or visible?
- Can a new operation replace the old one safely?
- What happens if cancellation occurs after the result arrives?
Adding AbortController without defining ownership only moves the ambiguity.
Debouncing delays the start of work
Debouncing waits for a pause in input before starting work.
It reduces the number of operations. It does not cancel a request that already started.
Debouncing and cancellation solve different problems
flowchart TD
A["Debounce: delay unnecessary starts"] --> B["Cancellation: stop outdated work in flight"]
B --> C["Freshness: reject obsolete results"]Search interfaces often need all three.
Error propagation needs an owner
flowchart TD
T["Transport layer: reports request failure"] --> D["Data layer: validates and normalizes data"]
D --> F["Feature layer: decides recoverable UI state"]
F --> S["Screen layer: presents feedback and retry"]Do not let every layer catch and replace the same error with a vague message.
Each boundary should add context or make a decision.
finally is for cleanup
Use finally for work that should occur after success, failure, or
cancellation: loading flags, controller references, locks, and temporary UI
state.
Avoid unhandled promise rejections
Every promise chain needs an intentional owner:
If a caller intentionally starts background work, make that intent visible and attach a failure path. A rejected promise should not disappear silently.
Internationalization is more than translation
Locale-aware interfaces consider:
- language and script;
- number formatting;
- currencies;
- dates and time zones;
- relative time;
- plural rules;
- segmentation;
- direction and layout.
Do not concatenate English assumptions into strings and call the result localized.
Format numbers with Intl.NumberFormat
The locale controls grouping, decimal conventions, digits, and other display rules. Formatting belongs near presentation, not in the domain value itself.
Currency is a meaning, not a symbol
Do not hard-code a dollar sign or assume that the currency code determines the entire visual format.
Keep the numeric amount and currency identity separate until presentation.
Dates need a locale and a time-zone decision
Formatting is not the same as deciding what instant or calendar meaning the application intends. Store and transmit a clear temporal representation.
Relative time communicates change
The correct unit and rounding policy are application decisions. A relative label should not hide the exact timestamp when the exact time matters.
Plural rules are grammatical rules
Different languages have different plural categories. Avoid building messages
with count === 1 ? "item" : "items" as a universal model.
Intl.Segmenter respects language boundaries
String length and substring boundaries are not always equivalent to visible words or user-perceived characters.
A cancelable search: step 1
Debounce the user’s input so each keystroke does not immediately start work:
This controls the rate of starts. It does not yet control in-flight requests.
A cancelable search: step 2
Cancel the previous request when a newer query becomes authoritative:
The feature owns the controller because it owns the current search intent.
A cancelable search: step 3
Handle cancellation separately from failure:
Cancellation is expected control flow. It should not flash an error message for the user.
A cancelable search: the complete contract
flowchart TD
A[Input event] --> B[Debounce delay]
B --> C[Create current AbortController]
C --> D[Cancel previous active request]
D --> E[Fetch network data with signal]
E --> F[Validate response & token freshness]
F --> G[Format locale-aware values]
G --> H[Render current results only]
H --> I[Cleanup controller in finally]Each arrow is a boundary where an error, race, or ownership decision can occur.
The practical lab
Build a cancelable locale-aware search.
The runnable Vite and TypeScript implementation will live in the separate Playground repository. This deck describes the JavaScript behaviour and the evidence to collect.
Practical stages 1–3
- Use closure state for a small search controller or cache.
- Apply immutable updates to the search state.
- Split the feature into modules with clear imports and exports.
Write down which module owns the current query, results, and cancellation controller.
Practical stages 4–6
- Dynamically import an optional result formatter or panel.
- Compare sequential and concurrent independent requests.
- Create a deliberate race condition by delaying responses differently.
Observe how a correct build can still show incorrect stale data.
Practical stages 7–9
- Prevent stale results from replacing newer results.
- Add
AbortControllercancellation. - Add input debouncing.
Test rapid typing, clearing the input, slow responses, cancellation, and network failure separately.
Practical stages 10–11
- Add locale-aware numbers, dates, plurals, and direction.
- Draw the asynchronous flow and label every ownership boundary.
The final diagram should show what starts work, what can cancel it, what can reject, and what is allowed to update the UI.
Try it yourself
Choose one asynchronous feature in an application you know.
Answer:
- Can two operations be in flight together?
- What makes an old result stale?
- Who owns cancellation?
- Which failures are expected user flow?
- Where should locale formatting happen?
Write one explicit contract before changing the code.
Troubleshooting questions (Part 1)
| Symptom | First question |
|---|---|
| A variable is unavailable | Which lexical scope owns it? |
| A callback uses old state | Which binding did the closure retain? |
A promise chain returns undefined | Did each handler return its next value? |
| Independent requests are slow | Are they accidentally awaited sequentially? |
Troubleshooting questions (Part 2)
| Symptom | First question |
|---|---|
| Old search results appear | What defines freshness and cancellation? |
| Loading never ends | Is cleanup in finally? |
| Cancellation shows as an error | Is AbortError handled separately? |
| Dates look different by machine | Is locale and time zone explicit? |
Completion check
- I can explain lexical scope and closures.
- I understand prototype delegation beneath class syntax.
- I can choose clear collection transformations.
- I can update nested state while preserving useful identity boundaries.
- I can design a module graph with explicit dependencies.
- I can distinguish sequential work from concurrent work.
- I can explain promise rejection and error ownership.
- I can identify and prevent stale asynchronous results.
- I can use cancellation and debouncing for different purposes.
- I can format numbers, dates, plurals, and text for a locale.
- I completed the cancelable locale-aware search practical.
Next: TypeScript, Runtime Contracts & Safe Data Boundaries
Chapter 5: static types, inference, runtime validation, trusted boundaries, and the difference between what a compiler knows and what a browser receives.
Modern Front-End Engineering
5 - TypeScript, Runtime Contracts & Safe Data Boundaries
TypeScript, Runtime Contracts & Safe Data Boundaries
Make the boundary explicit
Chapter 5
Polla Fattah
Today’s goal
Understand what TypeScript can prove, what it cannot know, and how runtime validation turns untrusted input into trusted application data.
We will connect:
- inference and explicit types;
- unions, narrowing, and exhaustive state design;
- generics and reusable contracts;
- strictness, DOM typing, and assertions;
- APIs, URLs, storage, forms, and configuration;
- parsers, schemas, branded values, and trusted domain data.
By the end of today you can
- model domain states instead of decorating arbitrary objects;
- use
unknownat an external boundary; - narrow values with evidence and discriminated unions;
- write generic result and collection APIs;
- keep strict TypeScript useful rather than ceremonial;
- distinguish a type assertion from a runtime conversion;
- validate API, URL, storage, form, and configuration data;
- separate transport types from domain types;
- create branded identifiers only after validation;
- explain why static types do not replace tests or runtime contracts.
The central lesson
flowchart LR
A["TypeScript Static Types\n(Accepted by Compiler)"] -.->|Type Erasure| B["Runtime JavaScript"]
C["Runtime Data\n(Delivered from Outside)"] --> D["Runtime Validation Boundary"]An annotation is useful for the compiler and editor.
It is not a force field around a value arriving from a network, URL, browser, or user.
The chapter’s progression
flowchart TD
A[Type Inference] --> B[Data Modeling]
B --> C[Unions & Narrowing]
C --> D[Generics]
D --> E[Strictness & DOM Typing]
E --> F[Trust Boundaries]
F --> G[Runtime Validation]
G --> H[Trusted Domain Data]The order matters: prove small facts first, then compose them into safe boundaries.
JavaScript has values; TypeScript adds a model
JavaScript executes the values that arrive.
TypeScript can describe the values we expect while developing, but its annotations are removed before the browser runs the code.
A type annotation is not validation
The assertion changes the compiler’s story.
It does not inspect input, check id, or create a missing name.
Start with inference
Inference keeps local code readable and lets the implementation remain the source of truth.
Add an annotation when it documents intent, constrains a boundary, or catches an important mistake.
Widening and literal values
Literal information can widen when a value must remain mutable.
Use literal types deliberately when a small vocabulary is part of the contract.
Model the domain, do not label a guess
This is a useful domain model for trusted application data.
It is not evidence that a random JSON object already satisfies the model.
Interfaces describe object shape
Interfaces work well for object contracts that may be extended or implemented across a codebase.
Type aliases compose precisely
Aliases are especially expressive for unions, tuples, mapped types, conditional types, and domain vocabulary.
Choose the form that communicates the design; both participate in structural typing.
Literal types create a controlled vocabulary
Small finite sets should be visible in the type rather than repeated as undocumented strings.
Unions represent alternatives
The safe question is not “what do I hope this is?”
It is “what evidence distinguishes the possible members?”
Narrow with typeof when the evidence fits
Narrowing is a proof step. After the branch, TypeScript allows operations supported by the proven type.
Narrow with property checks
Property checks are useful, but a stable discriminant is often clearer for important state machines.
Discriminated unions make state explicit
Each state carries exactly the data that makes sense in that state.
This prevents impossible combinations such as status: "loading" with stale error and data fields.
Render state by its discriminant
The branch gives the renderer the exact fields valid for that state.
Intersections combine capabilities
Use intersections when one value genuinely satisfies multiple independent contracts.
Do not use them to hide a domain model that is becoming difficult to understand.
unknown is the honest boundary type
unknown says: a value exists, but this function has not earned knowledge about its shape yet.
It is safer than any because every operation requires evidence.
A type guard is a proof function
The return type tells TypeScript what follows when the function returns true.
The implementation must justify that claim at runtime.
Type guards can still be wrong
This compiles, but it lies: arrays, dates, and empty objects pass the test.
A type predicate is not automatically verified by the compiler. Treat it as safety-critical code.
never protects exhaustive decisions
Adding a new state now creates a compile-time reminder at every incomplete decision.
Generics preserve relationships
Generics are not “any with extra syntax”. They carry a relationship between inputs and outputs.
A generic result keeps success data precise
The common result contract is reusable while T preserves the domain-specific payload.
Constrain a generic when it needs a capability
The constraint does not say every T is exactly an object with only id.
It says the function may safely rely on id while preserving additional fields.
Generic components should describe relationships
The table does not need to know whether a row is a product, lecture, or user.
The caller supplies the domain-specific relationship once.
Utility types transform an existing contract
Utility types are useful for views and transitions, but they should not replace clear domain states.
Strictness catches boundary mistakes early
Prefer a strict baseline:
The exact set depends on the project, but the principle is stable: make uncertainty visible where it matters.
Avoid implicit any
An implicit any turns off the very feedback TypeScript is meant to provide.
If a value is unknown, write unknown and decide how to prove it.
Null and undefined are part of the model
An empty result is not a failed type. It is a state the caller must handle.
Indexed access can be unsafe
Even when an array is typed as Product[], an arbitrary index may not contain a product.
noUncheckedIndexedAccess makes that possibility explicit.
DOM APIs are runtime boundaries too
The selector returns a nullable, broad value. Narrow it before using form-specific behavior.
Type the element and the event
Use the most specific safe DOM type available, and remember that the element can still be absent.
The dangerous API shortcut
This is convenient at the exact point where convenience is most dangerous.
The network response is outside the compiler’s control. Treat it as unknown until it is parsed.
Assertions, casts, and conversions are different
An assertion changes static knowledge.
A conversion or parser changes or inspects a runtime value.
What counts as a trust boundary?
flowchart TD
subgraph Boundaries["Untrusted External Trust Boundaries"]
API[API responses]
URL[URL & query string]
Storage[Browser storage]
Forms[Form input]
Post[postMessage]
SDK[Third-party SDKs]
endEvery source outside the current trusted function may be malformed, stale, partial, or surprising.
URL values are strings, even when they look numeric
Parsing and domain checks are separate decisions. Number("") becoming 0 may be a valid JavaScript conversion but an invalid page.
Browser storage is untrusted text
Storage may contain old versions, manually edited values, or invalid JSON.
Read it as an external payload, then validate a versioned shape.
Forms produce strings and absence
The HTML control type is not enough to guarantee the value your domain logic expects.
Manual validation is a parser
The returned Product is earned by checks and reconstruction, not by a blind assertion.
Parse, then trust the result
Keep parsing close to the boundary. Keep the rest of the application free from repeated shape checks.
Validation errors should help the caller
Useful errors identify:
- which boundary failed;
- which field was unexpected;
- what kind of value was expected;
- whether the response was malformed or the request failed.
Do not include tokens, passwords, or complete sensitive payloads in error messages.
Schema libraries make repetitive checks composable
A schema can centralize parsing, reuse nested rules, report paths, and derive static types.
The library does not remove the need to design the domain contract.
Safe parsing returns a controlled result
This makes invalid input an explicit branch instead of an unexpected exception deep in rendering code.
Transport data is not always domain data
The boundary can validate the transport shape and map it into the vocabulary used by the application.
A complete API boundary
Transport failure and schema failure are separate facts and should remain distinguishable.
Keep raw data unknown at the edge
flowchart TD
A[fetch / storage / URL / form] --> B[unknown]
B --> C[parse + validate]
C --> D[trusted transport data]
D --> E[map to domain]
E --> F[trusted domain model]
F --> G[components and features]The farther a value travels from the edge, the less useful it is to keep asking whether it is valid.
The boundary architecture
flowchart TD
Adapter["Adapter: knows external format"] --> Parser["Parser: checks runtime shape"]
Parser --> Mapper["Mapper: creates domain vocabulary"]
Mapper --> App["Application: consumes trusted values"]
App --> View["View: renders explicit state"]Each layer has a small responsibility and a clear reason to change.
The parse-then-trust principle
Trust should be established once at the boundary, then preserved by types and module boundaries.
Repeated checks in every component usually signal that the boundary is in the wrong place.
satisfies checks without widening useful literals
The object is checked against the contract while retaining precise keys and values for later inference.
This is often better than annotating the whole object as Record<string, string>.
When assertions are legitimate
An assertion can be reasonable when:
- a platform API is typed too broadly;
- a checked invariant is understood by the compiler but not expressible locally;
- a test fixture deliberately models a controlled case;
- the assertion is close to the proof and documented.
It is risky when it is used to skip parsing, null checks, or domain decisions.
Avoid double assertions
This pattern can force unrelated types together and hides the missing proof.
If the boundary is uncertain, keep the value as unknown and write the parser that makes the transition explicit.
Non-null assertions hide a missing state
The ! silences null without changing runtime behavior.
Prefer a checked lookup, a required-element helper, or an initialization path that makes absence impossible.
Structural typing is useful and subtle
Both shapes satisfy the same structure, even if the domain meanings differ.
Shape compatibility does not automatically communicate semantic identity.
Branded types protect semantic identifiers
The compiler can now distinguish two strings that represent different kinds of identifier.
The brand has no runtime representation.
Create a brand only after validation
The assertion is local and justified by the parser’s runtime rule.
Never export a public “brand anything” helper that bypasses the boundary.
Brands are useful when the cost is justified
Consider them for:
- identifiers that are frequently confused;
- validated URLs or paths;
- normalized currency codes;
- security-sensitive tokens with distinct lifecycles.
Do not brand every primitive. Extra vocabulary should reduce real mistakes, not create ceremony.
Static contracts are not shared runtime validation
flowchart TD
A["Shared TypeScript type: agreement during development"]
B["Runtime schema: verification of delivered data"]
C["Generated API types: synchronized description"]
A -.->|requires| B
C -.->|requires| BEven generated types can become stale or be bypassed by a broken server.
The receiving application still owns its trust boundary.
Choose validation proportionately
flowchart TD
A["Small local constant: direct type / simple check"]
B["Stable internal module: typed constructor or parser"]
C["External API: schema and useful errors"]
D["Security-sensitive data: strict validation & logging policy"]The goal is reliable boundaries, not maximum ceremony everywhere.
Typed errors preserve failure meaning
Callers can render, retry, report, or ignore failures based on their actual cause.
TypeScript does not replace tests
Types catch many inconsistencies before execution.
Tests still verify behavior, boundary cases, compatibility, and user-visible outcomes.
Test malformed payloads, missing fields, old storage versions, invalid query strings, and empty results.
TypeScript does not replace security
Compile-time types do not sanitize HTML, authorize a user, protect a secret, or prevent a malicious server response.
Security checks must happen at runtime and at the correct trust boundary.
A malformed response should fail clearly
Do not render the first item and silently accept the second.
Decide whether the contract is all-or-nothing, item-level recovery, or partial data with an explicit warning.
Replace the assertion with unknown
The compiler now forces the design to answer what happens when the payload is not valid.
Validate, then construct the trusted model
In production, make the item failure shape explicit and avoid allowing an exception to erase useful context.
The trusted core should stay boring
Once data is trusted, domain functions should focus on domain behavior.
If every function still checks whether priceCents is a number, the boundary has not done enough work.
Common misconceptions
| Misconception | Better mental model |
|---|---|
as Product validates JSON | It only changes static interpretation |
unknown is inconvenient | It marks work the boundary must do |
any is faster | It removes useful feedback |
| shared types guarantee the API | They describe an agreement, not delivery |
| schemas make domain design unnecessary | They validate a contract you still choose |
| strictness means more code everywhere | It makes important uncertainty visible |
Practical lab: Build a Safe Data Boundary
Create a small API boundary that parses unknown, validates the response, and exposes trusted product data to application code.
The practical moves from static modeling to malformed payloads, manual parsing, schema validation, URL and storage boundaries, and branded IDs.
Practical stages 1–3: model uncertainty
- Model the domain product and its load states.
- Test inference, literal types, and generic result types.
- Introduce
unknownwhere the API response enters.
Do not begin by asserting that the response is already a Product[].
Practical stages 4–6: make state and failure explicit
- Render every discriminated load state exhaustively.
- Add a generic
ApiResult<T>. - Create an intentionally unsafe API response with valid and malformed records.
Observe how the unsafe version spreads assumptions into the UI.
Practical stages 7–9: replace trust with parsing
- Replace the assertion with
unknown. - Write a manual parser with useful field errors.
- Compare the parser with a schema library.
Record the maintenance trade-off: readability, reuse, error paths, dependency cost, and change frequency.
Practical stages 10–13: finish the boundary map
- Validate URL state.
- Validate browser storage.
- Create branded IDs only after validation.
- Draw the trust architecture from raw input to trusted domain data.
Verification: invalid data cannot enter the trusted model silently.
Try this yourself
Add a new archived product state and update every renderer.
Then add a server field that is optional in the transport response but required by the domain model.
Where should the default be applied?
At the parser, mapper, or component? Explain the boundary decision.
Troubleshooting guide
| Symptom | Likely cause |
|---|---|
Everything became any | An untyped boundary leaked inward |
| Assertions appear everywhere | Parsing is missing or too far away |
| Components repeat shape checks | Trusted data is not established once |
| Union branches feel awkward | Add a discriminant or redesign the state |
| A guard compiles but fails in production | The predicate claims more than it checks |
| Old storage breaks after a release | Add versions and migration/validation rules |
Completion checklist
- local values rely on inference where appropriate;
- domain states use explicit, meaningful unions;
- external values enter as
unknown; - parsers return useful failure information;
- transport and domain models are separated where needed;
- branded values are created only after validation;
- tests cover malformed and stale inputs;
- trusted core functions do not repeat boundary checks.
The chapter in one sentence
Use TypeScript to describe and connect trusted program states; use runtime validation to earn trust at every external boundary.
Next: Chapter 6
The next chapter turns these contracts into component architecture:
- component boundaries;
- props and composition;
- controlled and uncontrolled state;
- reusable UI contracts;
- accessibility as a design constraint;
- testing behavior at the boundary.
Questions
Which value in your current application is trusted only because an assertion says it is?
That is the first boundary worth making explicit.
6 - Component-Driven Architecture & Design Patterns
Component-Driven Architecture & Design Patterns
Make responsibility visible
Chapter 6
Polla Fattah
Today’s goal
Learn to decide where a component should begin and end.
We will connect:
- responsibility and reasoning boundaries;
- inputs, outputs, props, events, and slots;
- controlled and uncontrolled state;
- compound and headless components;
- context and dependency injection;
- reuse, abstraction, and duplication;
- feature, layer, and domain-oriented organization;
- accessibility, performance, internationalization, and security.
By the end of today you can
- identify a coherent component responsibility;
- distinguish a component boundary from a state boundary;
- design intent-oriented public APIs;
- choose controlled or uncontrolled ownership deliberately;
- compose behavior without forcing styling decisions;
- use context or injection without hiding dependencies;
- recognize oversized and over-fragmented components;
- extract components in a safe sequence;
- organize a front-end by feature, layer, or domain;
- refactor a catalogue into a maintainable component architecture.
The central principle
A component should exist because it owns a coherent responsibility - not merely because some markup can be extracted into another file.
The file boundary is an implementation detail.
The responsibility boundary is the architecture.
The chapter’s progression
flowchart TD
A[Complex Interface] --> B[Identify Responsibilities]
B --> C[Find Boundaries]
C --> D[Define Inputs & Outputs]
D --> E[Compose Components]
E --> F[Decide State Ownership]
F --> G[Share Dependencies Carefully]
G --> H[Organize by Feature, Domain, or System]Every step reduces accidental coupling while preserving understandable flow.
Why components exist
Consider a catalogue page containing:
- navigation;
- search and filters;
- product cards;
- pagination;
- basket actions;
- loading and error states;
- keyboard behavior and analytics.
A single function can render it.
That does not mean one function is a good unit of reasoning.
Components reduce the amount we understand at once
flowchart TD
Page["CataloguePage"]
Page --> Search["SearchControls"]
Page --> Filter["FilterPanel"]
Page --> Grid["ProductGrid"]
Page --> Paging["Pagination"]
Grid --> Card["ProductCard"]
Card --> Price["Price"]
Card --> Stock["StockStatus"]
Card --> AddBtn["AddToCartButton"]If price formatting changes, we should not need to understand pagination, filters, and cart state at the same time.
Reuse is useful, but not required
It may appear once and still deserve a component because it owns:
- a coherent domain responsibility;
- substantial behavior;
- a meaningful state model;
- a natural testing boundary.
Reuse is one reason to create a component. Responsibility is stronger.
Isolation is encapsulation
A date picker can expose:
while hiding:
- calendar-grid generation;
- month navigation;
- keyboard behavior;
- focus management;
- date-cell implementation.
The caller gets what it needs without controlling every internal detail.
Testing boundaries follow behavior
A date picker can be tested for:
- selected-date behavior;
- keyboard navigation;
- disabled dates;
- accessible labels.
The surrounding page should not have to reproduce every internal interaction to test the date picker.
Test where behavior and risk are concentrated - not merely where files happen to exist.
Ownership makes team work clearer
A useful component boundary can make it clear:
- which team owns the behavior;
- which API is stable;
- which implementation is private;
- which changes require coordination.
Architecture is also communication between people, not only organization for the compiler.
Bad extreme: one giant component
Typical symptoms:
- everything can reach everything;
- a small change creates a large regression surface;
- tests require excessive setup;
- state ownership is ambiguous.
Bad extreme: component explosion
flowchart TD
Page["Page"] --> Section["Section"]
Section --> Wrapper["Wrapper"]
Wrapper --> Row["Row"]
Row --> Text["Text"]
Text --> Label["Label"]Small is not automatically good.
Too many trivial components increase navigation cost, obscure control flow, and make a simple change require a tour of the repository.
Find boundaries by asking five questions
- What changes together?
- What owns the behavior?
- What represents one domain concept?
- What is genuinely reusable?
- What should remain private?
The answers should describe responsibility, not just the shape of the markup.
What changes together?
If a visual structure, its interaction rules, and its tests always change together, they may belong in one component.
If a value, layout, and data-fetching policy change independently, forcing them into one unit creates coupling.
Change patterns are evidence for boundaries.
What owns the behavior?
Do not move state merely because a child renders the element.
The owner is the unit that makes the decision and coordinates its consequences.
What represents one domain concept?
Domain concepts are often stronger boundaries than generic visual fragments.
They give the API meaningful vocabulary and give tests a clear subject.
What should remain private?
Not every internal helper needs to become a public component.
Keep implementation details local when they:
- have no independent responsibility;
- are used in one place;
- would expose unstable structure;
- make the public API harder to understand.
Local components are valid architecture.
A deliberately monolithic shape
The problem is not that it renders JSX.
The problem is that unrelated responsibilities have no visible boundaries.
First refactoring: stable visual responsibilities
Extract the stable responsibilities first.
Keep coordination in the parent until the data and ownership decisions are understood.
A component boundary is not necessarily a state boundary
A child can be a useful rendering boundary while state remains higher in the tree.
Do not force every component to own every value it displays.
Inputs and outputs define the contract
flowchart LR
Inputs["Inputs / Props\n(Outside Values)"] --> Comp["Component\n(Encapsulated Logic)"]
Comp --> UI["Rendered Output\n(Visible Result)"]
Comp --> Outputs["Outputs / Callbacks\n(Events & Intent)"]The public contract should be smaller and more stable than the implementation.
Props describe data and capabilities
Good props answer:
- what does this component need?
- what decisions can it request?
- what state is intentionally owned elsewhere?
Vue props and events express the same boundary
The syntax differs from React.
The architectural question is the same: what enters, what leaves, and who owns the decision?
Callbacks and events should communicate intent
Prefer:
Over exposing implementation details:
Intent-oriented APIs let the component change its internal markup without breaking callers.
Children and slots enable composition
The panel owns structure and semantics.
The caller supplies content that varies.
Named slots make variation explicit
Named composition points communicate where variation belongs without adding a boolean prop for every possibility.
Composition over inheritance
flowchart TD
Card["Card Container"]
Card --> Header["Header (Slot / Child)"]
Card --> Body["Body (Slot / Child)"]
Card --> Actions["Actions (Slot / Child)"]Compose small capabilities and content rather than building a deep inheritance hierarchy.
Composition keeps variation near the caller and avoids base classes accumulating unrelated assumptions.
Wrapper components should add a reason
A wrapper is valuable when it adds:
- semantics;
- layout responsibility;
- accessibility behavior;
- a state boundary;
- a stable composition API.
If it only forwards every prop and renders one child unchanged, it may be navigation cost without architectural value.
Controlled components: the parent owns the value
The component renders and reports interaction.
The parent owns the source of truth and can synchronize it with URL state, data fetching, or another control.
Uncontrolled components: the component owns the value
The component manages its internal editing state.
This can simplify local interactions when the parent does not need every intermediate value.
Controlled versus uncontrolled is an ownership decision
flowchart TD
A{"State Ownership Need"}
A -- Synchronize with external state / URL --> B["Controlled Component\n(Parent owns state via props & callbacks)"]
A -- Only final submission value required --> C["Uncontrolled Component\n(Component manages internal state)"]
A -- Both local reactivity & parent control --> D["Deliberate Bridge Pattern\n(Explicit value + onChange synchronization)"]Neither mode is universally better.
Choose based on who must make decisions about the state.
Avoid half-controlled APIs
What wins if value is absent? What happens when it appears later?
Ambiguous ownership creates warnings, stale state, and surprising transitions.
Use separate APIs or document the controlled/uncontrolled contract precisely.
Good component APIs minimize invalid combinations
Instead of:
Model meaningful states and relationships.
Keep the public surface small and stable
flowchart TD
subgraph PublicContract["Public Interface (Contract)"]
Props["Props / Attributes"]
Events["Events / Callbacks"]
Slots["Slots / Children"]
end
subgraph PrivateImpl["Private Implementation (Encapsulated)"]
State["Internal State"]
Helpers["Private Helpers & Handlers"]
DOM["Internal DOM Nodes"]
end
PublicContract --> PrivateImplEvery public prop is a promise.
Every public event is a dependency for callers.
Expose the smallest contract that supports the component’s responsibility.
Compound components share a local vocabulary
The pieces are separate in markup but belong to one conceptual system.
Why compound components can help
They can provide:
- readable structure;
- explicit composition points;
- shared state without prop repetition;
- a constrained vocabulary;
- flexible placement of tabs and panels.
The API communicates the relationship between the parts.
The cost of compound components
They also add:
- hidden coordination rules;
- context or injection dependencies;
- more concepts to document;
- more invalid combinations to prevent;
- debugging work when pieces are used incorrectly.
Use the pattern when the relationship is real and repeated - not because the API looks advanced.
Headless components separate behavior from styling
flowchart TD
subgraph HeadlessLogic["Headless Behavior Layer"]
KB["Keyboard Navigation Rules"]
Sel["Selection State Machine"]
ARIA["ARIA Attributes & Relationships"]
Focus["Focus Management"]
end
subgraph ConsumerUI["Consumer UI Layer"]
Markup["Consumer-Owned Markup"]
Styles["Tailwind / Custom CSS Styles"]
Layout["Flexible Component Layout"]
end
HeadlessLogic --> ConsumerUIA headless component supplies interaction logic without imposing a visual design.
Why headless UI exists
It is useful when teams need:
- shared accessibility behavior;
- different visual systems;
- consistent keyboard interaction;
- application-specific layout;
- reusable state machines.
The abstraction is behavioral, not merely visual.
Dependency sharing: explicit props first
Explicit inputs make dependencies visible and make the component easy to render in isolation.
Prop passing becomes a problem only when it is genuinely burdensome or crosses unrelated layers repeatedly.
React context shares a local dependency
Context can keep compound parts coordinated without repeating the same props at every level.
Vue provide/inject expresses the same idea
The framework syntax changes; the architectural trade-off remains: the dependency is less visible at the call site.
Context is not automatically better than props
Context can:
- hide where a value comes from;
- make isolated rendering harder;
- widen a component’s implicit dependency surface;
- cause broad updates when the value changes.
Use it for a real shared relationship, not just to avoid writing one more prop.
Dependency injection makes variation explicit at setup
The feature depends on a capability, not on one concrete network implementation.
Why dependency injection helps
It can improve:
- testing with fakes;
- environment-specific adapters;
- separation of domain behavior from transport;
- migration between implementations.
The dependency remains a design decision rather than a hidden import.
Dependency injection can also be overused
If every helper receives a container with dozens of services, the dependency graph becomes harder to understand.
Inject the smallest capability needed.
Prefer a direct import for a stable, genuinely global constant when injection adds no meaningful variation.
Reusable UI components versus application components
Reusable UI components should avoid application-specific assumptions.
Application components should speak the domain language and may coordinate several reusable primitives.
Reuse has levels
flowchart TD
A["One Feature Boundary\n(Lowest cost, highest velocity)"] --> B["One Application Boundary\n(Shared across feature teams)"]
B --> C["Multiple Products Boundary\n(Multi-app shared packages)"]
C --> D["Global Design System Package\n(Highest contract cost, strict versioning)"]The wider the reuse boundary, the more expensive the public contract becomes.
Do not design a global package API for a problem that only exists in one feature.
Premature abstraction freezes assumptions
An abstraction created before the second use often encodes:
- accidental naming;
- the first layout’s constraints;
- one feature’s state model;
- props that do not generalize.
Wait until the shared behavior and variation are understood.
Duplication can be cheaper than the wrong abstraction
Two similar components may differ in:
- ownership;
- accessibility requirements;
- lifecycle;
- domain vocabulary;
- future change direction.
Temporary duplication preserves independent evolution.
Remove duplication when the shared concept - not only the current markup - is real.
Design methodologies are lenses, not laws
Atomic design, feature folders, layers, and domain modules can all be useful.
None can decide a boundary without understanding:
- change patterns;
- ownership;
- dependencies;
- product vocabulary;
- team constraints.
Use a methodology to ask better questions, not to avoid judgment.
Atomic design: strength and limitation
Atomic design encourages a vocabulary from primitives to composed interfaces.
It can help teams discover reusable visual patterns.
But visual size does not always match responsibility.
An “organism” may be a domain feature, while a tiny “atom” may still contain complex behavior.
Feature-oriented decomposition
Feature organization keeps the code that changes together close together.
It is often a strong default for application-scale work.
Layer-oriented structure
Layer organization makes technical roles easy to scan.
Its risk is scattering one feature across many directories and weakening domain ownership.
Domain-oriented decomposition
Domain organization keeps vocabulary, behavior, and tests near the concept they serve.
Technical reuse still crosses domains
A reusable technical primitive can be shared across domains without forcing domain components into one generic model.
The shared layer should remain intentionally small.
Domain components should speak domain language
Prefer:
Over:
The domain component hides coordination details and exposes the intent relevant to its caller.
Avoid boolean prop explosion
Many booleans create a combinatorial API and states nobody designed.
Prefer meaningful variants or modeled states:
Stable components should not depend on the whole application
A reusable component should not know:
- the entire route tree;
- the global store shape;
- the current user object;
- every feature’s analytics policy.
Pass the smallest data and capabilities needed.
Application coordination belongs above the reusable boundary.
Smart and presentational is a useful distinction
flowchart TD
Container["CatalogueContainer\n(Fetches data, owns state, coordinates features)"]
Grid["ProductGrid\n(Lays out collection, handles viewport)"]
Card["ProductCard\n(Presents one item, emits user intent)"]
Container --> Grid
Grid --> CardThe names are less important than the separation of coordination from presentation.
React catalogue architecture
The page coordinates. The children expose focused responsibilities.
React product grid
The grid owns collection layout.
The card owns one product’s presentation and interaction surface.
Vue can express the same architecture
Compare the responsibilities and data flow, not the framework punctuation.
Compare architecture, not syntax
flowchart LR
subgraph React["React Paradigm"]
RProps["Props + Callbacks"]
RChild["Children"]
RCtx["React Context"]
RHooks["Custom Hooks"]
end
subgraph Vue["Vue Paradigm"]
VProps["Props + Emits"]
VSlots["Slots"]
VPI["provide / inject"]
VComp["Composables"]
end
RProps <-->|Architectural Equivalent| VProps
RChild <-->|Architectural Equivalent| VSlots
RCtx <-->|Architectural Equivalent| VPI
RHooks <-->|Architectural Equivalent| VCompFrameworks provide mechanisms.
They do not decide ownership, coupling, or the right domain boundary for you.
Recognize an oversized component
Warning signs:
- many unrelated state variables;
- long conditional render branches;
- repeated markup with slightly different behavior;
- effects that coordinate unrelated systems;
- tests that require the entire application setup;
- a name that no longer describes one responsibility.
Extract by responsibility, not by line count alone.
Recognize over-fragmentation
Warning signs:
- a component has no meaningful API;
- understanding one behavior requires many file jumps;
- props are forwarded unchanged through several layers;
- every markup element has its own file;
- local details become global vocabulary.
Navigation cost is part of the architecture.
Use the extraction test
Ask:
- Does it have its own responsibility?
- Does it have a meaningful public API?
- Does it change for a different reason?
- Is it reused or likely to be reused for a real reason?
- Will extraction reduce or increase navigation cost?
If most answers are no, keep it local for now.
Colocation keeps private behavior near its owner
Colocation shortens the path from behavior to styles to tests.
Move files outward only when they become shared or when the repository’s structure makes ownership clearer elsewhere.
Component boundaries and performance
Boundaries can help with:
- reducing rerender scope;
- memoizing stable subtrees;
- lazy-loading feature code;
- isolating expensive calculations.
But splitting every element is not a performance strategy.
Measure the actual bottleneck and preserve a clear data flow.
Component boundaries and accessibility
An accessible pattern has a behavioral contract:
- roles and relationships;
- keyboard interaction;
- focus movement;
- names and descriptions;
- disabled and busy states.
Keep these rules with the component that owns the interaction, especially for dialogs, tabs, menus, and composite widgets.
Component boundaries and internationalization
Do not bury locale assumptions in generic components.
The component can format correctly while the feature decides which domain value and locale it is displaying.
Text, plural rules, direction, and date conventions are part of the public behavior.
Component boundaries and security
Keep trust-sensitive behavior near the boundary that understands it.
flowchart LR
External["External / User Content"] --> Sanitize["Sanitize & Validate Boundary"]
Sanitize --> Trusted["Trusted Display Component"]Do not make a generic renderer responsible for deciding whether arbitrary HTML is safe.
Make the safe path the easiest public API.
The public API is a contract
Changing a public prop can affect:
- callers;
- tests;
- stories and examples;
- accessibility behavior;
- analytics;
- downstream packages.
Treat component APIs with the same care as a service boundary: explicit, intentional, and proportionate.
A component is not automatically a design-system component
flowchart LR
App["Application Component"] --- AppDesc["Speaks one product's domain language"]
DS["Design-System Component"] --- DSDesc["Cross-product stable behavior & appearance"]Promoting a local component too early creates a public API before variation is understood.
Keep application components local until real reuse proves the broader boundary.
One possible layering model
flowchart TD
Shell["Application Shell"] --> Feature["Feature Components"]
Feature --> Domain["Domain Components"]
Domain --> Adapter["Shared Behavior & Data Adapters"]
Adapter --> UI["Reusable UI Primitives"]The dependency direction should be intentional.
Lower-level primitives should not import feature-specific decisions.
Practical refactoring strategy
- Identify responsibilities.
- Identify shared state.
- Extract stable visual responsibilities.
- Keep coordination in the parent initially.
- Refine APIs.
- Move truly reusable primitives downward.
- Add context or injection only when explicit passing is genuinely burdensome.
Refactoring is a sequence of smaller decisions, not a single rewrite.
Refactor the monolith in observable steps
flowchart TD
Monolith["Monolithic Interface"] --> Layout["1. Stable Layout Shell"]
Layout --> Collection["2. Product Collection Boundary"]
Collection --> Item["3. Product Item Boundary (ProductCard)"]
Item --> Search["4. Controlled Search & Filter Boundary"]
Search --> Composition["5. Composition Slot / Children Boundary"]
Composition --> Shared["6. Shared Primitives (Only where justified)"]After each step, preserve behavior and re-check ownership.
Practical lab: Compound Headless Tabs
Build a composable tabs system that separates keyboard behavior and state ownership from styling and content layout.
The practical applies the chapter’s component API, controlled state, compound composition, accessibility, and headless behavior principles.
Practical stages 1–3: name responsibilities
- Define the responsibility of the root, tab list, tab, and panel.
- Implement controlled and uncontrolled selection.
- Add
role,aria-selected,aria-controls, andtabindexbehavior.
Start with the contract before choosing context, slots, or styling.
Practical stages 4–6: test the interaction
- Test arrow-key navigation, disabled tabs, and dynamic panels.
- Compare explicit props, context/provide-inject, and a headless API.
- Verify that selected tab and visible panel remain synchronized.
The keyboard behavior should be independently testable from the visual treatment.
Practical extension: lazy panels
Add lazy panel loading.
Document the states explicitly:
flowchart LR
Unrequested["Not Requested"] --> Loading["Loading"]
Loading --> Ready["Ready"]
Loading --> ErrorState["Error"]
ErrorState -->|User Retry| LoadingThe tabs contract should make loading and failure visible rather than hiding them inside a boolean prop combination.
Try this yourself
Add a disabled tab that:
- cannot receive selection;
- remains represented in the tab order according to your chosen accessibility policy;
- exposes its disabled state;
- does not load its panel;
- does not break arrow navigation.
Write the behavioral contract before the implementation.
Troubleshooting guide (Part 1)
| Symptom | Likely cause |
|---|---|
| Props are forwarded through many layers | Ownership or a real shared dependency is unclear |
| Every component has many booleans | Invalid combinations are not modeled |
| A child and parent fight over state | The controlled contract is ambiguous |
| Context appears everywhere | Dependencies are hidden instead of designed |
Troubleshooting guide (Part 2)
| Symptom | Likely cause |
|---|---|
| Reusable component needs feature knowledge | The boundary is too low or too broad |
| Refactor creates dozens of files | Extraction followed markup, not responsibility |
| Keyboard behavior is duplicated | Interaction logic lacks one owner |
Completion checklist
- each extracted component has a coherent responsibility;
- public props and events express intent;
- state ownership is explicit;
- controlled and uncontrolled modes are not mixed accidentally;
- compound APIs have a real relationship to model;
- headless behavior is independent from styling;
- context or injection is used only where it clarifies composition;
- accessibility behavior belongs with its interaction owner;
- the final structure reflects feature or domain change patterns.
Misconceptions to leave behind (Part 1)
| Misconception | Better mental model |
|---|---|
| Every repeated markup needs a component | Responsibility and change patterns matter |
| Components exist mainly for reuse | Reasoning, isolation, and ownership matter too |
| Smaller components are always better | Navigation cost is part of quality |
| Each component owns all its state | The owner is the decision-making unit |
Misconceptions to leave behind (Part 2)
| Misconception | Better mental model |
|---|---|
| More props mean more flexibility | More public combinations mean more obligations |
| Context is better than prop drilling | Context trades repetition for visibility |
| Reusable means generic | Reuse should preserve meaningful vocabulary |
| A framework decides architecture | Frameworks provide mechanisms, not boundaries |
The chapter in one sentence
Design components around coherent responsibility, explicit ownership, and small stable contracts; let composition provide variation.
Next: Chapter 7
The next chapter will build on component architecture with:
- application state and data flow;
- URL and server state;
- local versus shared state;
- synchronization and derived state;
- predictable updates and debugging.
Questions
Which component in your current application owns too many decisions?
Which one has an API so generic that its domain responsibility is no longer visible?
7 - Reactivity & Rendering Mechanics
Reactivity & Rendering Mechanics
Trace state to the screen
Chapter 7
Polla Fattah
Today’s goal
Understand what happens between a state change and the UI the user sees.
We will connect:
- state-driven rendering;
- snapshots, batching, and identity;
- reconciliation and keys;
- derived values and memoization;
- effects and cleanup;
- React and Vue mental models;
- fine-grained reactivity and signals;
- scheduling, measurement, and unnecessary work.
By the end of today you can
- describe UI as a function of state;
- distinguish render calculation from DOM commitment;
- explain reconciliation and component identity;
- use keys to preserve or reset state intentionally;
- update from previous state safely;
- separate source state from derived values;
- reserve effects and watchers for external synchronization;
- explain React snapshots, Vue proxies, computed values, and batching;
- trace a reactive dependency graph;
- optimize only after measuring the actual work.
The central lesson
flowchart TD
A["Source State"] -->|dependencies| B["Render / Derivation"]
B -->|scheduling| C["Commit / Synchronization"]
C --> D["Screen & External Systems"]Reactivity is not magic.
It is a system that tracks relationships, schedules work, preserves identity, and crosses explicit side-effect boundaries.
The chapter’s progression
flowchart TD
A[Manual DOM Updates] --> B[State-Driven UI]
B --> C[Render & Reconciliation]
C --> D[Identity & Snapshots]
D --> E[Derived State & Effects]
E --> F[React Mechanics]
F --> G[Vue Mechanics]
G --> H[Fine-Grained Systems & Signals]
H --> I[Scheduling & Measurement]The same UI goal can be implemented with different reactive mechanisms.
Manual DOM updates are imperative
The event handler must remember every DOM node affected by the change.
As the interface grows, manual synchronization becomes a coordination problem.
State-driven UI describes a result
The UI is derived from state rather than updated through a growing list of unrelated DOM instructions.
UI as a function of state
The function is conceptual, not necessarily a literal full-page rewrite.
It gives the system a useful question:
Given this state, what should the interface represent?
Reactivity does not mean everything runs automatically
A reactive system must decide:
- what depends on what;
- what work is invalidated;
- when work is scheduled;
- what can be reused;
- what must be synchronized;
- when cleanup runs.
“Reactive” describes a mechanism, not a guarantee that every operation is free or immediate.
State is not the same as every variable
Use state when a change must participate in the UI’s update model.
Do not put every value in state just because it changes somewhere.
A small running example
The visible list depends on all three values.
The selected count may be derived.
The network request is an external operation.
The reactive architecture should make each relationship visible.
The React mental model
React calls component functions to calculate a description of UI.
Calling the function does not mean the browser DOM is immediately rewritten.
A React render is a calculation
flowchart TD
Update["State Update Triggered"] --> Call["React calls component function"]
Call --> Desc["New Element Descriptions (VNodes)"]
Desc --> Reconcile["Reconciliation / Diffing Engine"]
Reconcile --> Commit["Commit Necessary Host DOM Changes"]The render phase should be free of observable side effects.
Render should behave like a pure calculation
Avoid in render:
- network requests;
- subscriptions;
- DOM mutation;
- timers;
- analytics calls.
Reconciliation compares descriptions
flowchart LR
subgraph Prev["Previous VNode Tree"]
Old["<h1>Old</h1>"]
end
subgraph Next["Next VNode Tree"]
New["<h1>New</h1>"]
end
Prev -->|Diff / Reconciliation| Patch["DOM Mutation: textContent = 'New'"]
Next -.-> PatchThe framework compares what was previously described with what is now described, then determines the smallest host update required.
The component function may run even when the DOM change is tiny or absent.
Virtual DOM without the mythology
“Virtual DOM” is a representation and comparison strategy.
It does not mean:
- the entire real DOM is rebuilt every time;
- virtual work is automatically faster than all alternatives;
- every render creates a visible browser update;
- architecture no longer matters.
Measure the work that actually matters.
The commit phase changes external reality
After reconciliation, the framework commits required changes to the host environment.
flowchart LR
RenderPhase["Render Phase: Pure Calculation\n(Calculate VNodes; zero DOM side-effects)"] --> CommitPhase["Commit Phase: Host Mutation\n(Apply DOM diffs, attach refs, run effects)"]Keeping these phases conceptually separate explains why render code should remain predictable.
Parent rendering and child rendering
When a parent renders, its child functions may be called again.
That does not automatically mean:
- every child DOM node changed;
- every child state reset;
- every expensive calculation must rerun forever.
Rendering, reconciliation, commitment, and state preservation are related but distinct questions.
Component identity is part of behavior
flowchart TD
Check{"Component Check"}
Check -- Same component type + same key + same position --> Preserve["Preserve state & instance"]
Check -- Different component type OR different key --> Destroy["Destroy old instance & initialize new state"]Identity determines whether a component is treated as the same logical instance.
This affects focus, input values, animations, and user experience - not only performance.
State is associated with a position in the render tree
If the same component remains at the same logical position, its state may persist across parent renders.
Changing the structure or key can intentionally create a new identity.
Keys are about identity, not silence
The key tells the renderer which item is which across changes.
It is not merely a way to remove a warning.
Why array index keys can be dangerous
If items are inserted, removed, or reordered, the index can now identify a different item.
Local state may move to the wrong row.
Use a stable identity from the data whenever the list can change.
Keys can intentionally reset state
When documentId changes, the editor receives a new identity.
This can be useful when switching between records should discard local draft state.
Resetting is a deliberate product behavior, not a rendering trick.
State is a snapshot
The handler closes over the state value from the render that created it.
Calling a setter schedules a future render; it does not mutate the current snapshot in place.
Batching combines updates
Both updates may read the same snapshot and request the same next value.
Batching reduces unnecessary intermediate work, but it means update intent must be expressed correctly.
Use the previous state for dependent updates
Each updater receives the latest queued value.
Use this form whenever the next state depends on the previous state.
Derived values should usually be calculated
If a value can be directly calculated from current inputs, storing it separately creates another synchronization obligation.
Source state versus derived state
Store the source values.
Calculate the derived values during rendering or in a memoized calculation when measurement justifies it.
Duplicated state creates inconsistency
Now every product update and every filter update must keep both arrays synchronized.
One source of truth is usually simpler and safer.
Expensive derived values are a measurement question
First ask:
- is the calculation actually expensive?
- how often does it run?
- how large is the input?
- which dependencies change?
Do not add memoization because a calculation exists.
Memoization preserves a calculation result
The cache is valid only while its dependencies represent the same inputs.
Memoization is an optimization with cost, not a correctness requirement.
Referential identity affects memoization
This object is new on each render.
Passing it to a memoized child or using it as a dependency may invalidate the optimization even when its contents look unchanged.
Understand identity before optimizing around it.
Effects synchronize with something outside rendering
The document title is external to React’s render calculation.
The effect synchronizes it after the committed UI reflects the new state.
Effects are not for ordinary derivation
Avoid:
This creates an extra state update and an intermediate render for a value that can be calculated directly.
Use an expression or memoized calculation instead.
Events and effects answer different questions
Submit an order because the user activated submit.
Update a subscription because committed state says the subscription should exist.
Do not turn every event response into an effect chain.
Effect cleanup prevents stale work
Cleanup runs when dependencies change or the component leaves the tree.
It prevents old subscriptions, timers, and requests from outliving the state that created them.
React rendering summary
flowchart TD
A["State Update"] --> B["Snapshot-based Render Calculation"]
B --> C["Reconciliation & Identity Matching"]
C --> D["Commit Host DOM Changes"]
D --> E["Effects Synchronize External Systems"]This is a model for reasoning, not a promise that every implementation detail is synchronous or simple.
The Vue mental model
Vue tracks reactive dependencies more directly through refs, reactive proxies, computed values, and watchers.
The goal remains the same:
The tracking mechanism and timing model differ from React’s component recalculation model.
Vue ref() wraps a reactive value
The ref object gives Vue a stable reactive container.
In templates, Vue can unwrap refs for convenient access.
Vue reactive() proxies an object
The proxy intercepts reads and writes so Vue can track which reactive effects depend on which properties.
The proxy and original object differ
Use the reactive proxy consistently.
Identity assumptions become important when comparing, storing, or passing reactive objects.
Destructuring can break a reactive connection
The local query is no longer a reactive property reference in the same way.
Use toRefs, a computed value, or access through the reactive object when the connection must be preserved.
Vue rendering is a reactive effect
When a component renders, Vue records which reactive values it reads.
When one of those values changes, Vue knows that the component’s rendered output may need updating.
This is dependency tracking at the property level rather than a blanket statement that every component always reruns.
Vue DOM updates are scheduled
The state assignments are observable to Vue immediately, but DOM work is typically queued and batched.
Do not assume the DOM reflects a mutation on the next line.
Vue batching reduces intermediate work
Several synchronous mutations can be grouped into one update cycle.
This is why code that needs the updated DOM may need nextTick() or an equivalent lifecycle boundary.
The scheduling model should be part of the component’s reasoning, especially around focus and measurement.
Computed values represent derivation
Computed values declare a dependency relationship and can cache until their dependencies invalidate.
Computed values are not ordinary state
flowchart TD
P["products (source)"] --> C["computed: visibleProducts"]
Q["query (source)"] --> C
F["filter (source)"] --> CThe computed result should not be manually synchronized with every source update.
Keep derivation represented as derivation.
Vue watchers synchronize external systems
This is appropriate when a change in reactive state must update a router, storage layer, network operation, or other external system.
Watchers should not replace computed values
This stores a derived value and introduces a synchronization path.
Prefer computed when the result is a direct calculation.
watch() and watchEffect() differ
flowchart TD
W1["watch(source, callback)"] --- W1D["Explicit dependency source & controlled comparison"]
W2["watchEffect(callback)"] --- W2D["Automatic dependency discovery during execution"]Use the most explicit form that communicates the intended relationship.
Watcher cleanup prevents stale effects
If a newer query arrives, the old operation should not win after it becomes irrelevant.
React and Vue share the goal, not the mechanism
| Concern | React | Vue |
|---|---|---|
| primary model | component calculation | tracked reactive dependencies |
| derivation | expression / memo | computed |
| external sync | effect | watch / watchEffect |
| state update | setter schedules render | mutation invalidates dependencies |
| DOM timing | commit and effects | queued update and nextTick |
Learn the mechanism well enough to predict behavior; do not flatten the differences into slogans.
Reactivity has granularity
flowchart TD
Coarse["Coarse Reactivity\n(Rerun broad component calculation)"]
Fine["Fine-Grained Reactivity\n(Invalidate only computations reading changed property)"]Coarser systems can be simple to reason about.
Finer systems can reduce work by tracking smaller dependencies.
Neither choice removes the need for good state design.
Fine-grained reactivity
flowchart LR
SigA["Signal A"] --> Deriv["Derived C"]
SigB["Signal B"] --> Deriv
Deriv --> Eff["Effect (DOM / Output)"]Only computations that depend on invalidated sources need to be reconsidered.
This graph is explicit in the runtime rather than reconstructed from a broad component render.
Signals are a family of ideas
Different libraries use different APIs and scheduling rules.
“Signals” names a reactive primitive pattern, not one universal technology.
Signals are not automatically faster
Performance depends on:
- graph shape;
- update frequency;
- computation cost;
- scheduling;
- DOM work;
- memory and bookkeeping;
- developer usage.
A fine-grained mechanism can still perform unnecessary work if the graph or state model is poorly designed.
Scheduling is part of the model
flowchart TD
Mut["State Mutation"] --> Inval["Dependency Invalidation"]
Inval --> Queue["Job Scheduler Queue"]
Queue --> Flush["Microtask Flush"]
Flush --> Exec["Batch Render / Commit / Effects"]Scheduling determines what “immediately” means.
It affects batching, race conditions, DOM measurement, focus, and perceived responsiveness.
Why scheduling exists
Scheduling can:
- combine several changes;
- avoid repeated layout work;
- prioritize urgent interaction;
- defer expensive computation;
- coordinate asynchronous results;
- prevent recursive update storms.
The trade-off is that a state mutation and a visible result may be separated in time.
Unnecessary rendering work is not always a bug
Ask:
- did the calculation actually cost enough to matter?
- did the DOM change?
- did the user notice?
- did the work block input or layout?
- does optimization add more complexity than it removes?
Not every rerender is a problem.
Example: expensive filtering
This may help when products is large and the calculation is repeated.
It may do nothing useful when the array is small, dependencies change every time, or rendering dominates the cost.
Memoization should follow measurement
flowchart LR
Measure["Measure Performance"] --> Identify["Identify Repeated Work"]
Identify --> Optimize["Narrowest Targeted Optimization"]
Optimize --> Verify["Verify Behavior & Real Cost"]Memoization has costs:
- dependency maintenance;
- memory;
- identity management;
- cognitive overhead.
Component memoization is not a correctness fix
Memoization can skip a render when props are considered equal.
It cannot repair incorrect keys, duplicated state, stale closures, or a bad ownership boundary.
Vue’s selective tracking has its own costs
Property-level tracking can avoid broad updates.
But deep reactive objects, unstable identities, and unnecessary watchers can still make an application difficult to reason about.
Selectivity is a mechanism - not a substitute for a clear dependency graph.
Identity affects list rendering in every framework
flowchart TD
Identity["Stable Item Key / Identity"]
Identity --> State["Preserve correct row/sub-tree state"]
Identity --> Focus["Preserve active focus & form inputs"]
Identity --> Order["Deterministic list reordering & animations"]Stable keys are a user-experience concern as much as a rendering optimization.
State preservation is architectural
When state disappears unexpectedly, ask:
- did the component type change?
- did its position change?
- did its key change?
- did a conditional branch replace its identity?
- did the data identity change while the key stayed index-based?
The answer is often in the render tree, not in the state setter.
Conditional rendering can change identity
Switching between different component types replaces the identity at that position.
If two modes should preserve one shared draft, model them accordingly.
If switching should reset state, make that reset intentional.
Effects and watchers can create feedback loops
flowchart TD
SChange["State Change"] --> Eff["Effect Runs & Calls setState"]
Eff --> Loop["Triggers Re-render"]
Loop --> SChangeBefore adding synchronization, define:
- the external source of truth;
- the direction of the update;
- the stopping condition;
- cleanup and error behavior.
Think in reactive graphs
flowchart TD
Query["query (source)"] --> Visible["computed: visibleProducts"]
Products["products (source)"] --> Visible
Stock["stockFilter (source)"] --> Visible
Visible --> Grid["ProductGrid (View)"]
Query --> URLEffect["URL Sync Effect"]The graph reveals which values are sources, derivations, render consumers, and external effects.
It also reveals cycles and broad dependencies.
Compiler-assisted optimization
Compilers can sometimes infer:
- stable expressions;
- memoization opportunities;
- dependency relationships;
- component boundaries;
- update paths.
They cannot infer product intent, correct identity, or whether a value belongs in state.
Optimization tools reduce some manual work; they do not remove architectural decisions.
React Compiler is an example, not a new mental model
Compiler assistance may reduce the need for some manual memoization.
The fundamentals remain:
- pure render calculations;
- stable identity;
- correct dependencies;
- explicit effects;
- measured performance work.
Understand the behavior even when tooling automates an optimization.
Searchable product list: the shared dependency graph
flowchart TD
Q["query (source)"] --> F["filteredProducts (derived)"]
P["products (source)"] --> F
Flt["filter (source)"] --> F
F --> List["Product List View"]
Q --> URL["URL Sync Effect"]
List --> Analytics["Selection Analytics Effect"]The same product feature can be implemented in React, Vue, or a signal system while preserving this conceptual graph.
React version: calculate and synchronize
Derivation stays in the render model.
The router is synchronized in an effect.
React anti-pattern: derived state effect
This creates:
- duplicated state;
- an extra render path;
- a possible stale intermediate value;
- more code to test and debug.
Calculate the list directly unless there is a measured reason not to.
Vue version: computed and watch
The same graph is expressed with Vue’s dependency tracking primitives.
Trace React updates
For a state change, record:
- Which setter was called?
- Which snapshot created the handler?
- Which components calculate again?
- Which identities and keys are reused?
- Which host nodes commit changes?
- Which effects run and which clean up?
This sequence is more useful than saying “React rerendered everything.”
Trace Vue updates
For a reactive mutation, record:
- Which ref or proxy property changed?
- Which computed values depend on it?
- Which component render effects read it?
- Which watchers are triggered?
- When is the DOM queue flushed?
- Which cleanup functions run?
The goal is to expose the dependency graph and schedule.
What should we measure?
Measure questions such as:
- how long does the expensive calculation take?
- how many items are processed?
- how often does the calculation run?
- how many components commit changes?
- does input responsiveness degrade?
- does memory grow because of caches or subscriptions?
Choose a measurement that can change the decision.
Necessary UI change versus unnecessary work
flowchart TD
StateChange["State Changed"] --> Recalc["Recalculate Derived Descriptions"]
Recalc --> DiffCheck{"Are Descriptions Different?"}
DiffCheck -- Yes --> DOMMut["Necessary DOM Changes Committed"]
DiffCheck -- No --> Skip["Skip Host DOM Updates"]Do not optimize away work before identifying which part is unnecessary.
Correctness and a truthful dependency graph come first.
State location affects rendering scope
State placed high in the tree can coordinate many consumers but broaden update scope.
State placed close to one interaction can reduce unrelated work but may require a deliberate communication path.
Choose location based on ownership and synchronization - not on a universal rule to lift or localize state.
Derived state affects rendering scope
Keep derivation near the data and consumers that need it.
If a derived value is shared, expose one clear calculation rather than duplicating it in multiple components.
If it is cheap and local, a plain expression is often the best design.
Side effects should cross a boundary
flowchart LR
subgraph Pure["Pure Computation Zone"]
Render["Render: State → Desired UI Description"]
end
subgraph Boundary["Side Effect Boundary"]
Sync["Network, Storage, DOM, Timers, Subscriptions"]
end
Pure -->|Committed State| BoundaryThe boundary helps answer:
- when does this run?
- what invalidates it?
- how is it cleaned up?
- what happens if the value changes again?
Do not let the reactive system become a mystery network
A healthy graph has:
- visible sources;
- named derivations;
- few synchronization edges;
- bounded effects;
- explicit cleanup;
- tests for identity and timing.
If changing one field triggers an unexplained chain of watchers and effects, simplify the graph before optimizing it.
A practical decision model
Ask:
- Is another value directly calculable from this one?
- Does a user action cause an operation?
- Must an external system remain synchronized?
- Is a calculation expensive and repeated unnecessarily?
- Is the update scope too broad?
These questions map naturally to derived values, events, effects, memoization, and state placement.
Four reactive strategies
| Strategy | Main strength | Main risk |
|---|---|---|
| React render model | explicit component calculation | confusing render with DOM work |
| Vue dependency tracking | selective property updates | hidden reactive connections |
| signals | fine-grained graph | graph and lifecycle complexity |
| compiler assistance | less manual optimization | false confidence about architecture |
Use the model your team can explain and debug.
Practical lab: Reactive Computed Graph
Implement a tiny educational reactive graph to make source state, derived values, dependency tracking, and effects visible.
The implementation is for learning. It is not a production reactive runtime.
Practical stages 1–3: sources and derivation
- Implement a signal with subscribers.
- Add lazy computed values and invalidation.
- Add an effect with cleanup.
Make the graph observable with logs or counters so the execution order can be inspected.
Practical stages 4–5: cycles and comparisons
- Create a cycle deliberately and explain why production systems must guard against it.
- Compare the educational graph with React render calculation and Vue
computed().
Ask which mechanism discovers dependencies, when invalidation occurs, and how cleanup is represented.
Practical extension: visualize the graph
Add a graph visualizer showing:
flowchart LR
Source["Signal Source"] --> Comp1["Computed A"]
Comp1 --> Comp2["Computed B"]
Comp2 --> EffectNode["Effect Observer"]Highlight:
- invalidated nodes;
- execution order;
- cached nodes;
- unused computations;
- cycles.
The visualizer should make the invisible dependency model inspectable.
Try this yourself
Build a searchable list with:
- source products;
- query state;
- a derived filtered list;
- a URL synchronization effect;
- cancellation for stale searches.
Then explain which updates are necessary and which calculations can remain untouched.
Troubleshooting guide (Part 1)
| Symptom | Likely cause |
|---|---|
| State update seems one step behind | Snapshot or batching misunderstood |
| Input state moves to another row | Unstable or index-based keys |
| UI flashes an old derived value | Derived state synchronized through an effect |
| Effect runs forever | Effect updates one of its own dependencies |
Troubleshooting guide (Part 2)
| Symptom | Likely cause |
|---|---|
| Vue value stopped updating | Destructuring removed the reactive connection |
| DOM is old after a mutation | Update is queued; await the framework flush |
| Memoization changes nothing | Dependencies or calculation cost do not justify it |
| Async result wins after a newer query | Missing cleanup or cancellation |
Completion checklist
- source state is distinct from derived values;
- render calculations are free of external side effects;
- list keys represent stable identity;
- previous-state updates are used for dependent changes;
- effects and watchers synchronize external systems only;
- cleanup prevents stale subscriptions and requests;
- React and Vue timing differences are understood;
- optimization decisions are supported by measurement;
- the reactive graph is explainable from source to screen.
Misconceptions to leave behind (Part 1)
| Misconception | Better mental model |
|---|---|
| State changed, so the DOM changed immediately | Updates are calculated and scheduled |
| A React render rebuilt the DOM subtree | Render and commit are different phases |
| Virtual DOM is always faster | Performance depends on the measured work |
| Every rerender is a bug | Some recalculation is necessary |
| Keys only remove warnings | Keys preserve logical identity |
Misconceptions to leave behind (Part 2)
| Misconception | Better mental model |
|---|---|
| Derived values belong in state | Direct calculations usually stay derived |
| Effects respond to any state change | Effects synchronize external systems |
| Vue watchers calculate normal derivations | computed represents derivation |
| Signals are one standard technology | Signals are a family of mechanisms |
| Compiler optimization makes architecture irrelevant | Tools cannot choose ownership or intent |
The chapter in one sentence
Design a truthful reactive graph: keep source state minimal, derivation explicit, identity stable, scheduling understood, and side effects at the boundary.
Next: Chapter 8
The next chapter will apply these rendering and state principles to:
- asynchronous data and server state;
- loading, error, empty, and success states;
- request cancellation and stale results;
- caching and synchronization;
- resilient data-fetching architecture.
Questions
For one interaction in your application, can you draw the path from source state to derived value to committed UI?
Where does the first external side effect enter the graph?
8 - State Management, Routing & Form Architecture
State Management, Routing & Form Architecture
Put every value where it belongs
Chapter 8
Polla Fattah
Today’s goal
Stop treating all state as the same kind of problem.
We will connect:
- local, shared, domain, server, URL, form, persistent, and derived state;
- ownership, reducers, stores, and state machines;
- routing as application state architecture;
- paths, parameters, history, loading, and navigation UX;
- controlled, uncontrolled, and hybrid forms;
- validation, dirty state, dynamic fields, and multistep workflows;
- state placement in a routed administrative catalogue.
By the end of today you can
- classify a value before choosing a state tool;
- keep state close to its practical owner;
- explain why server data is not ordinary global state;
- model URL state as a serializable public view;
- design history and back-button behavior intentionally;
- separate form drafts from domain models;
- distinguish touched, dirty, validation, and submission state;
- use reducers and state machines for explicit transitions;
- choose local state, context, or a store proportionately;
- review a feature for duplicated or misplaced state.
The central principle
State architecture is the deliberate placement of values according to ownership, lifetime, sharing, persistence, and transition rules.
The right question is not “which state library should we use?”
The better question is “what kind of state is this, and who is responsible for it?”
The chapter’s progression
flowchart TD
A[Classify State] --> B[Assign Ownership]
B --> C[Choose Update Model]
C --> D[Design URL & Route State]
D --> E[Design Form State]
E --> F[Model Workflow Transitions]
F --> G[Test Navigation & Persistence]Tools come after the state model, not before it.
“State” is not one thing
flowchart TD
subgraph StateTaxonomy["Application State Taxonomy"]
S1["dialogOpen: Local UI State"]
S2["selectedTab: Local / URL State"]
S3["currentUser: Domain / Session State"]
S4["products: Server State"]
S5["query: URL / Local Draft State"]
S6["formDraft: Form State"]
S7["themePreference: Persistent Client State"]
S8["resultCount: Derived State"]
endThe category predicts who owns the value and how it should change.
A practical state taxonomy (Part 1)
| Category | Typical lifetime | Example |
|---|---|---|
| local UI | one component or feature | modal open |
| shared UI | several nearby components | selected tab |
| domain | business workflow | approval status |
| server | remote source | product list |
A practical state taxonomy (Part 2)
| Category | Typical lifetime | Example |
|---|---|---|
| URL | shareable view | page and filters |
| form | unfinished input | draft email |
| persistent | across sessions | theme preference |
| derived | calculated | filtered results |
Local UI state
Local state is usually:
- interaction-specific;
- short-lived;
- not useful to unrelated routes;
- safe to discard when the component leaves the tree.
Keep it local unless another owner genuinely needs it.
Shared UI state
Shared UI state coordinates a small group of related components.
Use a parent, compound component context, or a local feature store when the relationship is real.
Do not promote it globally merely because more than one component reads it.
Domain state
Domain state represents business meaning and rules.
It should not be confused with whether a modal is open or a request is currently loading.
Server state is a different category
Server state has:
- a remote owner;
- latency and failure;
- freshness and staleness;
- caching concerns;
- invalidation rules;
- multiple possible consumers.
Treating it as one ordinary global variable usually loses important behavior.
Server state is not “just another global variable”
The state includes the lifecycle of synchronization, not only the latest payload.
Cached data needs an ownership policy
Ask:
- who populated the cache?
- when is it stale?
- who invalidates it after a mutation?
- can two requests race?
- can an older response overwrite a newer one?
- what happens offline?
Caching is architecture, not merely a performance toggle.
URL state is public view state
URL state can survive:
- reloads;
- sharing;
- bookmarks;
- back and forward navigation;
- opening the view in another tab.
That makes it valuable - but also part of the public contract.
Form state is unfinished work
A draft is not automatically a valid domain object.
Form architecture should represent the journey from incomplete input to accepted data.
Persistent client state has a migration problem
Persisted values can outlive the code that created them.
Treat storage as an external boundary and design for old or malformed versions.
Derived state should usually be calculated
If visibleProducts is directly calculable from source state, storing it separately creates another value that can become stale.
Use memoization only when repeated computation is measurably expensive.
Ownership: who is responsible?
For every value, ask:
- who decides it?
- who needs to read it?
- who changes it?
- how long should it survive?
- should it be shareable or persisted?
- what external system owns the truth?
The answers identify the correct boundary better than a favorite library does.
Keep state as close as practical
Local ownership reduces coordination and update scope.
Move state outward only when a real consumer, lifecycle, or synchronization rule requires it.
State should move outward for a reason
Good reasons include:
- two siblings must coordinate;
- a route must own the view;
- a server cache is shared;
- a workflow spans multiple screens;
- another system must synchronize with the value.
“It might be useful later” is not an ownership rule.
Global state has a cost
Global state can create:
- invisible dependencies;
- broad update scope;
- unclear ownership;
- difficult isolated tests;
- stale values that survive too long;
- accidental coupling between features.
Make a value global because its lifetime and sharing demand it, not because global access is convenient.
Reducers make transitions explicit
The transition vocabulary makes state changes inspectable and testable.
Unidirectional data flow
flowchart LR
State["Current State"] --> Render["Render UI"]
Render --> Intent["User Action / Intent"]
Intent --> Action["Dispatch Action"]
Action --> Reducer["Transition Function"]
Reducer --> Next["New State Snapshot"]
Next -.-> StateOne direction makes it easier to answer:
- what caused this value?
- which action changed it?
- which consumers should update?
It does not mean every value belongs in one global store.
Stores are boundaries, not magic containers
A store should define:
- the state it owns;
- actions or methods that change it;
- selectors or derived values;
- effects and external dependencies;
- initialization and cleanup.
If a store becomes the home for every value, it has stopped communicating ownership.
Reducers versus stores
| Reducer | Store |
|---|---|
| transition function | longer-lived owner |
| often pure | may coordinate effects |
| easy to test as input/output | may expose selectors and subscriptions |
| useful inside local features | useful for shared domain workflows |
They can be combined. Neither is automatically the correct scale.
State machines make workflows visible
flowchart LR
Idle["Idle"] --> Editing["Editing"]
Editing --> Submitting["Submitting"]
Submitting --> Success["Success"]
Submitting --> ErrorState["Error"]
ErrorState -->|Retry| Submitting
ErrorState -->|Edit again| EditingState machines are useful when legal transitions matter more than storing a collection of booleans.
Replace impossible boolean combinations
Avoid:
Prefer one explicit status:
The model should make impossible states difficult to represent.
Routing is state architecture
A route decides more than the address.
It can determine:
- which feature is active;
- which data loads;
- which layout persists;
- what can be shared;
- what the back button restores;
- which code is loaded.
Routing is a state-lifetime and ownership decision.
Paths represent resource identity
Use a path segment when the value identifies the resource or nested location being viewed.
The route should speak domain language, not reveal component filenames.
Route parameters identify a resource
Validate the parameter before using it as a domain identifier.
The string from the URL is external input, not automatically a valid ProductId.
Query parameters describe a view
Query parameters are suitable for filters, sorting, pagination, search, and other shareable view choices.
They should be serializable, parseable, and stable enough to form a public contract.
Path parameter or query parameter?
Use the semantic distinction, not personal preference.
Nested routes express nested ownership
Parent routes can own layout, permissions, data context, and persistent navigation while child routes own the active view.
Layout routes preserve context
flowchart TD
Layout["AdminLayout (Persistent Shell)"]
Layout --> Sidebar["Sidebar Nav"]
Layout --> Header["Top Header & User Info"]
Layout --> Outlet["Child Route <Outlet />"]The layout can remain mounted while the child changes.
This preserves navigation context and avoids rebuilding shared structure unnecessarily.
Navigation changes state and history
flowchart TD
Push["push: New history entry (back returns here)"]
Replace["replace: Update current entry (no new history step)"]
Back["back / forward: Restore historical view snapshot"]Choose history semantics deliberately.
Typing a filter may replace the current entry.
A meaningful page transition may push a new entry.
Redirects are state transitions
Redirect when:
- a user lacks access;
- a resource moved;
- a form completed a workflow;
- a canonical URL should replace an invalid one.
Preserve useful context where appropriate, and avoid redirect loops that obscure the actual state.
File-based routing is a convention
File structure can make routes discoverable:
It helps organize the map, but it does not decide which state belongs in the route or how transitions should behave.
The URL is state
flowchart LR
URL["URL Address Bar"] <-->|"Parse / Serialize"| State["Parsed Route State"]
State <-->|"Render / Events"| View["Rendered View"]Do not copy URL values into local state without a clear ownership reason.
Two sources of truth create synchronization work and surprising back-button behavior.
Shareable state belongs naturally in the URL
Good candidates:
- search query;
- filters;
- sorting;
- pagination;
- selected public tab;
- view density when sharing it matters.
If another person opens the URL, the meaningful view should be reconstructable.
What usually should not go in the URL?
Avoid placing:
- passwords or secrets;
- sensitive personal data;
- large drafts;
- ephemeral hover state;
- internal implementation details;
- values that cannot be serialized safely.
The URL is visible, copyable, logged, and often shared.
URL state should have one owner
Avoid:
Define the stages explicitly:
flowchart LR
Draft["1. Local Draft Input\n(Immediate keystrokes)"] -->|"Commit (Enter/Blur/Debounce)"| URL["2. Committed URL Query\n(Shareable & Back-button aware)"]
URL -->|"Data Fetch / Cache"| ServerReq["3. Server API Request\n(Cancellable async query)"]Each stage has a different purpose and transition.
Parse URL state at the boundary
Parsing should:
- provide defaults;
- reject or normalize invalid values;
- clamp unsafe ranges;
- preserve only supported vocabulary;
- return a trusted view model.
Example: URL-based filtering
The view consumes state, not raw strings from location.search.
Routing UX is part of architecture
A route transition should define:
- pending feedback;
- loading behavior;
- error boundaries;
- scroll restoration;
- focus placement;
- unsaved-change handling;
- code-loading behavior.
The route is a user interaction, not merely a URL replacement.
Pending navigation needs visible feedback
flowchart TD
Req["Navigation Requested"] --> Indicator["Pending Feedback Indicator"]
Indicator --> Resolve["Data & Code Bundle Resolve"]
Resolve --> CommitRoute["New Route Commits to DOM"]Keep the current context understandable while the next view is loading.
Avoid showing a blank screen for an operation that can preserve useful layout.
Loading states should belong to the right scope
flowchart TD
S1["Route shell loading: Page-level fallback / spinner"]
S2["Panel data loading: Panel skeleton loader"]
S3["Button mutation: Local inline pending state"]One global spinner often hides which part of the interface is actually unavailable.
Route errors are part of the route contract
Model distinct causes:
- invalid route parameter;
- missing resource;
- permission failure;
- network failure;
- unexpected application error.
The user should receive the recovery action appropriate to the cause.
Scroll restoration is state restoration
Decide whether navigation should:
- restore the previous scroll position;
- start a new route at the top;
- preserve a nested panel’s scroll;
- maintain position while query filters update.
Surprising scroll behavior makes a correct route feel broken.
Focus after navigation is accessibility state
After a route changes, move focus to a meaningful landmark or heading when appropriate.
Do not leave keyboard users at an old location with no indication that the view changed.
Focus management belongs to the route transition boundary.
State preservation across routes is a choice
Ask:
- should the parent layout remain mounted?
- should the child form survive route changes?
- should a tab selection be encoded in the URL?
- should changing an ID reset local state?
Route nesting, component identity, and state ownership work together.
Route-level code splitting
flowchart TD
Shell["Initial Application Shell"] --> CatRoute["Load Catalogue Route Chunk"]
CatRoute --> EditRoute["Load Edit Route Chunk on Demand"]Splitting at route boundaries can reduce initial work and align code loading with user navigation.
The route model and build model should reinforce one another.
Form architecture is state architecture
A form includes more than input values:
Design these lifecycles explicitly instead of allowing them to emerge from scattered event handlers.
Controlled forms
Benefits:
- immediate application visibility;
- easy derived feedback;
- explicit formatting and validation;
- synchronization with other state.
Costs include rerenders, wiring, and more code for large forms.
Uncontrolled forms
The browser owns current input values until submission or a deliberate read.
Native form behavior can be simple, efficient, and accessible when the application does not need every keystroke.
Hybrid form architectures
flowchart LR
Native["Native Input Editing"] --> Draft["Local Form Draft State"]
Draft --> Validation["Controlled Validation Engine"]
Validation --> Command["Submitted Domain Command"]Use control where coordination is needed and native behavior where it is enough.
Do not choose one mode for ideological reasons.
Form values are not automatically domain values
Forms represent unfinished, string-heavy input.
Parse and validate before constructing a domain command.
Keep the form model separate when useful
The draft supports editing.
The command represents validated intent.
Touched state records interaction
Use it to decide when a field-level message should appear.
Touched does not mean the value changed and does not mean the value is valid.
Dirty state records change from an initial value
Dirty state answers whether there may be unsaved work.
It is different from touched state: a user can touch a field and return it to its original value.
Validation has several layers
flowchart TD
F1["1. Field-Level: Format, length, regex (e.g. Email syntax)"]
F2["2. Cross-Field: Relationship rules (e.g. End date after Start date)"]
F3["3. Domain/Business: Entity rules (e.g. Permit status editable)"]
F4["4. Server-Side: Authority validation (e.g. Unique registration number)"]
F1 --> F2 --> F3 --> F4Keep the layer visible so the UI can show the right message and recovery path.
Derive validation where practical
Do not store every validation message if it can be derived from the current draft and interaction state.
Store server results and asynchronous validation when they are not directly calculable.
Cross-field validation needs the whole draft
The validation owner should receive the relevant model rather than forcing fields to synchronize through unrelated global state.
Dynamic fields need stable identity
Use a stable draft key for rows that can be inserted, removed, or reordered.
Index identity can move input values to a different line item.
Dynamic field operations are domain actions
Represent these operations explicitly in a reducer or form API.
Scattered array mutations make dirty state, validation, and focus behavior harder to preserve.
Multistep forms are workflows
flowchart LR
Account["Account"] --> Profile["Profile"]
Profile --> Review["Review"]
Review --> Submit["Submit"]Each step needs:
- an entry condition;
- validation scope;
- persistence policy;
- back behavior;
- recovery from invalid or stale data.
Treat the wizard as a state machine when transitions are meaningful.
URL and multistep forms can work together
Put the current step in the URL when it should survive reload, sharing, or navigation.
Keep sensitive drafts and unfinished values in an appropriate private owner.
A wizard state machine
flowchart LR
Incomplete["profileIncomplete"] --> Complete["profileComplete"]
Complete --> InReview["reviewStep"]
InReview --> Submitting["submitting"]
Submitting --> Done["success"]
Submitting --> Fail["serverError"]
Fail -->|Retry| SubmittingExplicit transitions prevent the UI from entering a review step without the required data.
When native HTML is enough
Prefer native forms when:
- fields are simple;
- browser validation is adequate;
- every keystroke need not update application state;
- submission can read
FormData; - accessibility should follow platform behavior.
Native behavior is a feature, not an implementation failure.
When a form library is justified
A library can help with:
- many fields and nested structures;
- reusable validation patterns;
- touched and dirty tracking;
- field arrays;
- controlled/uncontrolled integration;
- performance and subscription granularity.
Adopt it for repeated complexity, not for a two-field form.
Administrative catalogue: classify before building
The map explains where each value should live.
Avoid duplicating URL state
Bad:
flowchart LR
subgraph Bad["Anti-Pattern: Duplicating URL State Across Stores"]
B1["URL Page"] --> B2["Local Page State"] --> B3["Global Store Page"] --> B4["Request Page"]
end
subgraph Good["Architectural Best Practice: Direct Pipeline"]
G1["URL Page"] --> G2["Parsed Route State"] --> G3["Request Query"]
endCreate a local draft only when editing and committing are intentionally different states.
Avoid copying server data into the form too early
This is a deliberate transition, not a continuous mirror.
After editing starts, the draft can diverge safely until save or reset.
Keep server freshness and local unsaved work conceptually separate.
Modal state should stay local unless navigation owns it
Put the modal in the URL only when the modal itself should be deep-linkable, restorable, or part of browser history.
Not every visible state deserves a route.
Selected tabs: local or URL?
Use local state when:
- the tab is ephemeral;
- sharing the selection is not useful;
- back navigation should not record every change.
Use URL state when:
- the selected view is meaningful to share;
- reload should preserve it;
- browser history should restore it.
Pagination, sorting, and search
These often belong in the URL because they define the collection view.
But distinguish:
The user can type freely without creating a history entry for every keystroke, then commit the meaningful query deliberately.
Persistent preferences are not route state
A preference follows the user across views.
A route value describes one navigable view.
Their lifetimes and ownership differ.
One value can change category over time
The value’s role changes through explicit transitions.
Do not force every stage into one state container.
A state placement decision tree
flowchart TD
Q1{"Is it calculable from existing data?"} -- Yes --> A1["Derive it (Pure computed / inline)"]
Q1 -- No --> Q2{"Is an external system the owner?"}
Q2 -- Yes --> A2["Synchronize at boundary (API/Storage)"]
Q2 -- No --> Q3{"Must it be bookmarkable / shareable?"}
Q3 -- Yes --> A3["URL Query Parameter"]
Q3 -- No --> Q4{"Must it survive reloads / sessions?"}
Q4 -- Yes --> A4["Persistent Client Storage (localStorage/IndexedDB)"]
Q4 -- No --> Q5{"Who needs it?"}
Q5 -- Single component --> A5["Local Component State"]
Q5 -- Component subtree --> A6["Context / Provide-Inject"]
Q5 -- Cross-feature domain workflow --> A7["Dedicated Store / State Machine"]This is a reasoning aid, not a mechanical law.
Common state smells
Watch for:
- state copied from props without a reset rule;
- derived values stored as state;
- everything placed in a global store;
- one value duplicated in URL and component state;
- effects used to synchronize local copies;
- server data copied into forms continuously.
Each smell suggests competing owners.
Routing smells
Warning signs:
- route reflects implementation rather than domain;
- important view state disappears on refresh;
- back button behaves surprisingly;
- child navigation destroys useful parent context;
- a rarely used route inflates the initial bundle;
- invalid parameters reach data-fetching code unparsed.
Form smells
Warning signs:
- every keystroke updates a global store;
- validation is duplicated across fields and submit handlers;
- draft and domain models are forced to be identical;
- dirty and touched flags are synchronized manually everywhere;
- a multistep form uses unrelated booleans instead of transitions.
React URL-state catalogue
The URL owns committed view state.
The component receives a parsed model rather than raw strings.
Vue URL-state catalogue
Again, compare ownership and transitions rather than framework syntax.
Forms in React and Vue
flowchart LR
subgraph ReactWorld["React Paradigm"]
R1["Controlled Input (value + onChange)"]
R2["useReducer (state + dispatch)"]
R3["useEffect (synchronization)"]
end
subgraph VueWorld["Vue Paradigm"]
V1["v-model (two-way binding syntax)"]
V2["reactive + action methods"]
V3["watch / watchEffect"]
end
R1 <-->|Equivalent| V1
R2 <-->|Equivalent| V2
R3 <-->|Equivalent| V3The important design questions remain:
- what is the draft?
- what is valid?
- what is submitted?
- who owns the transition?
A complex edit form model
The model separates current values, comparison baseline, interaction history, validation, and submission lifecycle.
A form reducer makes operations visible
Explicit actions make it possible to test dirty state, touched state, validation, and reset behavior as transitions.
Route plus form interaction
flowchart TD
Route["1. Route: /products/p-42/edit"] --> Load["2. Load Server Product into Cache"]
Load --> Init["3. Initialize Local Form Draft (Detached copy)"]
Init --> Edit["4. User Edits Draft (Zero mutation of server cache)"]
Edit --> Validate["5. Validate & Submit Command"]
Validate --> Inval["6. Invalidate / Update Server State"]
Inval --> Nav["7. Navigate to Canonical Resource View"]Each arrow is a deliberate ownership transition.
Unsaved changes need a policy
When a dirty form meets navigation, decide:
- block and confirm;
- autosave;
- preserve a draft;
- discard explicitly;
- allow navigation and make loss clear.
Do not let a route unmount silently destroy work the user believes is still present.
Persistence for long forms
Persist only what is appropriate:
- version the draft;
- exclude secrets;
- expire stale drafts;
- validate on restore;
- show the user what was restored;
- provide clear reset behavior.
Persistence is another external boundary with lifecycle and privacy decisions.
Local state versus context versus store
flowchart TD
Local["Local State: One component / feature owns it"]
Context["Context State: Related family shares it without prop drilling"]
Store["Store State: Cross-feature domain workflow / server data"]
Local --> Context --> StoreStart with the smallest scope that satisfies the real consumers.
Server state versus client state
flowchart TD
S1["Server Data: Remote owner, asynchronous, stale, refetchable"]
S2["Client UI State: Local user choices, immediate interaction, ephemeral"]
S3["Domain State: Business transitions, validated models, domain rules"]The same object may be represented in more than one layer, but each representation needs a clear owner and synchronization rule.
URL state versus persistent state
flowchart LR
URL["URL State: Current shareable view & navigation bookmark"]
Storage["Persistent Storage: Long-lived preference or offline draft across sessions"]The URL is visible and navigable.
Persistence survives beyond one route and may need migration.
Do not use one as a substitute for the other.
State ownership and testing
Clear ownership makes focused tests possible:
- parser tests for URL state;
- reducer tests for transitions;
- form tests for validation and dirty behavior;
- route tests for history and redirects;
- server-state tests for loading and stale responses.
If every test needs the whole application, ownership may be too global.
State ownership and team scale
As a team grows, implicit ownership becomes expensive.
Document:
- which module owns a value;
- which API changes it;
- what is public URL state;
- what is cached server data;
- what can be persisted;
- which transitions are legal.
Architecture reduces coordination cost when boundaries are explicit.
Practical lab: Routed Administrative Catalogue
Build a catalogue whose filters, sorting, pagination, and selected view are represented in the URL when they should survive reload, sharing, and history navigation.
Then add a routed edit form with explicit ownership for draft, validation, server state, and workflow transitions.
Practical stages 1–4: inventory and URL state
- Create the state inventory.
- Build URL-based filters.
- Parse and validate query parameters.
- Keep server state separate.
Verification: a copied URL reconstructs the same meaningful view without exposing sensitive data.
Practical stages 5–8: local UI and editing
- Add local modal state.
- Convert editing to a route.
- Create a complex edit form.
- Track touched and dirty state.
Do not copy server data into a continuously synchronized global form object.
Practical stages 9–12: validation and transitions
- Add cross-field validation.
- Add dynamic fields with stable keys.
- Add a reducer.
- Model the workflow as a state machine.
Make invalid combinations and illegal transitions visible in the model.
Practical stages 13–16: resilient navigation
- Add navigation UX.
- Add route-level code splitting.
- Add persistence with versioning and validation.
- Create the final state map.
Include loading, error, retry, unsaved-change, and back/forward behavior.
Practical extension: debounced server search
Add:
flowchart LR
Input["Local Keystrokes"] --> Debounce["Debounce Timer"]
Debounce --> Cancel["AbortController (Cancel in-flight)"]
Cancel --> Cache["Server-State Cache (Query Result)"]Keep the local suggestion draft separate from the committed URL query.
Ensure an old response cannot overwrite a newer query’s result.
Try this yourself
For a product editor, decide where these values belong:
Write the owner and lifetime beside each one before writing components.
Troubleshooting guide (Part 1)
| Symptom | Likely cause |
|---|---|
| URL and input disagree | Two owners exist for the same value |
| Back button feels noisy | Every transient edit pushed history |
| Refresh loses the view | Shareable state stayed local |
| Old API data overwrites new data | Server-state race lacks cancellation or identity |
Troubleshooting guide (Part 2)
| Symptom | Likely cause |
|---|---|
| Form loses work on navigation | Dirty policy is undefined |
| Validation flickers | Draft, touched, and errors are conflated |
| Store contains everything | State categories were never classified |
| Child route feels like a full reset | Parent context or layout is not preserved |
Completion checklist
- every important value has a category and owner;
- derived values are not duplicated unnecessarily;
- server state has freshness and failure semantics;
- URL state is serializable, validated, and shareable;
- history semantics are intentional;
- form drafts are separate from accepted domain commands;
- dirty, touched, and validation state have distinct meanings;
- reducers or state machines model meaningful transitions;
- persistence is versioned and safe;
- route and form behavior is tested from the user’s perspective.
Misconceptions to leave behind (Part 1)
| Misconception | Better mental model |
|---|---|
| State management means choosing a library | First classify ownership and lifetime |
| All shared state belongs in a global store | Share only across real consumers |
| Server data becomes ordinary client state | It has freshness, cache, and synchronization rules |
| The URL is just routing | It is public, serializable view state |
| Every option belongs in the URL | Ephemeral and sensitive values have other owners |
| Forms are collections of controlled inputs | Forms are workflows with drafts and transitions |
Misconceptions to leave behind (Part 2)
| Misconception | Better mental model |
|---|---|
| Controlled forms are always better | Choose based on coordination needs |
| Draft and domain models must match | Input representation can differ from accepted data |
| Dirty and touched mean the same thing | They record different interaction facts |
| A reducer is only for global state | Local complex transitions benefit too |
| Back/forward is only the router’s problem | History behavior is product architecture |
The chapter in one sentence
Classify state by ownership and lifetime, keep one clear source of truth, and make routing and forms explicit state machines for the user’s journey.
Next: Chapter 9
The next chapter will build on state architecture with:
- resilient asynchronous data flows;
- request lifecycle design;
- caching, invalidation, and optimistic updates;
- loading and error boundaries;
- race-free server synchronization.
Questions
Which value in your current application has two owners?
What should happen to it on reload, sharing, back navigation, and an interrupted request?
9 - Client–Server Communication, APIs & Cache Management
Client–Server Communication, APIs & Cache Management
Make remote data predictable
Chapter 9
Polla Fattah
Today’s goal
Build a client that treats the network as an unreliable, stateful boundary.
We will connect:
- HTTP methods, headers, status codes, and request helpers;
- runtime validation and transport/domain separation;
- cancellation, timeouts, retries, and backoff;
- REST, GraphQL, pagination, and API compatibility;
- loading, empty, stale, error, and recovery states;
- query keys, deduplication, freshness, and invalidation;
- pessimistic and optimistic mutations;
- server validation, authentication failures, and partial data.
By the end of today you can
- implement a fetch boundary that handles HTTP errors explicitly;
- distinguish network, transport, schema, domain, and authentication failures;
- choose safe retry behavior;
- model remote UI states without conflating loading and empty;
- design stable cache keys and freshness policies;
- deduplicate requests and cancel stale work;
- invalidate or update caches after mutations;
- use optimistic updates with rollback only when justified;
- preserve useful partial data during dashboard failures;
- draw the full client-server data architecture.
The central lesson
Remote data is not a value; it is a lifecycle of requests, cache entries, freshness, failures, mutations, and recovery.
The UI should make that lifecycle visible without forcing every component to understand HTTP details.
The chapter’s progression
flowchart TD
A[HTTP Boundary] --> B[Transport & Domain Layers]
B --> C[Request Lifecycle]
C --> D[Remote UI States]
D --> E[Query Cache]
E --> F[Freshness & Invalidation]
F --> G[Mutations & Rollback]
G --> H[Resilient App Architecture]Every layer answers a different question about remote data.
Browser and server have different responsibilities
The client cannot assume that the server is fast, available, current, or correct for every request.
HTTP is the communication foundation
Each part carries contract information.
Do not treat an HTTP exchange as merely “fetch some JSON”.
HTTP methods express intent
| Method | Typical intent |
|---|---|
| GET | read a representation |
| POST | create or trigger an operation |
| PUT | replace a resource representation |
| PATCH | partially modify a resource |
| DELETE | remove a resource |
The exact API contract still matters; method names are not permission to guess semantics.
Safe and idempotent are different properties
GET is generally safe and idempotent.
PUT is commonly idempotent but not safe.
POST is not automatically idempotent, so retries require care.
Request headers are part of the contract
Headers communicate representation, credentials, caching, conditional requests, and client capabilities.
Accept and Content-Type answer different questions
Confusing them can produce content negotiation and parsing bugs that look like application errors.
Status codes are part of the API contract
The client should classify status codes into user-visible recovery paths instead of displaying one generic failure for everything.
fetch() does not reject for every HTTP error
A 404 or 500 response can resolve normally.
Network failures typically reject; HTTP failure status still needs explicit handling.
A small fetch helper
The helper centralizes transport behavior, but it must not hide meaningful errors or bypass runtime validation.
Network and domain layers should not be confused
The adapter knows status codes and headers.
The domain feature knows what a valid product or workflow means.
Keeping those responsibilities separate makes change and testing cheaper.
JSON is a representation, not a type guarantee
JSON can represent an object that is:
- missing fields;
- using the wrong types;
- from an older server version;
- structurally valid but semantically invalid.
Parse at the boundary before trusting it.
Request bodies need an explicit representation
The request body is a transport representation, not automatically the same as the form draft or domain entity.
Credentials and cookies change the boundary
Authentication, CSRF protection, CORS, expiration, and logout behavior are part of the client-server contract.
Never treat a missing credential as an ordinary empty response.
Cancellation is correctness, not only optimization
Cancellation prevents obsolete work from consuming resources or updating a UI that now represents a different request.
Timeouts are application policy
Choose timeout behavior based on user action, operation cost, connectivity, and retry policy.
Retry carefully
Retrying can help transient failures.
It can also:
- duplicate a non-idempotent mutation;
- overload a struggling server;
- delay a useful error;
- hide an authorization failure;
- create a thundering herd.
Retry only when the operation and policy justify it.
Exponential backoff spreads retries
flowchart TD
Att1["Attempt 1: Fails (Network / 5xx)"] -->|"Wait 1,000ms"| Att2["Attempt 2: Fails"]
Att2 -->|"Wait 2,000ms + Jitter"| Att3["Attempt 3: Fails"]
Att3 -->|"Wait 4,000ms + Jitter"| Att4["Attempt 4: Abort / Surface Error"]Add jitter so many clients do not retry at the same instant.
Bound the number of attempts and expose a recovery action when automatic retry ends.
API design affects front-end architecture
The API determines:
- what can be fetched independently;
- how mutations are represented;
- which fields are stable;
- how pagination works;
- which errors can be recovered;
- which cache entries need invalidation.
Client architecture cannot fully compensate for an ambiguous server contract.
REST is resource-oriented
Resource-oriented URLs can map naturally to cache keys and route identities.
They are a style, not a promise that every API has identical semantics.
Collections and resources have different concerns
Do not assume that invalidating one product automatically answers what should happen to every filtered collection containing it.
Pagination is part of the cache model
Decide whether pages are:
- independent cache entries;
- concatenated into an infinite list;
- invalidated together after a mutation;
- prefetched based on navigation.
Filtering and sorting belong in the request identity
Changing any relevant input should produce a distinct query key or an explicit cache update.
Never reuse a cache entry for a request with different semantics.
API versioning protects compatibility
Versioning may be represented through:
- URL paths;
- headers;
- content negotiation;
- additive schema evolution.
The client should validate the response it actually receives, even when generated types say the version is known.
GraphQL asks for a shape of data
The client requests fields and receives a schema-described response.
This changes transport shape and caching questions; it does not eliminate them.
GraphQL schemas and mutations
Schemas can provide:
- typed fields;
- discoverable operations;
- nested data selection;
- validation at the operation boundary.
Mutations still need authorization, error handling, cache updates, conflict policies, and runtime behavior.
REST and GraphQL can coexist
Use the boundary that fits the resource and team needs.
A system may use REST for uploads and resource mutations, GraphQL for composed reads, and a separate stream for live updates.
The architecture should make each boundary’s semantics explicit.
Remote data has UI states
Represent the lifecycle instead of reducing it to data plus isLoading.
Loading is not the same as empty
The recovery and user message differ.
An empty catalogue may need onboarding.
A loading catalogue needs progress or preserved context.
Loading UI should match scope
A global spinner can erase useful context and make the application feel slower.
Spinner, skeleton, or existing content?
Choose based on what the user can still do and how much layout is known.
Empty state is product state
An empty result may mean:
- no records exist yet;
- filters are too restrictive;
- the user lacks access;
- a search has no matches;
- the resource was removed.
Show the cause and the next useful action when possible.
Error states need recovery
An error component should answer: what failed, what remains usable, and what can the user do next?
Client caching is a policy
A cache answers:
- may this result be reused?
- for how long?
- when should it revalidate?
- how is it invalidated?
- can stale data remain visible?
- who owns the cache?
Caching is not simply “store the last response”.
Freshness depends on the data
Freshness is a product and domain decision, not one global number.
Stale does not mean wrong
Stale data can still be useful while a background request checks for newer data.
The UI should communicate that it is refreshing when the distinction matters.
Discarding useful data on every refresh failure can create a worse experience than showing stale data with a clear warning.
Stale-while-revalidate
flowchart LR
CacheHit["Cache Hit"] --> RenderCached["Render Cached Data Immediately"]
RenderCached --> BackgroundReq["Background Revalidation Request"]
BackgroundReq --> Update["Update Cache & Silent UI Re-render"]This pattern separates immediate usefulness from eventual freshness.
It requires a policy for errors, timestamps, and concurrent requests.
Cache keys identify query meaning
A cache key must include every input that changes the response.
It should be deterministic, serializable, and stable across callers.
Missing key inputs cause incorrect reuse
Bad:
when the response varies by query, page, sort, account, or locale.
The cache then returns data that is valid for a different request but wrong for this view.
Deduplicate identical requests
flowchart LR
CompA["Component A"] --> Dedup["In-Flight Request Deduplicator\n(Promise Cache)"]
CompB["Component B"] --> Dedup
CompC["Component C"] --> Dedup
Dedup --> Single["Single HTTP Network Request"]
Single --> Shared["Shared Server Result Broadcast"]Deduplication reduces duplicate work and makes concurrent consumers observe one request lifecycle.
The cache must distinguish an in-flight promise from a completed fresh value.
Revalidation asks whether data changed
Revalidation can happen:
- when data becomes stale;
- when a route gains focus;
- after a mutation;
- on explicit refresh;
- after reconnecting;
- through a conditional HTTP request.
Choose triggers that match the user’s need for freshness.
Invalidation removes confidence, not necessarily data
flowchart TD
Mut["Mutation Succeeds on Server"] --> Inval["Mark Related Cache Keys Stale / Invalid"]
Inval --> Refetch["Trigger Active Query Revalidation"]
Refetch --> Fresh["Render Canonical Server Truth"]Invalidation is a statement that cached knowledge may no longer be current.
It does not always mean immediately deleting useful visible data.
Invalidation versus direct cache update
Use invalidation when:
- many related queries may change;
- server logic computes fields the client cannot reproduce;
- correctness is more important than avoiding a request.
Directly update one entry when:
- the mutation response is authoritative;
- the affected cache shape is known;
- the update is easy to verify.
Mutations have their own lifecycle
flowchart LR
Idle["Idle"] --> Submitting["Submitting"]
Submitting --> Success["Success"]
Submitting --> ErrorState["Error"]
ErrorState -->|Retry| SubmittingAdd:
- duplicate-submission protection;
- server validation handling;
- authentication behavior;
- cache reconciliation;
- draft preservation on failure.
Mutation UX communicates authority
flowchart TD
Pess["Pessimistic: Wait for authoritative server 200 OK before updating UI"]
Opt["Optimistic: Mutate UI immediately; rollback if server rejects request"]Choose based on reversibility, conflict risk, user expectations, and the cost of being briefly wrong.
Prevent duplicate submission at several layers
Client disabling helps the interface.
It does not guarantee that duplicate requests cannot arrive.
Use server-side idempotency keys or operation identifiers where repeating an operation has real cost.
Pessimistic updates are conservative
flowchart TD
Sub["Submit Command"] --> Pending["Display Inline Pending State"]
Pending --> Ok["Server Success: Update Canonical UI"]
Pending --> Err["Server Error: Preserve Draft & Show Recovery Action"]Use this for high-risk, non-reversible, or conflict-sensitive operations where showing unconfirmed state would mislead users.
Optimistic updates trade certainty for responsiveness
flowchart LR
Action["User Action"] --> Snapshot["1. Snapshot Previous Cache"]
Snapshot --> OptimisticUI["2. Update Cache & UI Instantly"]
OptimisticUI --> Network["3. Dispatch Server Request"]
Network --> Confirm["4a. Server 200: Confirm & Revalidate"]
Network --> Rollback["4b. Server Error: Restore Snapshot & Alert"]The client must know how to undo the change and what to show if the server rejects it.
Optimistic rollback needs a snapshot
In real systems, also consider concurrent edits, invalidation, and a final revalidation.
Not every mutation should be optimistic
Avoid optimism when:
- the action is irreversible;
- server rules are complex;
- conflicts are likely;
- authorization may fail often;
- rollback is ambiguous;
- the visible result depends on server-side computation.
Fast feedback is not worth misleading the user.
Conflict and concurrency need a policy
Two clients may update the same resource.
Possible strategies include:
- last write wins;
- version checks;
- conditional requests;
- conflict UI;
- server merge rules;
- explicit refresh before editing.
The cache cannot solve a domain conflict by itself.
Conditional requests use server versions
The server can reject a mutation when the resource changed since the client read it.
This prevents a stale edit from silently overwriting newer data.
Browser HTTP cache versus application cache
| Browser HTTP cache | Application/query cache |
|---|---|
| controlled by HTTP semantics | controlled by application policy |
| stores representations | stores parsed/query-aware data |
| keyed by request semantics | keyed by query and domain inputs |
| works below application code | exposes freshness and invalidation |
They can cooperate, but they are not the same layer.
Cache-Control is a server instruction
HTTP caching can reduce network work before application code runs.
The application still needs its own policy for query state, mutations, and visible stale/error behavior.
Forms and server mutations
flowchart TD
Draft["1. Form Draft State"] --> Validate["2. Validate Form Locally"]
Validate --> Transport["3. Submit Transport Command via Fetch"]
Transport --> ServerVal["4. Server-Side Validation & Authorization"]
ServerVal --> Commit["5. Commit or Preserve Draft on Error"]A successful request does not automatically mean every form field, cache entry, and route is now synchronized.
Native form submission remains useful
Native submission provides:
- keyboard behavior;
- browser integration;
- progressive enhancement;
- a clear submit event;
FormDataserialization.
Enhance it when application behavior requires it; do not discard platform behavior without a reason.
FormData is a transport boundary
Values may be strings, files, or null.
Validate and convert them before building a domain command or JSON payload.
JSON form submission is an explicit conversion
The parser is where unfinished input becomes a validated transport model.
Server validation errors are different from local errors
Both may appear near a field, but they have different owners, timing, and recovery paths.
Do not overwrite a useful server response with a generic client message.
Map server errors to fields carefully
Only map an error to a field when the server contract identifies that field reliably.
Some failures belong to the form or page, not one input.
Do not merge client and server validation carelessly
Track the sources distinctly:
Clear server errors when the relevant draft changes if the server response is no longer applicable.
Keep the message that explains the actual current failure.
Progressive enhancement for forms
flowchart TD
Native["1. Native HTML form submit works without JavaScript"] --> Enh["2. Progressive Enhancement with JS: Pending states, client validation & SPA transitions"]The enhanced path should preserve the core meaning of the native form rather than creating a completely different contract.
File uploads have different transport needs
Do not set a JSON content type for multipart data manually.
Model upload progress, cancellation, size limits, type validation, and partial failure explicitly.
Authentication failures are not ordinary validation errors
The correct response may be sign-in, permission explanation, or field correction - not a red error beneath an input.
Search needs query identity, cache, and cancellation
flowchart LR
QA["Query A (/search?q=ca)"] --> CancelA["Abort in-flight Controller A"]
CancelA --> QB["Query B (/search?q=cat)"]
QB --> Network["Execute Request B"]An old response must not replace results for a newer query.
The query key should include every input that changes the result.
Query keys and URL state fit naturally together
The URL provides the committed view identity.
The query cache stores the server result for that identity.
Parse once and derive both from the same validated model.
Avoid copying query data into another store
Copying into a general client store creates:
- duplicate ownership;
- stale copies;
- unclear invalidation;
- extra synchronization effects.
Keep server data in a server-aware cache unless there is a deliberate transformation or offline model.
Dependent queries form a data graph
flowchart LR
User["currentUser Query"] -->|"Resolves accountId"| Account["account Query"]
Account -->|"Resolves credentials"| Tx["transactions Query"]Start the dependent request only when its required input is known.
Represent the not-ready state instead of sending malformed requests with missing identifiers.
Request waterfalls cost time
If B and C do not depend on A, request them in parallel.
If they do depend, consider server composition, prefetching, or a route loader that understands the graph.
Parallel fetching
Use parallel work for independent requests.
Choose Promise.allSettled() or explicit result handling when partial success is useful.
Partial failure can preserve useful data
flowchart TD
Dash["Dashboard View"]
Dash --> Q1["Sales Widget (Status: Success)"]
Dash --> Q2["Inventory Widget (Status: Error + Retry Button)"]
Dash --> Q3["Alerts Widget (Status: Success)"]Do not blank the entire dashboard because one independent panel failed.
The architecture should allow panels to own their own request lifecycle.
Promise.all() versus partial failure
Promise.all() is appropriate when all results are required for one operation.
allSettled() or independent query states are appropriate when each panel can recover separately.
Background refresh preserves context
Keep current rows visible while a refresh runs when stale data remains useful.
Show that the data is refreshing without turning every background check into a blank loading screen.
Stale data with an error is still a meaningful state
Tell the user that the displayed data may be out of date and offer retry.
This is often more useful than replacing known content with an empty error page.
Mutation followed by revalidation
flowchart LR
Save["Mutation 200 OK"] --> Inval["Invalidate Related Query Keys"]
Inval --> Refetch["Background Refetch of Canonical Server Truth"]Revalidation is safest when the server computes fields, permissions, totals, or relationships the client cannot reproduce.
Mutation response as a cache update
flowchart LR
Patch["PATCH /products/:id"] --> Resp["Response Payload Contains Canonical Product"]
Resp --> Direct["Update ['product', id] Cache Directly (Zero extra GET)"]This can avoid a request when the response is authoritative and the affected query shapes are known.
Still consider list ordering, filters, totals, and related cache entries.
Optimistic cache update
flowchart LR
Snap["1. Cache Snapshot"] --> Prov["2. Provisional UI"]
Prov --> Req["3. Async Request"]
Req --> Resolution["4. Confirm / Rollback / Revalidate"]The cache update should be isolated, reversible, and tested for failure and concurrency.
Server-state libraries provide mechanisms
They may offer:
- query keys;
- caching;
- stale time;
- retries;
- deduplication;
- invalidation;
- mutation lifecycles;
- optimistic updates.
They do not decide API semantics, domain ownership, or whether a mutation is safe to retry.
A useful separation of responsibilities
flowchart TD
Adapter["Transport Adapter: HTTP methods, headers, status codes"]
Query["Query Cache Layer: Freshness, deduplication, staleTime"]
Mapper["Domain Mapper: Transforms untrusted JSON into domain entities"]
Feature["Feature Coordinator: User intent, pagination, mutations"]
View["View Layer: Renders data, skeletons, and error recovery"]
Adapter --> Query --> Mapper --> Feature --> ViewEach layer should expose the information its caller needs without leaking every lower-level detail.
Do not hide every HTTP detail behind one giant service
A universal api.requestEverything() wrapper often obscures:
- status-specific recovery;
- cancellation;
- request identity;
- cache behavior;
- mutation semantics.
Centralize repeated mechanics, but keep meaningful operation contracts visible.
Transport errors versus domain errors
Different causes require different UI and retry behavior.
Normalize errors without leaking secrets
Normalize enough for the UI to make a decision.
Do not expose stack traces, tokens, internal SQL, or sensitive payloads to users.
Administrative catalogue architecture
The catalogue becomes easier to reason about when each concern has one owner.
Product query lifecycle
flowchart TD
URL["1. Parse URL & Query State"] --> Key["2. Build Deterministic Query Key"]
Key --> Cache["3. Read Fresh Cache or Trigger Fetch"]
Cache --> Validate["4. Validate Schema at Trust Boundary"]
Validate --> Expose["5. Expose Loading / Stale / Error / Data"]
Expose --> UI["6. Render Data & Recovery Actions"]The component should consume a query model, not reinvent this lifecycle.
Editing a product
flowchart TD
Load["1. Load Canonical Product into Server Cache"] --> Draft["2. Create Detached Local Edit Draft"]
Draft --> Validate["3. Validate Inputs Locally"]
Validate --> Submit["4. Submit Transport Command via Fetch"]
Submit --> ErrorBranch["5a. Failure: Preserve Draft & Show Alert"]
Submit --> SuccessBranch["5b. Success: Update/Invalidate Cache & Navigate"]Never let a failed mutation erase the user’s unfinished work by default.
Search and cancellation
The cache and UI must also ensure that an older result cannot win after a newer query.
Pagination and cache retention
Decide:
- whether previous pages remain visible;
- whether the next page is prefetched;
- how long old pages remain cached;
- how a mutation changes page membership;
- what happens when filters change.
Pagination is a user experience and cache policy together.
Prefetching remote data
Prefetch when:
- the next destination is predictable;
- the request is safe;
- the cost is bounded;
- the data is likely to be used.
Do not prefetch every possible route and overwhelm the network or cache.
A client cache is not a database
Cache entries can:
- expire;
- be evicted;
- be incomplete;
- fail to persist;
- disagree with the server;
- belong to one user or permission context.
Never use a cache as the sole authority for durable business data.
Server state and offline state differ
Offline support requires decisions about:
- local persistence;
- queued mutations;
- conflict resolution;
- retry scheduling;
- user visibility;
- security and data expiry.
A query cache alone does not create an offline architecture.
Practical lab: Cached Administrative API Client
Build a small API client that separates HTTP caching, application/query caching, runtime validation, and user-visible failure states.
The practical makes request identity, freshness, invalidation, mutation behavior, and partial failure observable.
Practical stages 1–4: establish the boundary
- Implement a
fetch()wrapper that checksresponse.okexplicitly. - Validate response data at the boundary.
- Model remote UI states.
- Add search and cancellation.
Simulate 401, 403, 404, 422, 500, timeout, and malformed-data responses.
Practical stages 5–8: query cache fundamentals
- Build a simple query cache.
- Add freshness.
- Deduplicate requests.
- Add pagination.
Verify that query keys include every relevant input and that stale data remains distinguishable from fresh data.
Practical stages 9–13: mutations and cache truth
- Add a mutation.
- Handle server validation.
- Invalidate after save.
- Directly update one cache entry.
- Add an optimistic toggle with rollback.
Document why each mutation uses pessimistic, direct-update, or optimistic behavior.
Practical stages 14–17: resilience and architecture
- Simulate partial dashboard failure.
- Compare browser cache and query cache.
- Add native
FormDatasubmission. - Draw the full data architecture.
Verification: loading, empty, stale, error, recovery, and partial-success states are visible.
Practical extension: conditional requests
Add ETags and conditional requests:
Document which layer owns:
- the HTTP cache;
- the query cache;
- freshness;
- invalidation;
- the final domain model.
Try this yourself
Design the cache keys for:
Then list which mutations invalidate or directly update each key.
Troubleshooting guide (Part 1)
| Symptom | Likely cause |
|---|---|
| 500 response enters success code | response.ok was not checked |
| Old search result replaces new result | missing cancellation or request identity |
| Different filters show the same data | cache key omits an input |
| Every refresh blanks the screen | stale data is discarded unnecessarily |
Troubleshooting guide (Part 2)
| Symptom | Likely cause |
|---|---|
| Failed save loses the draft | mutation lifecycle owns the form incorrectly |
| Retry duplicates an operation | non-idempotent request has no safety policy |
| One panel breaks the dashboard | independent queries were coupled with Promise.all() |
| Cache never updates after save | invalidation or direct update is undefined |
Completion checklist
- HTTP errors and network errors are distinct;
- responses are validated before entering domain code;
- request cancellation prevents stale work;
- retries are limited and operation-aware;
- remote UI states distinguish loading, empty, stale, and error;
- cache keys include all relevant inputs;
- freshness and invalidation policies are explicit;
- mutations preserve drafts on failure;
- optimistic updates have rollback and revalidation rules;
- partial failures preserve independent useful data.
Misconceptions to leave behind (Part 1)
| Misconception | Better mental model |
|---|---|
fetch() rejects for every 4xx or 5xx | Check response.ok explicitly |
| TypeScript proves the server response | Runtime validation proves delivered data |
| Every request should be retried | Retry only safe, useful operations |
| Loading and empty are the same | One is pending; one is a successful zero result |
| Cache means correct forever | Cache means reusable knowledge under a policy |
Misconceptions to leave behind (Part 2)
| Misconception | Better mental model |
|---|---|
| Stale means unusable | Stale data may be useful while revalidating |
| Every mutation should be optimistic | Choose based on reversibility and conflict risk |
| Browser cache and query cache are identical | They are different layers with different owners |
| A server-state library replaces API design | It provides mechanisms, not semantics |
| Successful save fixes every cache | Related entries still need reconciliation |
The chapter in one sentence
Treat remote data as a lifecycle with explicit request identity, validation, freshness, failure, mutation, and recovery policies.
Next: Chapter 10
The next chapter will build on client-server architecture with:
- accessibility and inclusive interaction;
- semantic structure and assistive technology;
- keyboard, focus, and form behavior;
- robust component contracts;
- testing the experience rather than only the implementation.
Questions
For one request in your application, can you explain its key, freshness policy, cancellation rule, retry policy, failure states, and invalidation path?
10 - Real-Time Communication, Offline Systems & Client Persistence
Real-Time Communication, Offline Systems & Client Persistence
Design for disconnection
Chapter 10
Polla Fattah
Today’s goal
Build front-end systems that remain understandable when messages arrive, networks disappear, and local work must synchronize later.
We will connect:
- polling, long polling, SSE, WebSockets, and WebRTC;
- connection state, ordering, duplication, backpressure, and reconnection;
- cookies, browser storage, IndexedDB, and Cache Storage;
- service-worker lifecycle and caching strategies;
- offline reads, offline writes, outboxes, and synchronization;
- identity, conflicts, eventual consistency, and background sync;
- a field-inspection architecture that works without a permanent connection.
By the end of today you can
- choose the simplest communication model that satisfies the requirement;
- prevent overlapping polls and stale live updates;
- design SSE and WebSocket message envelopes;
- model connection state as user-visible state;
- distinguish transport delivery from application semantics;
- choose browser storage by data meaning and lifetime;
- explain service-worker install, activate, and fetch interception;
- design cache-first, network-first, and stale-while-revalidate policies;
- persist drafts and queued operations offline;
- synchronize with idempotency, retry limits, and conflict detection.
The central lesson
Real-time and offline behavior are not switches; they are explicit policies for communication, persistence, synchronization, identity, and recovery.
The network may be delayed, duplicated, reordered, unavailable, or partially available.
The application must make those conditions meaningful rather than pretending they cannot happen.
The chapter’s progression
flowchart TD
A["Communication Requirement"] --> B["Select Simplest Transport"]
B --> C["Model Connection Lifecycle"]
C --> D["Local Persistence Strategy"]
D --> E["Service Worker & Cache Storage"]
E --> F["Offline Reads & Writes"]
F --> G["Outbox Synchronization"]
G --> H["Conflict Resolution & Recovery"]Complexity should be earned by a real product requirement.
“Real-time” is a requirement, not one technology
Ask:
- how fresh must the data be?
- who sends updates?
- is communication one-way or two-way?
- can a short delay be accepted?
- must work continue without a network?
- are messages durable or disposable?
The answers narrow the transport and persistence choices.
Start with the simplest communication model
flowchart LR
A["Manual Refresh"] --> B["Short Polling"]
B --> C["Long Polling"]
C --> D["Server-Sent Events (SSE)"]
D --> E["WebSockets / WebRTC"]Use the least complex model that satisfies freshness and interaction needs.
Do not add a bidirectional socket when periodic reads are sufficient.
Polling is repeated HTTP
Polling is easy to deploy, observe, authorize, and cache.
Its cost is repeated requests even when nothing changed and delayed delivery between intervals.
Polling interval is a trade-off
Choose based on the business meaning of freshness, not an arbitrary “real-time” label.
Poll only when useful
Pause or reduce polling when:
- the document is hidden;
- the user leaves the relevant route;
- the device is offline;
- the data is not visible or actionable;
- a long-lived connection already supplies updates.
Resume with an explicit refresh when the user returns.
Avoid overlapping poll requests
If a request takes longer than the interval, a naive timer can create concurrent requests and out-of-order responses.
Long polling is still repeated HTTP
sequenceDiagram
autonumber
participant Client
participant Server
Client->>Server: HTTP GET /events (hangs open)
Note over Server: Server delays response until event occurs
Server-->>Client: 200 OK (Event payload)
Note over Client: Client processes event
Client->>Server: HTTP GET /events (immediately reopens)Long polling reduces empty responses while preserving an HTTP-shaped deployment model.
It still needs cancellation, timeout, reconnection, and duplicate handling.
Server-Sent Events are one-way streams
SSE is useful when the client sends commands through ordinary HTTP and the server streams updates back.
SSE architecture
flowchart LR
subgraph Client["Browser Client"]
Cmd["Mutation Action"]
Listener["EventSource Listener"]
end
subgraph Server["Server API"]
HTTP["POST /api/commands"]
Stream["GET /api/events (text/event-stream)"]
end
Cmd -->|Standard HTTP POST| HTTP
Stream -->|Unidirectional Stream| ListenerThe one-way direction simplifies some authorization and infrastructure concerns compared with a fully bidirectional socket.
When SSE fits well
Use SSE for:
- notifications;
- progress updates;
- monitoring dashboards;
- assignment changes;
- server-generated status events.
It is a poor fit when the client must exchange frequent messages in both directions over one connection.
Commands over HTTP, events over SSE
sequenceDiagram
autonumber
participant App as Browser Client
participant API as Municipal REST API
participant SSE as SSE Stream Server
App->>API: POST /assignments/a-1/complete (HTTP)
API-->>App: 200 OK { id: "a-1", status: "completed" }
API->>SSE: Broadcast domain event
SSE-->>App: event: assignment.updated
data: { id: "a-1", status: "completed" }Keeping commands and events separate can make authorization, retries, and audit behavior clearer.
The event remains an announcement; the server remains authoritative.
Named SSE events improve intent
Named events avoid one generic handler having to infer every message kind from an unstructured payload.
SSE reconnect is part of the contract
An interrupted stream should define:
- reconnect delay;
- maximum or bounded backoff;
- authentication refresh;
- missed-event recovery;
- duplicate-event handling;
- user-visible connection status.
Reopening the stream alone does not guarantee that no update was missed.
Events need identity when necessary
An event ID or entity version lets the client detect duplicates, gaps, and stale updates.
WebSockets are bidirectional transports
They fit interactive collaboration, presence, chat, live controls, and high-frequency two-way communication.
They also create more lifecycle and protocol responsibility.
WebSocket architecture
stateDiagram-v2
[*] --> Connecting: new WebSocket(url)
Connecting --> Authenticating: onopen
Authenticating --> Subscribed: auth token accepted
Subscribed --> Active: bidirectional framing
Active --> Active: message / ping / pong
Active --> Reconnecting: onclose / onerror
Reconnecting --> Connecting: exponential backoff
Active --> Closed: user disconnect / logout
Closed --> [*]The socket is one part of the application architecture, not the architecture itself.
A WebSocket is a transport, not an application protocol
The transport does not decide:
- message types;
- authentication refresh;
- ordering guarantees;
- idempotency;
- authorization;
- persistence;
- conflict resolution;
- missed-event recovery.
Those rules belong in the message protocol and domain design.
Design message envelopes deliberately
An envelope gives the client enough information to route, validate, order, and observe messages.
Runtime validation still applies
TypeScript cannot guarantee that a remote peer sent the expected data.
Connection state is UI state
Users need to know whether an action is live, queued, delayed, or unavailable.
Reconnection is not “just reopen the socket”
After reconnecting, the client may need to:
- refresh credentials;
- resubscribe;
- request a snapshot;
- replay safe commands;
- detect missed versions;
- discard obsolete local assumptions.
Reconnect is a synchronization sequence.
Snapshot plus events is a robust pattern
flowchart TD
A["1. GET /api/snapshot
(Baseline state at t0)"] --> B["2. Connect Event Stream / WebSocket
(Subscribe to mutations)"]
B --> C["3. Buffer In-Flight Events
(Queue events arriving during fetch)"]
C --> D["4. Apply Ordered Updates
(Discard events older than snapshot version)"]The snapshot provides a baseline.
Events provide changes after that baseline.
The protocol must define the race between reading the snapshot and subscribing.
Ordering cannot be assumed
Messages can arrive:
- late;
- out of order;
- duplicated;
- after a reconnect;
- from an earlier connection.
Use sequence numbers, versions, timestamps with care, or a server reconciliation step.
Duplicates are normal in reliable systems
At-least-once delivery can repeat a message.
Make handlers idempotent where possible:
flowchart TD
Ev["Incoming Live Event (version, eventId)"] --> Check{"Version > Current Local Version?"}
Check -->|No: Already applied or obsolete| Ignore["Drop / Acknowledge Safely (Idempotent)"]
Check -->|Yes: Exact next version| Apply["Apply update to UI state & cache"]
Check -->|Gap detected: Version > Current + 1| Resync["Buffer event & trigger snapshot reconciliation"]Exactly-once behavior is usually an application-level illusion built from IDs and state.
Backpressure protects the client
If messages arrive faster than the UI or storage can process them, define a policy:
- coalesce updates;
- drop obsolete intermediate states;
- pause subscriptions;
- apply batches;
- request a fresh snapshot;
- show degraded status.
Unbounded queues turn a temporary burst into a memory and responsiveness problem.
WebSocket reconnect needs bounded backoff
flowchart TD
Drop["Socket Disconnected"] --> Wait1["Wait Base Delay (e.g. 500ms + Jitter)"]
Wait1 --> Try1["Attempt Reconnect"]
Try1 -->|Failure| Wait2["Wait Exponential Delay (1000ms + Jitter)"]
Wait2 --> Try2["Attempt Reconnect"]
Try2 -->|Failure| Max["Cap at Max Delay (e.g. 10s) & Surface Disconnected Banner"]Add jitter and stop retrying when the failure is permanent, such as invalid credentials or forbidden access.
Polling, SSE, or WebSocket?
| Requirement | Good starting point |
|---|---|
| occasional freshness | polling |
| server-to-client stream | SSE |
| high-frequency two-way interaction | WebSocket |
| intermittent network with durable work | HTTP plus local outbox |
| peer media/data | WebRTC |
The product requirement should choose the transport.
WebRTC is a different kind of real-time
WebRTC is designed for peer media and data.
It still needs:
- signaling;
- identity and authorization;
- NAT traversal;
- relay infrastructure;
- connection lifecycle;
- application-level message rules.
It is not a replacement for ordinary server events.
Signaling establishes a connection
sequenceDiagram
autonumber
participant PeerA as Peer A
participant Sig as Signaling Server (HTTP/WS)
participant PeerB as Peer B
PeerA->>Sig: Send SDP Offer + ICE Candidates
Sig->>PeerB: Forward Offer
PeerB->>Sig: Send SDP Answer + ICE Candidates
Sig->>PeerA: Forward Answer
Note over PeerA,PeerB: Direct P2P Media / DataChannel Established
PeerA<<-->>PeerB: Direct P2P DataChannel / MediaThe signaling channel helps peers exchange connection information.
The actual media or data path may then be direct or relayed.
NAT traversal and relay
Some network environments prevent direct peer connectivity.
STUN can help discover a reachable address.
TURN can relay traffic when direct connection fails.
Real-time architecture must budget for the cases where the ideal path is unavailable.
Client persistence starts with meaning
Ask what must survive:
- reload;
- tab close;
- browser restart;
- route change;
- offline time;
- service-worker update.
Choose storage after defining lifetime, size, sensitivity, and access pattern.
Cookies
Cookies are sent with requests according to domain, path, and policy rules.
They are useful for server-managed sessions but require careful security attributes:
Do not use cookies as a general client database.
localStorage
It is convenient for small string preferences and simple drafts.
It is synchronous, string-only, quota-limited, and not a secure store.
localStorage is synchronous
Large reads and writes can block the main thread.
Avoid using it for large datasets, frequent updates, or high-volume logs.
For larger asynchronous data, consider IndexedDB or an application-specific persistence layer.
localStorage stores strings
Validate versions and shape on read. Data written by the application is still an external runtime boundary after reload.
sessionStorage has a shorter lifetime
It is scoped to a browser tab session.
It can suit temporary per-tab state that should survive reload but not a new tab or later session.
The same string-only and synchronous limitations still apply.
Storage events coordinate tabs
Storage events can notify other documents, but they are not a durable event bus or a complete synchronization protocol.
IndexedDB is an asynchronous local database
Use it for:
- larger structured data;
- offline records;
- drafts and outboxes;
- indexes and queries;
- data that should not block rendering.
The asynchronous API adds complexity, but supports a more appropriate data model.
IndexedDB is transactional
flowchart LR
Tx["db.transaction(['inspections', 'outbox'], 'readwrite')"] --> Ops["Execute Reads & Writes"]
Ops --> Success["All operations succeed
→ Automatic Commit"]
Ops --> Fail["Any error thrown
→ Automatic Complete Abort (Rollback)"]Group related updates so a draft and its outbox record cannot silently diverge.
Design transaction boundaries around invariants the application must preserve.
IndexedDB is asynchronous
Plan for:
- request errors;
- aborted transactions;
- blocked upgrades;
- unavailable storage;
- concurrent tabs;
- schema migration.
The database is local, but it is still a failure-prone boundary.
IndexedDB versioning is a migration contract
flowchart TD
Open["indexedDB.open('MunicipalApp', 2)"] --> Check{"Requested Version > Current DB Version?"}
Check -->|No| Ready["onsuccess: Database ready for transactions"]
Check -->|Yes| Upgrade["onupgradeneeded: Run migrations
- createObjectStore()
- createIndex()
- transform existing records"]
Upgrade --> ReadyTest upgrades from realistic previous versions.
Do not assume users start with an empty database after a release.
Cache Storage stores request/response pairs
It fits resources retrieved through the Fetch model, especially assets and selected HTTP responses.
It is not interchangeable with a domain database.
Cache Storage versus IndexedDB
| Cache Storage | IndexedDB |
|---|---|
| request/response pairs | structured application records |
| asset and HTTP-like retrieval | queries, indexes, transactions |
| service-worker friendly | domain/offline data friendly |
| cache strategy | persistence and synchronization model |
Choose based on data semantics, not only size.
Storage is not guaranteed forever
Data may be evicted because of:
- quota pressure;
- user settings;
- browser policy;
- private browsing;
- device storage constraints;
- application cleanup.
Offline architecture must tolerate missing local data and rehydrate from the server when possible.
Storage quotas affect product behavior
Large offline datasets need:
- size limits;
- eviction policy;
- user-visible storage status;
- cleanup rules;
- recovery when writes fail.
Do not promise indefinite offline history unless the platform and product support that promise.
Choose storage by data semantics
flowchart TD
subgraph BrowserStorageTaxonomy["Browser Storage by Purpose & Scope"]
S1["Session Identity → Cookies (HttpOnly, Secure)"]
S2["Small User Preferences → localStorage (<5MB, sync API)"]
S3["Structured Offline Data → IndexedDB (Async, indexed, large quota)"]
S4["HTTP Asset Responses → Cache Storage API (Request/Response pairs)"]
S5["Durable Queued Operations → IndexedDB Outbox (Transactional)"]
endOne application can use several stores with explicit ownership.
Service workers run outside the page
flowchart LR
Page["Active Browser Window / Page"] <-->|fetch() / Navigation| SW["Service Worker
(self.addEventListener('fetch'))"]
SW <-->|Cache Match / Put| Cache["Cache Storage API"]
SW <-->|Network Request| Net["Remote Network Server"]The worker has a different lifecycle and execution context.
It can intercept fetches and support caching, but it does not automatically understand domain data or user intent.
Secure context is required
Service-worker features generally require a secure context such as HTTPS, with localhost development commonly treated specially.
Deployment configuration is part of the offline architecture.
Service-worker lifecycle
stateDiagram-v2
[*] --> Installing: navigator.serviceWorker.register()
Installing --> Waiting: self.skipWaiting() / new SW downloaded
Waiting --> Activating: Old SW clients closed / skipWaiting
Activating --> Active: clients.claim()
Active --> Active: Intercepting network requests
Active --> Redundant: Replaced by updated script
Redundant --> [*]An updated worker may not control the current page immediately.
Design update messaging and cache compatibility rather than assuming an instant replacement.
Installation prepares resources
Installation should be bounded and versioned.
Do not make one missing optional asset prevent the entire application from installing unless that is intentional.
Activation cleans up and takes ownership
Activation is a migration boundary for caches and control behavior.
Coordinate old and new asset versions so the page does not combine incompatible files.
Fetch interception is policy code
The worker must decide which requests it can safely handle and which should pass through.
Do not cache credentials, private data, or mutations accidentally.
A service worker does not mean offline automatically
Offline support requires:
- a cache strategy;
- data persistence;
- a fallback UI;
- offline write behavior;
- synchronization;
- conflict policy;
- update and eviction handling.
Registration is only the start.
Cache-first strategy
flowchart TD
Req["Incoming fetch(event.request)"] --> Cache{"Cache.match(request)?"}
Cache -->|Hit| Fast["Return cached Response (Instant)"]
Cache -->|Miss| Net["Fetch from Network"]
Net --> Put["cache.put(request, clone)"]
Put --> Res["Return fresh Response"]Good for versioned static assets or content where immediate availability is more important than freshness.
Risk: stale content can persist if versioning and invalidation are weak.
Network-first strategy
flowchart TD
Req["Incoming fetch(event.request)"] --> Net{"Network fetch()"}
Net -->|Success| Put["cache.put(request, clone)"]
Put --> Res["Return fresh server response"]
Net -->|Failure / Offline| Cache{"Cache.match(request)?"}
Cache -->|Hit| Stale["Return cached offline fallback"]
Cache -->|Miss| Err["Return custom offline error page"]Good for data that should be fresh when connectivity exists but remain readable offline.
Risk: slow networks can delay the fallback unless timeouts are defined.
Stale-while-revalidate
flowchart TD
Req["fetch(event.request)"] --> Cache{"Cache.match(request)?"}
Cache -->|Hit| Ret["Return cached response immediately"]
Cache -->|Miss| WaitNet["Await network response"]
Ret --> BG["Async background fetch()"]
BG --> Put["Update Cache Storage for next load"]
WaitNet --> PutGood when fast display and eventual freshness are both useful.
The UI should communicate meaningful staleness.
Network-only and cache-only
flowchart LR
subgraph NetworkOnly["Network-Only Strategy"]
R1["fetch(request)"] --> N1["Server API"]
N1 -->|Never Cache| UI1["Critical Mutations / Auth"]
end
subgraph CacheOnly["Cache-Only Strategy"]
R2["fetch(request)"] --> C1["Cache Storage API"]
C1 -->|Zero Network| UI2["Pre-cached Static Assets / Offline Fonts"]
endDo not apply one strategy to every request.
Mutation requests should not be silently treated like cacheable reads.
Strategy matrix
| Data | Useful strategy |
|---|---|
| versioned assets | cache-first |
| current account data | network-first / query policy |
| readable offline catalogue | stale-while-revalidate |
| mutation command | network-only plus outbox when offline |
| private durable draft | IndexedDB, not public cache |
The right choice depends on authority, freshness, and recovery.
Offline fallback is a product surface
An offline screen should explain:
- what remains available;
- which work is saved locally;
- what is waiting to sync;
- what cannot be done now;
- how the user can retry.
Offline is a mode with capabilities, not simply an error page.
Offline-capable versus offline-first
Offline-first is a larger product commitment involving conflict, identity, local storage, and synchronization.
Do not adopt it for a feature that only needs cached reading.
navigator.onLine is only a hint
It can indicate a network interface state.
It does not prove:
- the server is reachable;
- authentication works;
- the request will succeed;
- the route is available;
- the network has useful bandwidth.
Treat actual requests and failures as stronger evidence.
Offline reads
flowchart LR
IDB["IndexedDB Store"] --> Read["Read Cached Record"]
Read --> Valid["Check freshness & validity"]
Valid --> Render["Render view with Offline/Stale indicator"]Offline reads should identify:
- when data was last synchronized;
- whether it is incomplete;
- whether an update is pending;
- when the next refresh can occur.
Offline writes need durable intent
flowchart LR
User["User Submits Form"] --> Split["Atomic Transaction"]
Split --> Rec["Update Local Record"]
Split --> Box["Enqueue Outbox Operation"]
Box --> Sync["Background Sync Engine"]Do not keep the only copy of a user’s work in memory while waiting for the network.
Persist the draft or operation before telling the user it is safely queued.
The outbox pattern
flowchart TD
Cmd["User Action: Submit Inspection"] --> Tx["Atomic IndexedDB Transaction"]
Tx -->|Write| Domain["Write local 'inspections' store (Status: PendingSync)"]
Tx -->|Write| Outbox["Write 'outbox' store (Operation: CREATE_INSPECTION)"]
Outbox --> Sync{"Sync Trigger
(Online event, page load, visibility)"}
Sync --> Post["POST /api/inspections with Idempotency-Key"]
Post -->|200 Ack| Done["Remove outbox record, update local status to Synced"]
Post -->|Network Drop| Retry["Increment attempt count, schedule backoff retry"]
Post -->|4xx Fatal| Dead["Mark outbox item FAILED, notify inspector"]The transaction keeps local work and its synchronization intent together.
Local identity versus server identity
An offline record may need:
classDiagram
class InspectionRecord {
+UUID localId "Generated client-side immediately (crypto.randomUUID())"
+String serverId "Canonical ID assigned or confirmed by server (null while offline)"
+UUID operationId "Unique idempotency token sent in outbox request"
+String syncStatus "draft | pending_sync | syncing | synced | conflict"
+Number version "Optimistic concurrency version tag"
}Do not use one identifier for three different meanings.
Pending synchronization is user-relevant state
Show whether work is:
stateDiagram-v2
[*] --> Draft: User edits record
Draft --> PendingSync: User commits inspection
PendingSync --> Syncing: Network available & outbox flush
Syncing --> Synced: Server returns 200 OK
Syncing --> PendingSync: Transient 5xx / timeout (retry)
Syncing --> Failed: 422 validation / 403 forbidden
Failed --> Draft: User edits data to fix validation
Synced --> [*]Users need confidence about whether their work is safe, not only whether the browser currently has a connection.
Sync is not the same as retry
Retry repeats an operation after a transient failure.
Synchronization reconciles local intent with remote authority after time and possibly other changes.
It may need identity, ordering, conflict detection, and a new domain decision.
Conflict example
When device A reconnects, blindly overwriting the server may destroy a newer decision.
The system needs a conflict policy.
Conflict policies
flowchart TD
subgraph ConflictResolution["Conflict Resolution Policies"]
C1["Last Write Wins (LWW)
Clock timestamp determines winner (Risky)"]
C2["Server Wins
Canonical authority; client state overwritten"]
C3["Client Wins
Local user decision always takes precedence"]
C4["Field-Level 3-Way Merge
Combine non-overlapping field edits"]
C5["Manual User Resolution
Side-by-side visual diff prompt"]
endChoose based on data meaning and harm, not implementation convenience.
Version-based conflict detection
The client submits the version it edited.
The server rejects or resolves the operation when its current version differs.
Operation logs make synchronization inspectable (Part 1)
| Field | Type | Description |
|---|---|---|
operationId | UUID | Unique idempotency key for this synchronization action |
entity | string | Target domain entity (e.g. 'inspection') |
command | string | Action type (e.g. 'SUBMIT_REPORT') |
payload | JSON | Complete serialized mutation payload |
Operation logs make synchronization inspectable (Part 2)
| Field | Type | Description |
|---|---|---|
createdAt | timestamp | Client timestamp when user performed action |
attempts | number | Retry counter with maximum threshold |
status | enum | 'pending_sync' | 'syncing' | 'failed' |
Eventual consistency is a user experience
After an offline action, local state may say “completed” while the server has not confirmed it.
Represent that distinction:
Do not present unconfirmed local intent as permanent server truth.
Background Sync is an enhancement
Background Sync can help flush queued work when the browser decides conditions are suitable.
The application must still work when it is unavailable, delayed, or denied.
Provide a visible manual retry and foreground synchronization path.
Periodic background sync needs restraint
Periodic work affects:
- battery;
- data usage;
- privacy;
- server load;
- freshness expectations.
Use it only when the product meaningfully benefits from background refresh.
Service-worker update risk
An old page and a new worker can temporarily coexist.
Plan for:
- compatible cache names;
- atomic asset versions;
- schema migration;
- user messaging;
- safe activation timing.
Updating a worker is a deployment and data-migration concern.
Cache versioning
Version names make cleanup and incompatible asset replacement explicit.
Do not let old bundles and new runtime assumptions share one unbounded cache.
Do not cache every API response forever
For each response, define:
- sensitivity;
- freshness;
- user scope;
- size;
- invalidation;
- offline usefulness;
- eviction.
Caching private data without a lifecycle can create security and correctness problems.
User identity and offline data
When a user signs out or changes account, decide what happens to local data:
- delete it;
- encrypt or isolate it;
- keep only non-sensitive preferences;
- mark it for a specific identity;
- prevent another account from reading it.
Local persistence must respect authorization boundaries.
Cache API is not authorization
A cached response existing on the device does not prove the current user may still access it.
Authorization remains a server and application decision.
Clear or revalidate private cached data when identity or permissions change.
Progressive web apps are a capability set
A PWA may include:
- installability;
- service workers;
- offline assets;
- local data;
- push notifications;
- background synchronization.
Installability does not automatically imply offline data or reliable synchronization.
The app shell is only one layer
flowchart TD
subgraph OfflineFieldSystem["The Four Pillars of Offline Resilience"]
P1["1. App Shell (Cache Storage)
HTML, CSS, JS bundles cached for 0ms offline boot"]
P2["2. Local Data Store (IndexedDB)
Read-only municipal permits, checklists, inspector profiles"]
P3["3. Durable Outbox (IndexedDB)
Queued operations preserved across reboots & browser closes"]
P4["4. Synchronization Engine
Background queue consumer with exponential backoff & idempotency"]
endCaching HTML, CSS, and JavaScript does not solve domain data or mutation conflicts.
Offline-first interaction
This can produce excellent responsiveness, but it requires durable local models, conflict handling, and clear confirmation states.
Network-first interaction
This is often appropriate for authoritative records where stale edits are risky and connectivity is usually available.
Select the offline scope
Possible scopes:
flowchart TD
L1["Level 1: Offline Shell Only
App boots to static frame with 'No Connection' banner"]
L2["Level 2: Offline Read Cache
User can browse previously viewed permits and checklists"]
L3["Level 3: Offline Drafts
Unsaved form inputs persist across reboots in IndexedDB"]
L4["Level 4: Offline Queued Outbox
Inspectors complete inspections offline; synced upon reconnect"]
L5["Level 5: Full Local-First Workflow
CRDTs / multi-device peer synchronization with zero central locks"]
L1 --> L2 --> L3 --> L4 --> L5Choose the smallest scope that solves the user problem.
Real-time plus offline together
stateDiagram-v2
state Online {
[*] --> Streaming
Streaming: WebSocket / SSE live events update local store
}
state Offline {
[*] --> Autonomous
Autonomous: Reads from IndexedDB; writes queued in Outbox
}
state Reconnecting {
[*] --> Heartbeat
Heartbeat: Egress probe succeeds
Heartbeat --> FlushOutbox: Send queued operations with Idempotency-Key
FlushOutbox --> InvalidateQueries: Refresh canonical server state
}
Online --> Offline: Connection lost
Offline --> Reconnecting: Network regained
Reconnecting --> Online: All operations settledOne local data model can be the bridge between live updates and offline work.
A reconnect sequence
flowchart TD
A["1. Connectivity Hint (online event / window focus)"] --> B["2. Heartbeat Ping (verify real internet egress)"]
B --> C["3. Validate Auth Token / Refresh Session"]
C --> D["4. Fetch Server Version Vector / Changes"]
D --> E["5. Flush Pending Outbox Operations with Idempotency"]
E --> F["6. Detect & Resolve Concurrent Conflicts"]
F --> G["7. Reconcile UI & Invalidate Fresh Server Queries"]The order matters. Sending stale operations before understanding current server state can create avoidable conflicts.
Avoid double-applying events
Use operation IDs or server versions to recognize that the local optimistic change and remote event refer to one operation.
Data freshness after reconnect
After a long offline period, cached data may be:
- outdated;
- structurally migrated;
- revoked by permissions;
- superseded by a conflict;
- incomplete.
Revalidate and communicate the result rather than silently labeling it current.
Practical project: Offline-Capable Field Inspections
Build a field-inspection application that can read assignments, save drafts, queue completed inspections, and recover when connectivity returns.
The practical combines live communication, persistence, service workers, outbox sync, idempotency, and conflict decisions.
Practical stages 1–3: communication models
- Implement polling.
- Prevent poll overlap.
- Replace polling with SSE.
Measure freshness, request count, cancellation, and behavior when the tab is hidden or offline.
Practical stages 4–7: live connection design
- Compare SSE and WebSocket.
- Add connection state.
- Design reconnection.
- Explore WebRTC conceptually.
Write down which message guarantees the application actually needs: ordering, identity, duplication handling, and recovery.
Practical stages 8–10: persistence foundations
- Add local preferences.
- Add IndexedDB.
- Register a service worker.
Test reload, tab close, unavailable storage, upgrade, and a missing or malformed local record.
Practical stages 11–13: caching and offline reads
- Implement cache-first assets.
- Implement network-first data.
- Add an offline fallback.
Make the strategy visible in the UI and document which requests are safe to cache.
Practical stages 14–17: offline writes and sync
- Build an offline draft.
- Create the outbox.
- Synchronize on reconnect.
- Add conflict detection.
Verification: records survive reload, duplicate submission is prevented or detected, and permanent failures remain recoverable.
Practical stages 18–20: operational hardening
- Treat Background Sync as an enhancement.
- Version the service-worker cache.
- Draw the full architecture.
The app must remain useful without Background Sync and must explain what is waiting to synchronize.
Practical extension: conflict resolution
Simulate a record changed on the server while the client was offline.
Compare:
Record which policy is safest for inspection evidence and why.
Try this yourself
Design the local records for one inspection:
Then define the transaction that persists the draft and the queued operation together.
Troubleshooting guide (Part 1)
| Symptom | Likely cause |
|---|---|
| Poll responses arrive out of order | Requests overlap without identity or cancellation |
| Reconnected socket misses updates | No snapshot or event-version recovery |
| Same event changes data twice | Handler lacks idempotency identity |
| Offline work disappears on reload | Only in-memory state was used |
| Service worker is registered but offline fails | No request strategy or local data model |
Troubleshooting guide (Part 2)
| Symptom | Likely cause |
|---|---|
| Old assets break with new worker | Cache versioning and activation are unsafe |
| Duplicate inspection is created | No idempotency key or server deduplication |
| Local data leaks across accounts | Persistence is not scoped or cleared on identity change |
| Sync retries forever | Permanent failures lack a terminal state |
Completion checklist
- the transport matches the actual freshness and direction requirements;
- polling and live connections have cancellation and reconnection policy;
- messages have validation and identity where needed;
- connection state is visible to users;
- storage is chosen by semantics and lifetime;
- service-worker caches are versioned and scoped;
- offline writes are durable before being acknowledged locally;
- outbox operations have IDs, retry limits, and terminal failures;
- synchronization detects duplicates, ordering issues, and conflicts;
- the app remains useful without optional background capabilities.
Misconceptions to leave behind (Part 1)
| Misconception | Better mental model |
|---|---|
| Real-time means WebSocket | Choose the simplest transport that fits |
| Polling is outdated | Polling is often clear and sufficient |
| SSE and WebSocket are the same | Direction and protocol responsibilities differ |
| Reopening a socket solves reconnect | Reconnect also needs resync and recovery |
| Messages arrive exactly once | Design for duplicates and reordering |
localStorage is a database | It is synchronous string storage |
Misconceptions to leave behind (Part 2)
| Misconception | Better mental model |
|---|---|
| App-written local data is trusted | Reloaded storage is a runtime boundary |
| Service-worker registration means offline | Strategies, persistence, and recovery are still needed |
navigator.onLine proves connectivity | It is only a hint |
| Retry and synchronization are identical | Sync reconciles local intent with remote state |
| Background Sync is guaranteed | It is an enhancement, not the foundation |
| Every application should be offline-first | Choose an offline scope from user need |
The chapter in one sentence
Design communication, persistence, and synchronization as explicit stateful systems that remain safe when connectivity is slow, absent, duplicated, or restored.
Next: Chapter 11
The next chapter will build on resilient client architecture with:
- security boundaries and threat modeling;
- authentication and authorization;
- browser security policies;
- safe handling of untrusted content;
- defensive application design.
Questions
When the network disappears during an important user action, what exactly is saved, what is queued, what is visible, and what happens when the server has changed meanwhile?
11 - Rendering Topologies: CSR, SSR, SSG & Beyond
Rendering Topologies: CSR, SSR, SSG & Beyond
Choose where the work happens
Chapter 11
Polla Fattah
Today’s goal
Understand rendering as a placement and timing decision rather than a framework label.
We will connect:
- client-side rendering, server-side rendering, and static generation;
- hydration, serialization, handoff, and browser work;
- partial hydration, islands, streaming, and hybrid routes;
- revalidation, edge rendering, server components, and resumability;
- cacheability, personalization, authentication, and failure modes;
- performance as server, build, network, and device cost;
- a route-specific rendering decision matrix.
By the end of today you can
- explain where rendering happens and when it occurs;
- compare CSR, SSR, and SSG without slogans;
- identify hydration mismatch causes and handoff costs;
- isolate interactive regions with islands or client boundaries;
- design streaming boundaries that preserve UX and accessibility;
- distinguish server components from server-rendered HTML;
- reason about revalidation, edge execution, and resumability;
- select a rendering strategy per route;
- keep secrets and server-only data on the server;
- measure the rendering cost triangle instead of optimizing one metric.
The central principle
Rendering architecture is the deliberate placement of work across build time, request time, and browser time according to freshness, personalization, interaction, cacheability, and device cost.
CSR, SSR, and SSG are tools in a continuum - not application-wide identities.
The four main questions
Ask:
- Where does rendering happen?
- When does it happen?
- What is transferred to the browser?
- What must execute in the browser?
The answers reveal the actual topology beneath a framework’s terminology.
Build time, request time, and browser time
flowchart LR
B["1. Build Time
(Static precomputation,
CI/CD compilation)"] --> R["2. Request Time
(On-demand server render,
data resolution per user)"]
R --> BR["3. Browser Time
(Client JavaScript execution,
hydration, reactivity)"]Every topology moves work among these three moments.
Rendering is work
Rendering may include:
- fetching data;
- transforming content;
- producing HTML;
- serializing state;
- parsing JavaScript;
- hydrating event handlers;
- committing DOM updates;
- running device-side interaction.
“Rendered” does not mean “free”. It means the cost moved somewhere.
Client-side rendering
flowchart TD
A["1. Browser receives empty HTML shell (<div id='root'>) & JS bundle"] --> B["2. Browser downloads & executes JavaScript"]
B --> C["3. Client fires fetch() to API server"]
C --> D["4. Client calculates virtual DOM / reactive graph"]
D --> E["5. DOM created and painted (Content visible at last)"]CSR moves much of the initial rendering work to the device.
CSR strengths
CSR can provide:
- rich interaction after startup;
- simple browser-side state ownership;
- app-like navigation;
- direct access to browser APIs;
- a consistent client runtime.
It is often appropriate for authenticated tools and highly interactive workspaces.
CSR costs
The browser may need to:
- download a larger JavaScript bundle;
- parse and execute before content appears;
- fetch data after startup;
- render on lower-powered devices;
- manage loading and error states before the first useful view.
The network and device become part of the initial critical path.
SPA and CSR are related, not identical
A site can use client rendering without being one large SPA.
An SPA can include server-rendered initial HTML.
Do not use the terms as interchangeable architecture decisions.
CSR and search engines
Search engines may execute JavaScript, but discoverability, timing, metadata, content quality, and operational behavior still matter.
Public content often benefits from HTML being available before client execution.
The correct choice depends on the content and delivery requirements.
Server-side rendering
flowchart TD
A["1. Browser requests URL (GET /permits/104)"] --> B["2. Server fetches database data & executes components to HTML"]
B --> C["3. Browser receives full HTML document (Instant FCP)"]
C --> D["4. Browser downloads client JavaScript bundle"]
D --> E["5. Client hydrates HTML: attaches listeners & reconciles state"]SSR moves initial rendering work to request time and improves the first document’s content availability.
SSR changes the critical path
The request may now wait for:
- server data;
- server rendering;
- serialization;
- network transfer;
- browser parsing;
- hydration.
SSR can improve first content while adding server latency and operational complexity.
SSR does not mean no JavaScript
If the page must respond to:
- clicks;
- typing;
- menus;
- client navigation;
- live state;
some browser code still needs to execute.
SSR determines how initial output is produced, not whether interaction exists.
SSR does not automatically improve every page
SSR can be a poor fit when:
- content is highly personalized and uncacheable;
- server data is slow;
- the page is mostly an internal interactive tool;
- hydration cost dominates;
- deployment cannot support reliable request rendering.
Measure the whole request-to-interaction path.
Time to first byte matters
A server-rendered page with a slow first byte may feel worse than a carefully designed static or client-rendered page.
Rendering location alone does not determine speed.
Static site generation
flowchart LR
A["Build Command
(Fetch data at compile)"] --> B["Generate Static HTML
(/docs/intro.html)"]
B --> C["Deploy to Edge CDN
(Globally distributed)"]
C --> D["Instant Edge Serve
(<20ms TTFB globally)"]SSG moves rendering cost to deployment or build time.
It works especially well for public content with predictable freshness and high cacheability.
SSG strengths
SSG can provide:
- fast delivery from a CDN;
- simple serving infrastructure;
- strong cacheability;
- stable public HTML;
- low request-time computation.
The trade-off is build-time freshness and build complexity.
Static generation moves cost to deployment
Large sites may pay through:
- long builds;
- many pages;
- content invalidation;
- preview workflows;
- deployment coordination.
“Static” is operationally simple at request time, not necessarily cheap everywhere.
Static does not mean non-interactive
A generated page can include search, menus, forms, and other browser behavior.
Static describes when initial output was produced, not the absence of JavaScript.
Static data can become stale
Define a freshness policy:
- rebuild on content change;
- revalidate periodically;
- regenerate on demand;
- add client refresh;
- show the publication time.
SSG requires a plan for change, not an assumption that published data never changes.
CSR, SSR, and SSG describe initial rendering
After initial delivery, all three may:
- fetch new data;
- navigate without a full document;
- run client components;
- update the DOM;
- synchronize with external systems.
Do not infer the entire runtime architecture from the first render label.
Hydration reuses existing HTML
flowchart LR
HTML["Server-Rendered HTML
(Passive DOM structure)"] & JS["Client JavaScript Bundle
(Component definitions & handlers)"] --> Hydrate["Hydration Step
- Walk DOM tree
- Attach event listeners
- Initialize reactive state"]
Hydrate --> Interactive["Fully Interactive Page"]Hydration is a handoff from server-produced markup to an interactive browser runtime.
It is not simply a second independent render with no cost.
Hydration mismatches are signals
A mismatch can come from:
- random values;
- current time or timezone;
- browser-only APIs;
- different data on server and client;
- unstable IDs;
- conditional rendering based on client detection.
The warning often indicates an unclear rendering boundary or nondeterministic initial model.
Browser-only APIs need a boundary
This cannot run during server rendering.
Options include:
- client-only component;
- post-hydration effect;
- server-safe default;
- explicit capability check outside initial render.
Deterministic initial rendering
The server and client should agree on the initial representation.
Avoid using during initial render:
- random values;
- current time without a shared value;
- locale-dependent formatting without a fixed locale;
- browser state unavailable to the server.
Make the handoff data explicit.
Hydration has CPU cost
The browser may need to:
- download JavaScript;
- parse and compile it;
- execute component code;
- attach listeners;
- reconstruct client state;
- process event boundaries.
HTML existing on screen does not mean the page is ready for interaction.
The server-to-client handoff
flowchart TD
A["Server fetches Database Entities"] --> B["renderToString(App) + JSON.stringify(data)"]
B --> C["Network Transfer (HTML payload + <script id='__DATA__'>)"]
C --> D["Browser parses HTML and JSON data"]
D --> E["Client framework reconstructs identical Component State"]
E --> F["Hydration completes (TTI reached)"]Design the handoff as a contract.
Transfer only what the browser needs and keep secrets server-side.
Serialization has constraints
Not every server-side value can cross the boundary safely:
- functions;
- database connections;
- secrets;
- file handles;
- cyclic structures;
- framework-specific objects;
- sensitive records.
Use explicit transport models rather than passing arbitrary server objects.
Secrets must stay server-side
flowchart TD
subgraph ServerZone["Secure Server Boundary"]
Sec["Database Credentials / API Secrets / Private Columns"]
Comp["Server Component / Data Loader"]
Sec --> Comp
end
Comp -->|Explicit Public DTO / Props ONLY| Wire["Network Wire (HTML + Client State)"]
Wire --> BrowserZone["Browser Client Runtime"]
Sec -.->|BLOCKED: Never pass credentials or unscrubbed DB rows| WireSSR does not make a secret safe if the secret is embedded in HTML or serialized props.
Rendering on the server is not the same as authorization.
Hydration is not free because HTML exists
Measure:
- time to first byte;
- first contentful paint;
- time to interactive;
- JavaScript transfer;
- hydration CPU;
- input delay;
- post-hydration data work.
The right topology minimizes the cost that matters for the route and its users.
Partial hydration
Partial hydration reduces browser work when most of the page is content or static structure.
It requires clear boundaries for state and events.
Islands architecture
flowchart TD
subgraph StaticDocument["Zero-JS Static HTML Page"]
H["Header & Layout HTML"]
Content["Article / Product Text HTML"]
Footer["Footer HTML"]
end
StaticDocument -.-> I1["Island 1: Navigation Menu
(client:load)"]
StaticDocument -.-> I2["Island 2: Search Autocomplete
(client:idle)"]
StaticDocument -.-> I3["Island 3: Interactive Comments
(client:visible)"]Each island can hydrate independently.
The default becomes HTML with isolated interactive regions rather than one full-page client runtime.
Islands change the default
Instead of asking “how do we render the whole app in the browser?” ask:
This can reduce JavaScript, but introduces coordination decisions when islands need shared state.
Island hydration timing
Possible policies:
Choose timing based on user value and interaction urgency.
Islands have boundaries
Cross-island communication may need:
- URL state;
- custom events;
- shared server state;
- browser storage;
- a client coordinator.
Do not recreate a hidden global application just to make every island know about every other island.
Streaming changes delivery timing
sequenceDiagram
autonumber
participant Browser
participant Server
Browser->>Server: GET /dashboard
Note over Server: Fast data ready immediately
Server-->>Browser: Flush HTTP headers + App Shell + Fast HTML
Note over Browser: Browser parses & renders App Shell (Fast FCP!)
Note over Server: Slow analytics query resolves after 400ms
Server-->>Browser: Streamed HTML chunk <template id='chunk-1'>
Server-->>Browser: Inline script swaps placeholder with chunk
Note over Browser: Analytics widget renders without page reloadStreaming can improve perceived progress and reduce waiting for the slowest dependency.
It does not automatically eliminate total server, network, or browser work.
Streaming boundaries are UX boundaries
Each boundary should define:
- what appears first;
- what loading state is shown;
- what remains interactive;
- what happens on failure;
- how layout remains stable.
The stream is part of the product experience, not only a transport optimization.
Suspense as a boundary concept
A suspense-like boundary lets the system reveal available content while a slower region resolves.
Use meaningful fallbacks rather than arbitrary spinners everywhere.
Streaming and accessibility
As content arrives progressively:
- preserve logical document order;
- do not move focus unexpectedly;
- announce important updates appropriately;
- keep headings and landmarks coherent;
- avoid making keyboard users wait for a needed control without feedback.
Delivery timing must preserve usable structure.
Streaming and layout stability
Reserve predictable space where possible.
Unexpected late content can:
- shift reading position;
- move a focused control;
- cause accidental clicks;
- increase cumulative layout shift.
Loading UI is a layout contract.
SSR and streaming can mix
Streaming is a delivery strategy layered onto server rendering.
It does not create a separate universal topology that replaces SSR or SSG.
Static generation and streaming can also mix
A statically generated shell can load or stream a dynamic region later through client or edge behavior.
The route may combine:
- prebuilt public content;
- request-time personalization;
- browser-side interactivity.
Hybrid is often a precise choice, not architectural inconsistency.
Hybrid rendering is route-level strategy
Choose per route and sometimes per region within a route.
Revalidation and incremental generation
Incremental strategies reduce full rebuild cost while preserving static delivery for many requests.
Incremental static regeneration
An incrementally generated page has:
- a cached representation;
- a freshness window or invalidation event;
- regeneration work;
- a fallback while regeneration occurs;
- failure behavior.
The important design is the stale and failure policy, not the product name of the mechanism.
Stale-while-revalidate rendering
sequenceDiagram
autonumber
actor User1
participant CDN as Edge CDN / Cache
participant Server as Origin Server / Static Generator
actor User2
User1->>CDN: GET /permits/104 (Cache Stale, past max-age)
CDN-->>User1: Return cached stale HTML instantly (0ms latency)
CDN->>Server: Trigger background regeneration
Server->>Server: Re-fetch database & re-render HTML
Server-->>CDN: Update cache with fresh HTML version
User2->>CDN: GET /permits/104 (5 seconds later)
CDN-->>User2: Return fresh pre-rendered HTMLThis works well for content where slightly stale output is acceptable and fast delivery matters.
Edge rendering
Edge execution moves request-time work closer to users or data paths.
Potential benefits:
- lower network distance;
- request-aware personalization;
- distributed cache integration.
It is still request-time server work with runtime and deployment constraints.
Edge runtime trade-offs
Edge environments may constrain:
- Node-specific APIs;
- native modules;
- filesystem access;
- connection lifetimes;
- debugging and observability;
- regional consistency.
“Edge” is not automatically faster or simpler.
Render near the data or near the user?
The best location depends on data source, cacheability, user distribution, and runtime capabilities.
Server components are not the same as SSR
A server component may participate in different rendering and navigation strategies.
Do not collapse execution location and HTML delivery into one concept.
Server components can run at different times
Depending on the framework and route, server-side component work may happen:
- at build time;
- at request time;
- during a client navigation that requests a new payload;
- during revalidation.
The route’s data and cache policy still determine the actual behavior.
Server components can reduce client JavaScript
They can keep data access and noninteractive rendering on the server.
Only interactive boundaries need client behavior.
This can reduce browser bundle responsibility, but the handoff and client boundaries still have costs.
Client components own browser behavior
Use a client boundary for:
- event handlers;
- local interactive state;
- browser APIs;
- effects and subscriptions;
- client-only libraries.
Keep the boundary as small as the interaction allows.
Server/client boundaries are architectural boundaries
flowchart TD
subgraph ServerComponent["Server Component (PermitDetail.server.tsx)"]
DB[("Direct SQL / Internal Microservice")] --> SC["Executes exclusively on Server
- Zero client bundle footprint
- Keeps SQL drivers & secrets on server"]
end
SC -->|Passes Serialized Props| CC["Client Component ('use client')
(PermitActionButtons.tsx)
- Handles onClick, hover, local UI state"]The boundary controls:
- what crosses the network;
- what JavaScript ships;
- where authorization occurs;
- which state is reconstructed in the browser.
Data can be loaded near server rendering
Loading data close to server-rendered components can avoid duplicating the initial request in the browser.
But make the client handoff explicit:
- is the data serialized?
- is it cached?
- does the client revalidate?
- can the browser receive only a projection?
The server component payload has a cost
The browser may receive a structured description of server/client boundaries and data needed for the client tree.
Measure:
- payload size;
- serialization work;
- parsing;
- client boundary setup;
- repeated requests during navigation.
Moving logic server-side does not make all transfer cost disappear.
Hydration still matters with server components
Client components still need browser execution and interaction setup.
Server components can reduce the amount that hydrates, but they do not remove the need to design client boundaries, identity, and state handoff.
Next.js as a case study
Next.js can combine:
- static generation;
- server rendering;
- route-level revalidation;
- server and client components;
- streaming;
- client navigation.
Treat these as mechanisms for route decisions, not as a substitute for the rendering mental model.
Nuxt as a case study
Nuxt can combine:
- universal rendering;
- static generation;
- route rules;
- server handlers;
- client hydration;
- hybrid deployment.
Again, understand where data and rendering work occur rather than memorizing framework labels.
Frameworks should not become the mental model
Ask framework-independent questions:
- when is the HTML produced?
- where is data loaded?
- what crosses to the browser?
- which region hydrates?
- how is freshness controlled?
- what happens on navigation and failure?
Then map the answers to framework configuration.
The “double data” oversimplification
SSR does not always mean the same data is fetched twice.
Possible designs include:
- server output only for static content;
- serialized data reused by the client;
- a client revalidation request by policy;
- server component payloads with selective transfer.
Inspect the actual network and handoff path.
Streaming does not eliminate handoff cost
Even if HTML arrives progressively, the browser may still need to:
- parse client JavaScript;
- hydrate interactive regions;
- receive additional data;
- attach event behavior.
Streaming improves timing of availability; it does not erase client responsibility.
Resumability changes the handoff shape
flowchart TD
subgraph HydrationModel["Hydration Model (React / Vue)"]
H1["Download all component JS"] --> H2["Execute all component functions"]
H2 --> H3["Rebuild VDOM & attach listeners"]
end
subgraph ResumabilityModel["Resumability Model (Qwik)"]
R1["Zero JS executed on initial load"] --> R2["DOM serialized with state & handler symbols"]
R2 --> R3["User clicks button → Download & execute ONLY that event handler"]
endResumability can reduce eager browser execution by preserving more execution context across the server boundary.
Hydration versus resumability
| Hydration | Resumability |
|---|---|
| eagerly reconstructs client behavior | resumes behavior on demand |
| predictable component startup cost | more serialization and constraints |
| familiar event setup model | finer-grained lazy execution |
| browser executes more initially | browser may do less initial work |
Both require stable boundaries and explicit data transfer.
Resumability has constraints
The system must preserve enough information to resume:
- event handlers;
- closures or state references;
- component identity;
- serialized data;
- dependency relationships.
Reducing initial work can move complexity into build output and programming constraints.
Rendering topologies form a continuum
flowchart LR
A["Pure Static
(SSG)"] <--> B["Incremental
(ISR)"]
B <--> C["Server-Side
(SSR)"]
C <--> D["Streaming
(SSR + Suspense)"]
D <--> E["Islands
(Partial Hydration)"]
E <--> F["Pure Client
(CSR / SPA)"]Real applications can occupy several points at once.
Route decision: marketing homepage
Often:
- public;
- highly cacheable;
- content changes on a publishing schedule;
- limited interaction.
SSG or revalidated static output is often a strong starting point, with small client islands for interaction.
Route decision: product catalogue
Often:
- public or semi-public;
- filterable and paginated;
- freshness matters;
- URL state matters;
- some interaction is client-side.
Hybrid SSR/SSG with URL-driven client behavior or cached server queries may fit better than one whole-site rule.
Route decision: account page
Often:
- personalized;
- permission-sensitive;
- less cacheable publicly;
- form and interaction heavy.
Request rendering or a client application with a secure data boundary may be appropriate.
Route decision: rich internal editor
Often:
- authenticated;
- interaction-dense;
- stateful;
- device-aware;
- not valuable to search engines.
CSR or a hybrid shell with focused server data boundaries may be simpler and more effective.
A route decision matrix (Part 1)
| Question | Push toward |
|---|---|
| public and rarely changing? | SSG / revalidation |
| personalized per request? | SSR / client data |
| high interaction density? | client boundaries / CSR |
| slow independent region? | streaming |
A route decision matrix (Part 2)
| Question | Push toward |
|---|---|
| mostly static with small interactions? | islands |
| expensive browser startup? | server rendering / partial hydration |
| strict freshness and authorization? | request-time server boundary |
Personalization reduces cacheability
Separate:
Do not make an entire page uncacheable when only one small region depends on identity.
The boundary can be server-rendered, client-loaded, or streamed according to risk and UX.
Authentication and rendering
Server rendering can check authentication before producing private output.
But ensure:
- private data is not cached publicly;
- redirects do not leak resource existence;
- serialized props do not contain excess data;
- client boundaries enforce appropriate behavior too.
Rendering location does not replace authorization policy.
Error handling across topologies
Each topology moves failure to a different phase.
Route loading across topologies
Choose loading states that match the actual phase the user is waiting through.
Navigation after initial load
Client navigation can avoid a full document request, but it still may require:
- route data;
- server component payload;
- code chunks;
- cache lookup;
- state preservation;
- scroll and focus management.
Initial rendering strategy does not fully determine navigation behavior.
Full document navigation still has value
Full navigation can provide:
- a clean server boundary;
- lower client runtime assumptions;
- reliable reset of page state;
- progressive enhancement;
- simpler failure recovery.
Do not remove it merely because client navigation is fashionable.
Progressive enhancement and rendering
The basic task should remain understandable and usable when client code is delayed or unavailable, when the product permits it.
Performance is multidimensional
Measure:
- server latency;
- build time;
- HTML size;
- JavaScript transfer;
- hydration CPU;
- device memory;
- interaction readiness;
- cache hit rate;
- freshness and error recovery.
Optimizing one number can worsen the actual experience.
JavaScript budget
Client JavaScript has costs in:
- transfer;
- parse and compile;
- memory;
- event setup;
- hydration;
- battery;
- later navigation.
Ship browser code because the user needs the behavior, not because the framework can render it there.
HTML is already a runtime format
HTML provides:
- document structure;
- links;
- forms;
- semantic controls;
- accessibility relationships;
- progressive behavior.
Do not replace platform capabilities with client code without a clear benefit.
Server work has operational cost
Request rendering consumes:
- compute;
- memory;
- connection capacity;
- data-source capacity;
- observability and deployment complexity.
SSR is not free work merely because the browser does less.
Build work has operational cost
Static generation consumes:
- build time;
- CI resources;
- content pipeline capacity;
- deployment storage;
- invalidation and preview complexity.
Choose build-time rendering when its freshness and delivery benefits justify the workflow.
Client work has device cost
Browser work varies with:
- device CPU;
- memory;
- battery;
- network conditions;
- browser capability;
- competing applications.
Do not treat a fast developer laptop as a universal client.
The rendering cost triangle
flowchart TD
subgraph CostTriangle["The Web Rendering Cost Triangle"]
SC["Server & Request Cost\n(CPU, DB connections, edge compute bill)"]
BC["Build Cost\n(CI/CD minutes, deploy frequency, build queue)"]
CC["Client & Device Cost\n(Battery, main-thread blocking, RAM, TTI latency)"]
SC --- BC
BC --- CC
CC --- SC
endMoving work away from one corner usually adds cost or constraints to another.
Content freshness changes the choice
Freshness is a route requirement, not a framework preference.
User-specific state changes the choice
Private data may require:
- request-time authorization;
- private caching;
- client-side loading after a public shell;
- isolated personalized regions.
Separate public and private regions when possible to preserve cacheability.
Interaction density changes the choice
The amount and locality of interaction matter more than whether a page is “modern”.
Practical project: one platform, several topologies
Compare the same product catalogue requirements through CSR, SSR, SSG, streaming, and an island-like boundary.
Record data freshness, interaction, handoff, JavaScript, caching, and failure assumptions for each version.
Practical stages 1–4: establish the comparison
- Establish route data and interaction requirements.
- Build a CSR baseline.
- Create a minimal server-rendered version.
- Create a static version with explicit freshness assumptions.
Do not compare only screenshots. Compare the work and data path behind each result.
Practical stages 5–9: handoff and streaming
- Add streaming or a delayed independent region.
- Record the server-to-client handoff and hydration cost.
- Design streaming boundaries.
- Compare one large boundary with many tiny boundaries.
- Identify which regions truly need browser execution.
Measure timing and layout stability as well as total output.
Practical stages 10–14: hybrid framework concepts
- Compare Next.js conceptually.
- Compare Nuxt conceptually.
- Build a hybrid route plan.
- Isolate personalization.
- Explore islands.
The final choice should be route-specific rather than application-wide dogma.
Practical stages 15–19: reduce and justify browser work
- Delay island hydration.
- Explore resumability conceptually.
- Build a rendering decision matrix.
- Measure JavaScript responsibility.
- Draw the final hybrid architecture.
Verification: each topology has a stated freshness and interaction model, and server-only data stays server-side.
Practical extension: authenticated account route
Compare the public catalogue with an authenticated account route.
Explain why the decision differs in terms of:
- cacheability;
- authorization;
- personalization;
- interaction density;
- data freshness;
- handoff and failure behavior.
Try this yourself
For one application, classify these routes:
For each, choose where rendering happens, when it happens, what crosses to the browser, and what must be interactive.
Common rendering smells
Watch for:
- entire site forced into CSR because one page is interactive;
- entire application SSR because “SSR is faster”;
- every component marked client-side;
- huge framework bundle for a static page;
- personalized region preventing the whole page from caching;
- slow widget blocking all server HTML;
- same data fetched on server and immediately fetched again on client.
Hydration mismatch is usually a modeling signal
Instead of suppressing the warning, ask:
- which value differs?
- who owns it?
- was it available on both sides?
- should it be client-only?
- should the server serialize a stable value?
- is the initial output deterministic?
Fix the rendering model before hiding the symptom.
Avoid browser detection during render
The server and client may choose different trees.
Prefer CSS for presentation, a stable server-safe structure, or a client-only boundary when behavior truly depends on browser capability.
Dates, time zones, and random values
These are common mismatch sources:
Use a shared value, deterministic seed, explicit locale/timezone, or post-hydration update.
Streaming and failure isolation
A streamed region can fail after the shell is visible.
Define:
- a local fallback;
- retry behavior;
- whether prior content remains;
- logging and observability;
- accessible announcement of the failure.
Progressive delivery requires progressive recovery.
SSG and content publishing
A publishing workflow should answer:
- when is a page rebuilt?
- how are previews generated?
- how is stale content invalidated?
- can a failed build leave the previous version live?
- which content requires request-time freshness?
Rendering strategy includes the authoring and deployment system.
Rendering topology and deployment
Deployment capabilities are part of the architecture, not an afterthought.
Rendering topology and failure modes
Design the failure path at the same time as the happy path.
Practical decision checklist
Ask:
- Is the content public?
- How often does it change?
- Is it user-specific?
- How interactive is it?
- Does it need browser-only APIs?
- Can the output be cached?
- Are data dependencies slow?
- Can interaction be isolated?
- How much state must cross to the browser?
- What happens after initial navigation?
Completion checklist
- route strategy is chosen from requirements rather than slogans;
- build, request, and browser work are identified;
- hydration output is deterministic;
- secrets and server-only data stay server-side;
- interactive boundaries are no larger than necessary;
- streaming boundaries have meaningful loading and error behavior;
- cacheability and personalization are separated where possible;
- freshness and revalidation are explicit;
- JavaScript and device cost are measured;
- deployment and failure behavior match the topology.
Misconceptions to leave behind (Part 1)
| Misconception | Better mental model |
|---|---|
| CSR means no server | It moves initial rendering work to the browser |
| SPA and CSR are identical | Navigation model and rendering location differ |
| SSR means no JavaScript | Interactive regions still need browser code |
| SSR is always faster | Server latency, hydration, and cacheability matter |
| SSG cannot be interactive | Static output can include client islands |
| Static means always fresh | Publication and revalidation define freshness |
Misconceptions to leave behind (Part 2)
| Misconception | Better mental model |
|---|---|
| Hydration rebuilds the DOM | It attaches client behavior to existing output |
| Streaming removes handoff cost | It changes delivery timing, not all work |
| Server components are just SSR | Execution location and HTML timing differ |
| Islands are automatically better | Coordination and boundary costs still exist |
| Edge is always faster | Runtime, data location, and cache behavior matter |
| Resumability is just faster hydration | It changes the server-client handoff model |
The chapter in one sentence
Choose a rendering topology per route and region by balancing freshness, personalization, interaction, cacheability, handoff, deployment, and device cost.
Next: Chapter 12
The next chapter will build on rendering architecture with:
- testing strategy and quality boundaries;
- unit, integration, and end-to-end tests;
- browser behavior and accessibility verification;
- performance and failure testing;
- confidence in evolving front-end systems.
Questions
For your home page, catalogue, account page, and admin editor: where should the first useful HTML come from, and what is the smallest region that truly needs browser execution?
12 - Modern Build Systems, Development Tooling & Team Workflows
Modern Build Systems, Development Tooling & Team Workflows
Understand the path from source to delivery
Chapter 12
Polla Fattah
Today’s goal
Make the front-end toolchain visible as an architecture rather than a collection of commands.
We will connect:
- package management, lockfiles, scripts, and modules;
- resolution, transformation, development servers, and HMR;
- production bundling, tree shaking, minification, and source maps;
- asset fingerprints, CSS, environment variables, and code splitting;
- Vite, Rollup, esbuild, Rolldown, and Turbopack;
- linting, formatting, type checking, testing, Git, and CI;
- workspaces, monorepos, package boundaries, and dependency direction;
- an inspectable, reproducible build workflow.
By the end of today you can
- explain why development and production optimize differently;
- read a module graph and identify dependency direction;
- distinguish resolution, transformation, bundling, and serving;
- understand what tree shaking can and cannot remove;
- choose code-splitting boundaries based on route behavior;
- keep build-time configuration separate from runtime secrets;
- design fast local feedback and shared CI verification;
- use workspaces without confusing them with architecture;
- expose package APIs without leaking private files;
- inspect emitted assets instead of guessing about performance.
The central principle
A build toolchain is a delivery architecture: it transforms source modules into environment-specific artifacts while preserving reproducibility, boundaries, and useful feedback.
The tool is not only a compiler.
It resolves dependencies, serves development code, creates production artifacts, and shapes how teams work.
The source-to-delivery pipeline
flowchart LR
A["1. Source Files
(TS, TSX, CSS)"] --> B["2. Package Resolution
(node_modules, exports)"]
B --> C["3. Module Graph
(Static dependency DAG)"]
C --> D["4. Transformation
(TypeScript/JSX stripping)"]
D --> E["5. Bundler / Tree Shaking
(Chunk splitting & minification)"]
E --> F["6. Production Artifacts
(Hashed JS, CSS, Source Maps)"]Each stage answers a different question and has different failure modes.
Development and production optimize differently
flowchart TD
subgraph DevelopmentMode["Development Environment (Inner Loop)"]
D1["Unbundled Native ESM
(Instant server boot)"]
D2["Hot Module Replacement (HMR)
(<50ms stateful updates)"]
D3["Detailed Source Maps & Error Overlays"]
end
subgraph ProductionMode["Production Environment (Delivery Artifacts)"]
P1["Aggressive Dead-Code Elimination (Tree Shaking)"]
P2["Route-Level Code Splitting & Chunking"]
P3["Content-Hashed Filenames for Immutable CDN Caching"]
P4["Byte Minification & Production Tree Stripping"]
endA development server is not automatically a production server.
The two modes can share configuration while serving different goals.
Package management is the beginning of the toolchain
Packages define:
- source dependencies;
- executable scripts;
- development tools;
- transitive dependency graphs;
- version constraints;
- reproducible installation inputs.
The package manifest is an architectural document, not only an install list.
Runtime and development dependencies
Misclassifying a dependency can:
- inflate production output;
- break a library consumer;
- make a production build depend on local tooling;
- hide a missing runtime requirement.
Semantic version ranges are constraints
The range describes acceptable versions.
The lockfile records the concrete resolution used by an installation.
Do not confuse the manifest promise with the installed graph.
Lockfiles belong in application repositories
A lockfile records:
- concrete versions;
- resolved locations;
- integrity information;
- transitive dependency choices.
It lets local development, CI, and deployment start from the same dependency graph.
Update it intentionally and use one package manager consistently.
Package scripts create a common interface
Scripts hide machine-specific command details and give the team a shared workflow vocabulary.
ES modules are the structural foundation
Static module syntax gives tools information about dependencies, exports, and boundaries.
Static imports are analyzable
The tool can use static imports to:
- build a dependency graph;
- detect missing modules;
- identify unused exports;
- split and optimize code;
- pre-bundle dependencies.
Dynamic imports create asynchronous boundaries
Dynamic imports can create:
- lazy routes;
- optional features;
- smaller initial chunks;
- delayed dependency loading.
They also add loading, error, caching, and prefetch decisions.
CommonJS still exists
Interoperability can affect:
- static analysis;
- default exports;
- tree shaking;
- runtime loading;
- package conditions.
Know the module format at the boundary you are integrating.
Module resolution answers “which file?”
Resolution may consider:
- relative paths;
- package exports;
- file extensions;
- aliases;
- conditions for browser, node, import, or require;
- workspace links;
- type declarations.
The module graph begins only after resolution succeeds.
Aliases can improve architecture - or hide it
Aliases can make stable package boundaries readable.
They can also conceal deep dependencies and make imports appear more independent than they are.
Use aliases to communicate architecture, not to avoid designing it.
Transformation is not one operation
flowchart LR
Src["Source Code
(TypeScript / JSX)"] --> Parse["Parser
(Generate AST)"]
Parse --> AST["Abstract Syntax Tree
(AST Data Structure)"]
AST --> Transform["Transforms
(Strip types, lower syntax)"]
Transform --> Emit["Code Generator
(Standard ES2022 JavaScript)"]Transformation may include:
- TypeScript syntax removal;
- JSX transformation;
- language lowering;
- macro or plugin transforms;
- CSS processing;
- asset URL rewriting.
TypeScript compilation has separate jobs
A fast transpiler can remove types without proving the program correct.
Run type checking as its own explicit quality step when the build tool does not perform it.
Transpilation and polyfills are different
Transpilation can rewrite syntax:
It does not automatically provide missing runtime APIs such as a browser feature.
Polyfills, browser targets, and runtime support are separate decisions.
Development servers optimize feedback
A development server may provide:
- native module serving;
- on-demand transformation;
- dependency pre-bundling;
- source maps;
- error overlays;
- HMR or Fast Refresh.
It should make the smallest useful update quickly and explain failures clearly.
Hot Module Replacement
sequenceDiagram
autonumber
actor Dev as Developer
participant FS as File System Watcher
participant Server as Dev Server (Vite)
participant Browser as Browser Client
Dev->>FS: Saves Button.tsx
FS->>Server: File change detected
Server->>Server: Re-transform single module (15ms)
Server-->>Browser: WebSocket: { type: 'update', path: '/src/Button.tsx' }
Browser->>Server: HTTP fetch(/src/Button.tsx?t=171000)
Server-->>Browser: Fresh module code
Note over Browser: HMR Runtime replaces module without full page reload!
Component state preserved.HMR shortens feedback loops, but it is not identical to a fresh application start.
Fast Refresh is framework-aware HMR
Framework-aware refresh can preserve component state when a change is safe.
It may reset state when:
- module exports change shape;
- boundaries are not refresh-safe;
- initialization semantics change;
- the framework cannot preserve identity.
Test both preserved and reset behavior when it matters.
HMR correctness matters
Hot updates can expose bugs hidden by full reloads:
- duplicate subscriptions;
- stale module state;
- missing cleanup;
- side effects at import time;
- global registration repeated on every update.
Write initialization and cleanup so reload behavior remains predictable.
Vite as a reference development tool
Vite’s development model emphasizes:
- fast startup;
- native ESM-like module requests;
- on-demand transforms;
- dependency pre-bundling;
- a separate production build pipeline.
The exact internals can evolve; the development versus production distinction remains useful.
Dependency pre-bundling
Third-party packages may contain many modules or formats that are expensive to request individually.
Pre-bundling can:
- reduce browser request count;
- normalize dependency formats;
- improve development startup after caching.
It does not mean the production bundle and development graph are identical.
Native ESM is an important mental model
flowchart TD
subgraph NativeESM["Unbundled Dev Server (Vite / Dev)"]
B["Browser requests /src/main.ts"] --> S["Dev Server compiles on demand"]
S --> M1["main.ts imports App.ts"]
M1 --> M2["App.ts imports Button.ts"]
M2 --> M3["Browser requests individual modules via native HTTP/2"]
end
subgraph ProductionBundling["Bundled Production Output (Rollup / esbuild)"]
G["Module Graph Crawler"] --> T["Tree Shaking & DCE"]
T --> C1["Chunk 1: app.8f2a.js (120KB)"]
T --> C2["Chunk 2: vendor.3c1d.js (80KB)"]
endDevelopment tools can preserve a module-oriented model while adding transforms, caching, and dependency optimization.
Do not assume that a development request corresponds one-to-one with a production asset.
Production bundling is graph transformation
Bundling includes:
- module combination;
- dependency ordering;
- tree shaking;
- code splitting;
- asset rewriting;
- minification;
- source-map generation.
It is much more than concatenation.
Tree shaking removes unreachable exports
If the graph and package semantics allow it, the unused export may be removed from a production chunk.
The tool needs analyzable module structure and correct side-effect information.
Tree shaking depends on analyzable code
Dynamic behavior can limit removal:
- unpredictable property access;
- CommonJS patterns;
- runtime module discovery;
- side effects hidden in imports;
- package metadata that is too broad.
The tool can only prove what the code structure makes visible.
Side effects matter
An import may be valuable even when it has no exported binding.
Do not mark a package as side-effect-free unless removing such imports is safe.
Minification changes representation
Minification can:
- shorten identifiers;
- remove whitespace;
- simplify expressions;
- eliminate unreachable code;
- compress repeated structures.
It improves transfer and sometimes execution, but it makes debugging harder without source maps.
Source maps connect artifacts to source
flowchart LR
Err["Runtime Error in Browser
(vendor.8f31c.js:1:4210)"] --> Map["Source Map File
(vendor.8f31c.js.map
VLQ Mappings)"]
Map --> Src["Original Source Code in DevTools
(src/services/permitApi.ts:42:15)"]Source maps are an operational policy:
- public or private?
- uploaded to an error service?
- exposed in production?
- retained for which release?
Asset fingerprinting enables safe caching
Content-based names let immutable assets use long cache lifetimes.
HTML or manifests must point to the current fingerprinted files.
Code splitting creates delivery choices
flowchart TD
Entry["Main Entrypoint (main.ts)"] --> AppChunk["Initial Core Bundle
(Header, Nav, Router, Theme)
[app.8f31c.js - 45KB]"]
AppChunk -.->|Static Import| Shared["Shared Vendor Chunk
(React, Query Cache)
[vendor.2e1a.js - 75KB]"]
AppChunk -->|Dynamic import('./Reports')| RouteA["Async Route: Reports & Charts
[reports.6d4b.js - 180KB]"]
AppChunk -->|Dynamic import('./Admin')| RouteB["Async Route: Admin Console
[admin.9c2e.js - 95KB]"]Split points affect:
- initial transfer;
- later navigation latency;
- cache reuse;
- request count;
- failure and loading behavior.
Route-level splitting is often a strong default
Routes usually represent meaningful user journeys and can isolate code that is not needed initially.
They also provide a natural loading and error boundary.
More chunks are not always better
Tiny chunks can create:
- request overhead;
- scheduling complexity;
- waterfall risk;
- poor cache reuse;
- difficult debugging.
Split around behavior and navigation, not every file.
Shared chunks are a trade-off
Shared dependencies can be downloaded once and reused.
But a large shared chunk can become part of every route’s initial cost even when only one feature uses it.
Analyze the actual graph and route traffic.
Preloading and prefetching split code
Use hints based on confident navigation and device/network conditions.
Prefetching code the user never needs is still work.
CSS participates in the build pipeline
The pipeline may:
- process imports;
- scope or extract styles;
- rewrite asset URLs;
- split CSS by entry or route;
- minify output;
- preserve source maps.
CSS loading order and extraction affect rendering and layout stability.
Asset imports are module dependencies
The tool can fingerprint assets, generate URLs, and include only resources reachable from the graph.
The runtime receives a reference appropriate to the target environment.
Environment variables have a boundary
flowchart TD
subgraph BuildTime["Build-Time Replacement (Vite / Bundler)"]
E1["import.meta.env.VITE_API_URL"] --> R1["Replaced during build with string literal
'https://api.erbil.gov.krd'"]
Note1["WARNING: Embedded into public client bundle!
Never place database passwords here."]
end
subgraph RuntimeConfig["Runtime Environment Configuration"]
E2["process.env.DATABASE_PASSWORD"] --> R2["Read dynamically on server execution"]
Note2["Safe: Stays inside secure container / server."]
endBrowser-exposed variables are public.
.env naming does not make a value secret once it is included in client output.
Build-time and runtime configuration differ
Build-time values require rebuilding to change.
Runtime values can vary per deployment or request without producing new assets.
Choose based on:
- environment portability;
- caching;
- deployment frequency;
- secrecy;
- per-request variation.
Vite production builds
A production build typically performs:
- module resolution;
- transformation;
- dependency graph optimization;
- code splitting;
- asset fingerprinting;
- minification;
- source-map output according to policy.
Inspect the generated result instead of inferring it from development requests.
Rollup
Rollup is a graph-oriented bundler known for:
- ES module analysis;
- library and application bundling;
- tree shaking;
- plugin-based transformation;
- controllable output formats.
Use it directly when its lower-level control matches the responsibility you need.
esbuild
esbuild emphasizes very fast native transformation and bundling.
It can be useful for:
- fast development transforms;
- custom build integration;
- straightforward application builds;
- tooling that needs quick parsing and emission.
Fast transformation does not remove the need for application architecture or type checking.
Rolldown and Turbopack
Modern tools evolve their internals to improve:
- incremental builds;
- graph computation;
- parallelism;
- native performance;
- framework integration.
Learn the responsibility each tool solves instead of treating its implementation name as the architecture.
Vite and Turbopack solve related problems differently
Tools can differ in when they bundle, how they cache, and how they integrate with a framework.
The stable mental model is the source graph and emitted responsibility.
The framework often chooses the tool
Framework defaults can determine:
- development server;
- route build model;
- server/client boundaries;
- asset handling;
- production output;
- plugin lifecycle.
Developers still need to understand the underlying responsibilities when diagnosing output or performance.
Plugins extend the pipeline
Plugins may add:
- syntax transforms;
- virtual modules;
- asset handling;
- route generation;
- environment integration;
- development middleware.
Every plugin adds behavior to resolution, transformation, or output. Keep plugin purpose and order understandable.
Avoid toolchain configuration as a hobby
Configuration should answer a delivery or developer-experience need.
Before adding a plugin or custom transform, ask:
- which problem does it solve?
- who owns it?
- what does it change in production?
- how is it tested?
- what is the fallback if it breaks?
Complexity without a user or team benefit is a liability.
Linting catches a class of problems
Linting can detect:
- suspicious patterns;
- unused bindings;
- import restrictions;
- unsafe APIs;
- inconsistent architectural rules.
It is not formatting, type checking, or a substitute for tests.
Formatting reduces diff noise
A formatter creates a shared representation so reviews focus on behavior and design.
Format consistently and avoid repeatedly reformatting unrelated files in feature commits.
Formatting should support review, not dominate it.
Type checking verifies static contracts
flowchart TD
subgraph QualityPillars["The Five Quality Gates of Front-End Delivery"]
Q1["1. Linter (ESLint / Biome)
Flags bug-prone anti-patterns & security flaws"]
Q2["2. Formatter (Prettier / Biome)
Guarantees deterministic code style across team"]
Q3["3. Type Checker (tsc --noEmit)
Verifies compile-time type safety & API contracts"]
Q4["4. Test Runner (Vitest / Playwright)
Verifies runtime behavioral correctness"]
Q5["5. Production Bundler (Vite build)
Verifies module graph resolution & asset generation"]
endThese gates overlap in value but do not answer the same question.
Local feedback should be fast
flowchart LR
IDE["1. IDE / Editor
(<50ms inline feedback)"] --> GitHook["2. Pre-Commit Hook
(lint-staged on changed files)"]
GitHook --> PR["3. Git Pull Request"]
PR --> CI["4. CI Automation Pipeline
(Full clean install, typecheck, test, build)"]
CI --> Prod["5. Verified Deployment Artifact"]Fast feedback catches cheap mistakes close to the change.
CI provides a shared clean-environment check.
Git is part of the engineering toolchain
Git supports:
- reviewable change boundaries;
- reproducible history;
- rollback;
- release traceability;
- collaboration across branches.
Treat commit and merge practices as part of delivery quality.
Keep commits reviewable
Prefer commits that separate:
- mechanical formatting;
- dependency updates;
- tool configuration;
- feature behavior;
- generated artifacts.
Small coherent changes make toolchain failures easier to bisect and understand.
Generated files need a policy
Decide which artifacts are:
- committed;
- generated in CI;
- published separately;
- reproducible from source;
- ignored locally.
Inconsistent policies create noisy diffs and uncertain releases.
Environment reproducibility
Reproducibility depends on:
- lockfile;
- package-manager version;
- Node/runtime version;
- build configuration;
- environment inputs;
- operating-system assumptions;
- clean install behavior.
Document and automate the parts that affect artifacts.
CI is the shared verification environment
CI should run from a clean state and verify the contracts that matter:
The exact order can vary, but the environment should not depend on one developer’s machine.
Build once, deploy predictably
Rebuilding separately for each environment can produce different output and weaken release confidence.
Keep truly runtime-varying configuration outside immutable assets where possible.
Workspaces solve package coordination
Workspaces can:
- install multiple packages together;
- link internal packages;
- share scripts and dependency policy;
- coordinate builds;
- make local package changes visible quickly.
They are a package-management feature, not automatically a monorepo architecture.
A workspace is not automatically a monorepo architecture
A workspace can host one application plus tools.
A monorepo still needs package boundaries, ownership, and dependency direction.
Monorepo benefits
Potential benefits include:
- shared code with local changes;
- coordinated releases;
- consistent tooling;
- cross-package refactors;
- shared CI infrastructure.
The benefits appear only when boundaries and workflows remain understandable.
Monorepo costs
Costs include:
- larger dependency graph;
- longer or more complex CI;
- ownership ambiguity;
- cross-package build ordering;
- accidental coupling;
- difficult versioning decisions.
Do not choose a monorepo solely because it is fashionable.
Package boundaries should represent architecture
The directory is useful when it reflects a responsibility and public API.
Internal packages need public APIs too
Prefer an intentional package entry point over:
Private file imports bypass encapsulation and make internal restructuring expensive.
Dependency direction matters
flowchart TD
App["apps/citizen-portal (Applications)"] --> Feat["packages/feature-permits (Feature Libraries)"]
Feat --> Domain["packages/domain-licensing (Domain Models & Rules)"]
Domain --> Core["packages/ui-components (Design System & Primitives)"]
Core --> Util["packages/utilities (Pure helpers & math)"]Lower-level packages should not import higher-level application decisions.
Direction makes ownership and reuse possible.
Circular dependencies are design feedback
Even if the build succeeds, cycles can create:
- partial initialization;
- undefined exports;
- confusing evaluation order;
- difficult testing;
- architecture that cannot be layered cleanly.
Break the cycle by clarifying ownership or extracting a stable lower-level contract.
Analyze the dependency graph
Look for:
- unexpected large imports;
- application code imported by shared packages;
- cycles;
- duplicate versions;
- route chunks containing unrelated features;
- a dependency that dominates initial transfer.
The graph turns performance and architecture discussions into evidence.
Development and production differ
Always run a production build before making claims about production artifacts.
dev is not a production server
Development servers may:
- transform on demand;
- tolerate missing optimizations;
- expose source files;
- use different caching;
- allow permissive CORS or proxies;
- rely on local filesystem behavior.
Use a production build and production-like serving environment for delivery verification.
Build targets define browser assumptions
Targets influence:
- syntax transformation;
- polyfill decisions;
- output size;
- supported APIs;
- debugging expectations.
Choose a browser policy deliberately and keep it visible to the team.
Baseline and browser policy
A supported-browser policy should answer:
- which browsers and versions are supported;
- which features are assumed;
- when a feature needs a fallback;
- how support changes are reviewed;
- how real devices are tested.
The build cannot invent a product support policy.
Library builds versus application builds
Library output needs stable formats, declarations, exports, and consumer compatibility.
Application output can make more assumptions about its deployment.
Dependency externalization
Libraries may leave dependencies external so consumers provide them.
This avoids duplicating framework code but creates peer-version and runtime compatibility obligations.
Decide which code belongs in the library artifact and which belongs to the consuming application.
Tree shaking and package design
Packages are easier to optimize when they:
- use analyzable ES modules;
- expose focused entry points;
- avoid import-time side effects;
- declare side-effect behavior accurately;
- avoid pulling a whole framework for one helper.
Package API design directly affects emitted application code.
Development proxying
A development proxy can route browser requests to an API and reduce local cross-origin friction.
It should not hide production differences in:
- origin;
- cookies;
- headers;
- HTTPS;
- path rewriting;
- error behavior.
Document what the proxy changes.
HTTPS in development
HTTPS may be required to reproduce:
- secure cookies;
- service workers;
- browser APIs requiring a secure context;
- mixed-content behavior;
- realistic origin policy.
Use a consistent local certificate workflow when the application depends on these capabilities.
Environment parity
Compare development and production for:
- origin and base path;
- asset URLs;
- API proxying;
- environment variables;
- compression;
- source maps;
- service-worker behavior;
- caching headers.
Parity is a debugging tool, not a demand that every environment be identical.
Team tooling should be boring
Good tooling is:
- documented;
- reproducible;
- fast enough to use;
- hard to misconfigure;
- easy to upgrade deliberately;
- understandable by the whole team.
The goal is dependable delivery, not admiration for configuration cleverness.
Avoid global tool dependencies
Prefer project-local versions for:
- bundlers;
- formatters;
- linters;
- test runners;
- type tooling;
- code generators.
Global tools can silently differ from CI and another developer’s environment.
Editor integration is developer experience
Editor support can provide:
- type feedback;
- import navigation;
- formatting;
- lint diagnostics;
- test discovery;
- refactoring support.
The editor should use the repository’s configuration rather than a parallel personal toolchain.
Pre-commit hooks are a feedback boundary
Use hooks for fast checks that should block obviously broken commits.
Keep them bounded.
Long full builds in every commit encourage bypassing the hook; put broad verification in CI.
Dependency updates need policy
For updates, consider:
- security advisories;
- lockfile changes;
- peer compatibility;
- bundle impact;
- behavior changes;
- migration notes;
- rollback path.
Toolchain dependencies can change emitted artifacts even when application code is unchanged.
Toolchain dependencies have supply-chain risk
Build tools execute code in a privileged development and CI context.
Reduce risk with:
- lockfiles and review;
- trusted registries;
- minimal dependencies;
- update monitoring;
- restricted CI credentials;
- artifact inspection.
The toolchain is part of the application’s attack surface.
A reference workflow
flowchart TD
A["1. Developer edits component in IDE"] --> B["2. Instant HMR verification in browser"]
B --> C["3. Local targeted lint and test execution"]
C --> D["4. Git commit & push to pull request"]
D --> E["5. CI executes clean install (npm ci with locked dependencies)"]
E --> F["6. Automated verification: lint + tsc + test suite"]
F --> G["7. Production build & asset budget check"]
G --> H["8. Immutable content-hashed artifacts deployed to CDN"]Each stage should add confidence without repeating every earlier cost.
Practical project structure
The exact files vary.
The important point is that source, configuration, scripts, and lockfile form one reproducible project boundary.
Build pipeline for the storefront
flowchart LR
S1["1. Resolve Imports"] --> S2["2. Transform TS / JSX"]
S2 --> S3["3. Route Code Splitting"]
S3 --> S4["4. Tree Shaking & DCE"]
S4 --> S5["5. Minification"]
S5 --> S6["6. Asset Fingerprinting
(app.8f31c.js)"]
S6 --> S7["7. Emit Manifest & Maps"]Trace a source module through this pipeline to understand what the browser receives.
Lazy loading reports
The reports route should not appear in the initial chunk when the build and route architecture permit splitting.
Verify with emitted assets and a network trace rather than trusting the source syntax alone.
Inspect the build instead of guessing
Useful evidence includes:
- asset sizes;
- chunk composition;
- duplicate dependencies;
- source-map module lists;
- route request waterfalls;
- cache headers;
- compressed transfer sizes.
Build analysis is most useful when connected to a user journey.
Toolchain smells
Watch for:
- build configuration nobody understands;
- every project using different lint rules;
- lockfiles regenerated by multiple managers;
- production builds rarely run;
- huge initial bundles despite route structure;
- hundreds of tiny chunks;
- secrets in front-end environment variables;
- internal package imports through private paths;
- circular dependencies everywhere.
Choose tools by responsibility
Choose the smallest toolset that serves the responsibility.
Current tooling changes; architecture remains
Tool internals evolve.
The stable questions remain:
- what is the module graph?
- where are transformation boundaries?
- what ships initially?
- what loads later?
- what is public configuration?
- how is output verified?
- who owns each package boundary?
Practical lab: Inspect a Modern Front-End Toolchain
Use a modern build tool to observe resolution, transformation, module graphs, code splitting, and deployment artifacts.
The practical compares source modules with development requests and production output.
Practical stages 1–4: create and inspect
- Create a small TypeScript application.
- Add a dynamically imported reports route.
- Inspect development requests and production chunks.
- Compare source modules with emitted assets and source maps.
Record which decisions belong to the toolchain and which belong to application architecture.
Practical stages 5–9: build and split
- Separate type checking.
- Create a production build.
- Explore tree shaking.
- Explore minification.
- Add route-level code splitting.
Verify that reports code is not in the initial route chunk when the build permits splitting.
Practical stages 10–14: optimize and verify
- Prefetch a lazy route.
- Analyze a large dependency.
- Add linting and formatting.
- Add CI-style verification.
- Compare build-time and runtime configuration.
Interpret bundle size alongside route behavior and field performance.
Practical stages 15–18: boundaries and architecture
- Explore package encapsulation.
- Compare Vite and low-level bundler responsibilities.
- Research Turbopack through a Next.js example.
- Create a build architecture diagram.
Do not present a dependency-graph visualizer as a production bundler.
Practical extension: educational graph visualizer
Build a deliberately limited visualizer showing:
Label it educational.
The goal is to make resolution and dependency direction visible, not to recreate a production bundler.
Try this yourself
Pick one initial route and answer:
- which source modules does it need?
- which dependency dominates its size?
- which feature can split at a route boundary?
- what can be tree-shaken?
- what must remain a side effect?
- which environment values are public?
Then verify every answer against the emitted build.
Troubleshooting guide (Part 1)
| Symptom | Likely cause |
|---|---|
| Dev works, production fails | Different transforms, paths, or environment assumptions |
| Reports code is in the initial chunk | Import is static or split boundary is ineffective |
| Unused package code remains | Side effects, module format, or graph opacity |
| CI differs from local | Lockfile, runtime, or global tool mismatch |
| Secret appears in client output | Public build-time variable was treated as private |
Troubleshooting guide (Part 2)
| Symptom | Likely cause |
|---|---|
| Package consumers import private files | Public API is incomplete or undocumented |
| HMR behaves strangely | Import-time side effects or cleanup missing |
| Tiny chunks hurt navigation | Split points follow files rather than user journeys |
| Build graph contains cycles | Dependency direction is unclear |
Completion checklist
- package versions and lockfile are reproducible;
- resolution and module boundaries are understandable;
- type checking is an explicit quality step;
- development and production workflows are distinguished;
- code splitting follows route or feature behavior;
- tree-shaking assumptions are validated by output;
- source maps and environment values have policy;
- package APIs protect internal files;
- CI verifies a clean production build;
- emitted artifacts are inspected rather than guessed at.
Misconceptions to leave behind (Part 1)
| Misconception | Better mental model |
|---|---|
| A build tool is just a compiler | It resolves, transforms, serves, bundles, and emits |
| Bundling means concatenation | It is graph transformation and asset design |
TypeScript must emit JavaScript through tsc | Type checking and transformation can be separate |
| Transpilation adds missing browser APIs | Polyfills and runtime support are separate |
| More code splitting is always better | Split around user journeys and costs |
| Tree shaking removes anything not called | It depends on analyzable graphs and side effects |
Misconceptions to leave behind (Part 2)
| Misconception | Better mental model |
|---|---|
| Minification makes source maps unnecessary | Debugging still needs source policy |
.env values are secret | Client-exposed values are public |
| A workspace is a monorepo architecture | Coordination tooling and boundaries differ |
| A successful circular build is healthy | Cycles are dependency-design feedback |
| The dev server is production | Production artifacts and serving behavior differ |
| Quality gates are interchangeable | Lint, format, types, tests, and build answer different questions |
The chapter in one sentence
Treat the toolchain as an inspectable delivery graph that transforms source into reproducible artifacts while preserving package, environment, and team boundaries.
Next: Chapter 13
The next chapter will build on tooling and architecture with:
- testing strategy and confidence boundaries;
- unit, integration, and end-to-end verification;
- browser behavior and accessibility tests;
- performance and failure testing;
- quality workflows for evolving applications.
Questions
Can you trace one user-visible route from its source entry point to the emitted assets, and explain which toolchain decisions affect its first-load cost?
13 - Front-End Security, Authentication & Browser Isolation
Front-End Security, Authentication & Browser Isolation
Design the boundary before the attack
Chapter 13
Polla Fattah
Today’s goal
Treat the browser as a security runtime with explicit trust boundaries.
We will connect:
- origins, same-origin policy, and CORS;
- XSS, encoding, sanitization, CSP, and Trusted Types;
- CSRF, cookies, sessions, and logout;
- authentication, authorization, bearer tokens, OAuth, and PKCE;
- BFF architecture and token storage trade-offs;
- secrets, SRI, dependencies, and third-party scripts;
- clickjacking, framing, COOP, COEP, and CORP;
- secure messaging, iframes, logging, and a review checklist.
By the end of today you can
- identify the origin and trust boundary of a browser request;
- explain what CORS does and does not protect;
- trace untrusted input from source to dangerous sink;
- prefer safe text rendering and reviewed sanitization;
- use CSP as defense in depth;
- distinguish XSS from CSRF and authentication from authorization;
- choose cookie, BFF, or browser-token architecture deliberately;
- explain OAuth authorization code flow with PKCE;
- keep secrets out of browser bundles and URLs;
- review third-party, framing, messaging, and isolation risks.
The central principle
Front-end security comes from layered trust boundaries, not from one framework feature, header, token format, or browser flag.
Ask where data comes from, who can read it, who can change it, and which layer must enforce the decision.
The security review progression
flowchart TD
A["1. Identify Origins & Trust Boundaries"] --> B["2. Trace Untrusted Input (Sources)"]
B --> C["3. Secure Rendering Sinks (XSS Defense)"]
C --> D["4. Protect State-Changing Requests (CSRF)"]
D --> E["5. Separate Authentication from Authorization"]
E --> F["6. Secure Credentials & Token Storage (BFF)"]
F --> G["7. Enforce Browser Isolation (CSP, COOP, COEP)"]
G --> H["8. Audit Dependencies & Third-Party Scripts"]Security is a system of related controls, not a checklist of isolated switches.
The browser is a multi-origin runtime
The browser applies different rules to interactions among these origins.
Your architecture must make those relationships intentional.
What is an origin?
An origin is the combination of:
flowchart LR
subgraph OriginTuple["The Web Origin Definition (RFC 6454)"]
S["Scheme (e.g. https://)"] --- H["Host (e.g. app.erbil.gov.krd)"] --- P["Port (e.g. :443)"]
endExamples:
Small differences can produce different origins.
Same-origin examples
These share an origin when scheme, host, and port match.
Paths differ, but origin policy is not path policy.
Origin is not the same as site
Cookies and browser policies sometimes reason about a broader “site” concept based on registrable domains.
Do not use “same site” and “same origin” interchangeably.
They affect different browser mechanisms.
Same-origin policy
The same-origin policy restricts how a document can read or interact with resources from another origin.
It is a browser protection boundary.
It does not mean cross-origin activity is impossible.
Same-origin policy is not “no cross-origin activity”
Browsers may allow controlled cross-origin actions such as:
- loading images;
- submitting forms;
- embedding frames;
- sending requests;
- loading scripts under policy rules.
The key question is often whether the initiating page can read the response or control the embedded context.
Cross-origin reads are the core concern
CORS controls whether browser JavaScript may read a cross-origin response.
It is not a universal firewall around the server.
Browser storage is origin-scoped
Storage isolation helps protect data between origins.
It does not protect an origin from XSS that executes inside that origin.
CORS is a server permission
The server tells the browser which origin may read a response.
The browser enforces the permission for script access.
CORS is not authentication
CORS does not prove who the user is.
It does not issue a session, validate an access token, or decide whether a user may update a record.
Authentication and authorization remain server responsibilities.
CORS does not protect the server from non-browser clients
Command-line clients, native apps, scripts, and attackers can send requests without browser CORS enforcement.
The server must authenticate, authorize, validate, and rate-limit independently.
Simple and preflighted CORS requests
Some cross-origin requests can proceed with a browser permission check.
Others trigger an OPTIONS preflight describing:
- intended method;
- requested headers;
- requested origin.
The server must explicitly allow the operation before the browser sends the actual request.
Preflight is not a failure
sequenceDiagram
autonumber
actor Browser as Browser Client (https://app.erbil.gov.krd)
participant API as Municipal API (https://api.erbil.gov.krd)
Browser->>API: OPTIONS /permits/104 (Preflight)
Origin: https://app.erbil.gov.krd
Access-Control-Request-Method: PUT
Access-Control-Request-Headers: Content-Type
Note over API: API verifies origin in allowed whitelist
API-->>Browser: 204 No Content
Access-Control-Allow-Origin: https://app.erbil.gov.krd
Access-Control-Allow-Methods: GET, PUT, POST
Access-Control-Allow-Headers: Content-Type
Browser->>API: PUT /permits/104 (Actual Mutation Request)
API-->>Browser: 200 OK (Resource updated)Preflight is a safety mechanism.
If it fails, inspect origin, method, headers, credentials, and server policy rather than disabling security blindly.
Credentials and CORS need alignment
Credentialed requests require coordinated policy:
- explicit allowed origin;
- allowed credentials;
- cookie attributes;
- server authentication;
- CSRF protection where relevant.
Wildcard origins are not a substitute for a deliberate credential policy.
CORS errors often reveal architecture problems
A CORS failure may indicate:
- API and app origins were not designed together;
- development proxy hid production behavior;
- credential policy is unclear;
- a public/private boundary is ambiguous;
- the browser is being asked to call a server that should be behind a BFF.
Fix the boundary, not only the console message.
Development proxies can hide CORS
Test production-like origins before shipping.
Otherwise the first real CORS behavior appears in deployment.
XSS is more than <script> tags
Cross-site scripting occurs when attacker-controlled data becomes executable or dangerous browser content.
Possible paths include:
- HTML injection;
- event-handler attributes;
- dangerous URLs;
- script-capable SVG;
- template or expression injection;
- DOM APIs that interpret strings as markup.
Safe rendering by default
Prefer APIs and framework bindings that treat values as text.
Escaping by default is useful, but still inspect escape hatches and URL/style contexts.
Dangerous escape hatches
Treat these as security-sensitive:
Every escape hatch needs a documented trust boundary and a reviewed input policy.
Encoding and sanitization are different
Encoding is context-specific.
Sanitization is a policy for permitting a restricted subset of content.
Neither should be applied blindly to every output context.
Prefer text over HTML
If the product requirement is text, do not create an HTML parsing problem.
Rich HTML should be an explicit feature with an explicit content policy.
DOM-based XSS
flowchart TD
subgraph Sources["Untrusted Sources"]
S1["location.search / hash"]
S2["API JSON responses"]
S3["localStorage / cookies"]
S4["postMessage events"]
S5["User form inputs"]
end
subgraph DangerousSinks["Dangerous DOM Sinks (Vulnerabilities)"]
D1["element.innerHTML"]
D2["dangerouslySetInnerHTML / v-html"]
D3["eval() / new Function()"]
D4["<a href='javascript:...'>"]
D5["document.write()"]
end
Sources -->|Direct assignment without sanitization| DangerousSinks
DangerousSinks --> XSS["Cross-Site Scripting (XSS)
Attacker script executes with full user privileges!"]The server does not need to be involved.
Client code can create XSS by moving an untrusted value into a dangerous sink.
Sources and sinks
Security review traces values from source to sink and asks what validation or encoding occurs between them.
textContent is safer for text
It creates a text node rather than parsing markup.
Use the simplest API that matches the content requirement.
URL handling needs context
Validate:
- allowed schemes;
- allowed hosts when appropriate;
- relative versus absolute behavior;
- redirect policy;
- display text separately from destination.
Text escaping alone does not make every URL safe.
Avoid eval()-style execution
Never turn untrusted strings into code through:
eval();Function();- string-based timers;
- dynamic script construction;
- template expression interpreters without a trusted boundary.
If the product needs expressions, design a constrained language and parser rather than executing JavaScript.
Sanitizing rich HTML
If rich content is required:
- define the allowed elements and attributes;
- sanitize with a maintained, reviewed library;
- sanitize near the trust boundary;
- preserve the sanitized representation;
- render through one controlled component;
- test dangerous payloads and URL contexts.
Do not assume a generic “clean HTML” label communicates the policy.
Sanitization should happen near the trust boundary
Repeated ad hoc sanitization in every component creates inconsistent policies and missed sinks.
Content Security Policy is defense in depth
CSP can restrict what the browser may execute or load if an injection reaches the page.
It should support safe architecture, not justify unsafe rendering.
Start from the resources the application actually needs.
default-src establishes a baseline
More specific directives can control:
The policy should be reviewed with application dependencies and deployment origins.
Script policies matter most
Avoid broad script permissions where possible.
Prefer:
- external scripts from known origins;
- nonces or hashes for deliberate inline code;
- removal of inline handlers;
- restricted dynamic execution.
unsafe-inline and unsafe-eval should be treated as explicit trade-offs, not defaults.
Nonces authorize specific inline scripts
The server places the same unpredictable nonce on an intentionally permitted script.
Never reuse a predictable or long-lived nonce.
CSP report-only mode
Report-only mode helps discover violations before enforcement.
Use it to:
- inventory real dependencies;
- identify inline scripts;
- find unexpected connections;
- observe third-party behavior;
- refine the policy.
Then move to enforcement intentionally.
CSP and third-party scripts
Third-party scripts expand the policy and trust surface.
For each script, ask:
- what data can it read?
- what can it send?
- what happens if it changes?
- can it be removed or isolated?
- is its origin and integrity controlled?
Allowing a script is granting code execution in the page’s origin.
Trusted Types
Trusted Types can require dangerous DOM sinks to receive approved trusted values rather than arbitrary strings.
They can make unsafe paths harder to reach accidentally.
They are a design constraint and enforcement layer, not a substitute for understanding the content policy.
Trusted Types are not sanitization by themselves
A policy can create trusted HTML only after applying a reviewed sanitizer or construction rule.
The policy is where the security decision lives.
CSRF and XSS are different
flowchart TD
subgraph XSS_Threat["Cross-Site Scripting (XSS)"]
X1["Attacker injects malicious script into trusted origin"]
X2["Script runs with full DOM access: reads tokens, steals cookies, logs keystrokes"]
end
subgraph CSRF_Threat["Cross-Site Request Forgery (CSRF)"]
C1["Attacker tricks authenticated browser into issuing request to target origin"]
C2["Browser automatically attaches ambient credentials (cookies)"]
C3["Attacker cannot read response, but executes unauthorized side-effects"]
endThey can interact, but the defenses and threat paths differ.
CSRF tokens
For cookie-authenticated mutations, a server can require a token that an attacker site cannot read.
Validate the token on the server for state-changing operations.
CSRF applies to state-changing requests
Protect operations such as:
- create;
- update;
- delete;
- change email;
- change password;
- transfer or purchase.
Do not rely on a request method name alone; define which operations change state.
SameSite cookies reduce cross-site sending
flowchart TD
subgraph SameSiteDirectives["SameSite Cookie Attribute Policies"]
ST["SameSite=Strict
Cookie NEVER sent on cross-site requests
(Even clicking an external link to the portal)"]
LX["SameSite=Lax (Modern Browser Default)
Cookie sent on top-level safe GET navigations
Blocked on cross-site POST / PUT / fetch mutations"]
NN["SameSite=None; Secure
Cookie sent across all cross-site requests (Requires HTTPS)
High CSRF exposure without explicit tokens"]
endSameSite is valuable defense in depth, but consider legacy behavior, integrations, and the operation’s risk.
Origin and Referer checks
For sensitive mutations, the server may verify request metadata such as Origin or Referer according to a documented policy.
These checks complement, rather than replace, appropriate session and CSRF design.
Cookie authentication
Cookies can keep session credentials out of JavaScript.
They still require:
- CSRF consideration;
- expiration and rotation;
- scope control;
- logout and revocation behavior.
HttpOnly
HttpOnly prevents JavaScript from reading the cookie value.
It can reduce token theft through direct storage access.
It does not prevent XSS code from making requests as the user while the page is open.
Secure
Secure tells the browser to send the cookie only over HTTPS.
It protects transport confidentiality for the cookie, but does not solve application authorization or compromised client code.
Cookie Domain and Path
Narrow scope where possible.
Broad domain cookies increase the set of subdomains and applications that participate in the session boundary.
Path is routing scope, not a complete security boundary for all cookie behavior.
Session fixation and rotation
Rotate session identity when privilege changes, such as after login.
This prevents an attacker from setting or learning a session identifier that remains valid after authentication.
Expire and revoke sessions according to risk and product requirements.
Logout is a security transition
Logout may need to:
- revoke or expire the server session;
- clear client state;
- clear private caches;
- stop live connections;
- remove pending user-specific data;
- prevent back-navigation from revealing sensitive content.
It is more than hiding a button.
Authentication versus authorization
The front end can reflect permissions.
The server must enforce them.
Front-end authorization is not security enforcement
This improves user experience.
It does not protect the delete endpoint if an attacker sends the request directly.
Every sensitive server operation must enforce authorization independently.
Permissions should come from an authoritative model
Avoid deriving critical permissions solely from:
- hidden UI controls;
- route names;
- client booleans;
- decoded but unverified payloads;
- stale local storage.
The server should return or enforce a permission model the client can safely present.
Bearer tokens
Anyone who possesses a bearer token may use it within its scope and lifetime.
Protect issuance, storage, transport, scope, expiry, and revocation.
Browser token storage is a trade-off
flowchart TD
subgraph LocalStorageOption["Option A: localStorage / Memory Bearer Token"]
L1["Readable by JavaScript in same origin"]
L2["Immune to CSRF (not ambiently sent)"]
L3["CRITICAL RISK: A single XSS flaw exposes token to theft!"]
end
subgraph HttpOnlyCookieOption["Option B: HttpOnly, Secure, SameSite Cookie"]
C1["Completely inaccessible to JavaScript (XSS cannot steal)"]
C2["Ambiently sent by browser on matching requests"]
C3["DEFENSE REQUIRED: Enforce SameSite=Lax + Anti-CSRF Token headers"]
endThere is no universal slogan that replaces threat modeling and architecture.
Prefer architecture over token folklore
Choose based on:
- XSS exposure;
- CSRF exposure;
- same-origin or cross-origin deployment;
- backend control;
- mobile/native clients;
- refresh and revocation needs;
- third-party integrations.
A BFF can change the browser’s credential boundary substantially.
Backend-for-Frontend topology
flowchart LR
subgraph BrowserZone["Browser Runtime"]
SPA["Single-Page App"]
end
subgraph InternalBoundary["Same-Origin Boundary"]
BFF["Backend-for-Frontend (BFF)"]
end
subgraph SecureBackend["Internal Protected Network"]
IDP["OAuth / OIDC IDP"]
APIs["Microservices / APIs"]
end
SPA <-->|HttpOnly, Secure Cookie| BFF
BFF <-->|Bearer Tokens| APIs
BFF <-->|PKCE Exchange| IDPA BFF acts as an application-specific gateway for the front-end.
Operational capabilities of a BFF
A Backend-for-Frontend can:
- Hold server credentials: keep sensitive API keys and tokens out of browser memory;
- Manage sessions: issue encrypted
HttpOnly,SameSite=Strictcookies to the SPA; - Aggregate responses: combine multiple downstream microservice calls into one tailored payload;
- Perform edge transformations: translate internal protocols without client complexity.
It replaces direct client token management with a hardened same-origin boundary.
OAuth is delegated authorization
OAuth allows a client to obtain access to resources through an authorization server.
It is not automatically a login protocol.
OpenID Connect adds an identity layer for authentication scenarios.
OAuth roles
Keep the roles distinct when reasoning about tokens and trust.
Authorization Code flow
The code is exchanged rather than delivering an access token directly through the browser redirect.
PKCE Phase 1: Authorization request
sequenceDiagram
autonumber
actor User as Citizen / User
participant App as Browser SPA
participant Auth as Authorization Server (IDP)
App->>App: 1. Generate code_verifier (random secret)
App->>App: 2. Compute code_challenge = SHA256(verifier)
App->>Auth: 3. Redirect to /authorize?code_challenge=...
User->>Auth: 4. User logs in & grants consent
Auth-->>App: 5. Redirect with auth_codeThe browser receives an authorization code, not a sensitive token.
PKCE Phase 2: Token redemption & verification
sequenceDiagram
autonumber
participant App as Browser SPA
participant Auth as Authorization Server (IDP)
participant API as Protected API Server
App->>Auth: 1. POST /token with auth_code + code_verifier
Note over Auth: 2. Verifies SHA256(verifier) == code_challenge!<br/>Prevents code interception attacks.
Auth-->>App: 3. Emits Access Token (+ ID Token)
App->>API: 4. GET /api/data with Bearer token
API-->>App: 5. Returns protected resourcePKCE ensures an intercepted code cannot be redeemed without the original secret verifier.
Avoid legacy implicit token delivery
Delivering tokens directly in a browser redirect has a broader leakage and handling surface.
Use modern authorization code patterns with PKCE where appropriate, following the identity provider and platform guidance.
Browser-based OAuth has its own threat model
Consider:
- redirect interception;
- authorization-code injection;
- open redirects;
- state and nonce validation;
- browser history and referrer leakage;
- token exposure to scripts;
- malicious extensions or compromised dependencies.
The browser is not a confidential client environment.
OpenID Connect
OIDC adds identity claims and an ID token to OAuth-style authorization.
Use it when the application needs to authenticate the user through an identity provider.
Still validate issuer, audience, signature, nonce, time claims, and flow context on the server or trusted verifier.
ID token versus access token
Do not send an ID token to an API as if it were an access token.
Do not use a client-decoded payload as proof of authorization.
JWT is a format, not an architecture
A JWT can be:
- signed or unsigned by a chosen algorithm policy;
- short-lived or long-lived;
- intended for one audience or another;
- used in different trust models.
The string shape does not make the authentication design secure.
Token validation belongs at the resource server
The resource server must validate:
- signature and key policy;
- issuer;
- audience;
- expiry and not-before;
- scopes or permissions;
- token type and context.
The front end may display decoded information, but display is not verification.
Front-end token decoding is not verification
This can read a payload.
It does not prove that the token is authentic, current, intended for this API, or authorized for this action.
Refresh tokens need stronger protection
Refresh tokens can create long-lived access.
Consider:
- whether the browser should receive them;
- rotation and reuse detection;
- secure cookie or BFF storage;
- revocation;
- device and session binding;
- logout behavior.
Do not treat refresh credentials like ordinary UI state.
OAuth state and nonce
Validate both in the correct flow and preserve them through redirects safely.
Redirect URI validation
Authorization servers should require registered redirect URIs.
Avoid:
- wildcard redirect patterns;
- open redirect chains;
- accepting attacker-controlled return URLs;
- mixing trusted and untrusted redirect targets.
Redirect handling is a credential boundary.
Open redirects
Unvalidated redirects can:
- enable phishing;
- leak codes or tokens through chains;
- make trusted links misleading.
Allowlist internal destinations or use opaque server-side state.
Secrets do not belong in front-end bundles
Anything shipped to the browser can be inspected by:
- users;
- browser extensions;
- automated tools;
- copied source maps;
- network observers under the user’s control.
Server credentials and private keys must remain on trusted server infrastructure.
Public API keys are a separate category
Some client identifiers are intentionally public and protected by:
- origin restrictions;
- quotas;
- limited scopes;
- server-side enforcement;
- monitoring.
Calling a value a “public key” does not make every key safe to expose.
Environment variables are not a vault
Use server-side secret management for confidential values.
Review compiled assets and source maps for accidental leakage.
Subresource Integrity
SRI lets the browser verify that a fetched resource matches an expected cryptographic digest.
It is useful for fixed external resources with stable content.
What SRI protects against
SRI can detect a changed resource at the browser boundary.
It does not:
- make a trusted third-party script safe by itself;
- protect inline code;
- validate dynamic resource selection;
- replace CSP or dependency review;
- prevent a trusted script from doing harmful things.
SRI and CORS
Cross-origin integrity checks require the resource to be delivered with appropriate CORS behavior.
Coordinate:
integrityattribute;crossoriginmode;- resource response headers;
- CDN deployment policy.
Dependency supply-chain risk
Dependencies can execute code during:
- installation;
- build;
- development;
- application runtime.
Review packages by capability, maintenance, provenance, update behavior, and access to credentials.
Build-time dependencies can be highly privileged
A build plugin may read:
- source files;
- environment variables;
- CI credentials;
- generated artifacts;
- deployment configuration.
Keep CI secrets scoped and avoid installing unnecessary tooling into privileged environments.
Reduce supply-chain exposure
Use:
- minimal dependencies;
- lockfiles and review;
- trusted registries;
- vulnerability and behavior monitoring;
- restricted scripts where appropriate;
- separate build and deploy credentials;
- reproducible artifacts.
Pinning helps reproducibility but does not eliminate malicious or compromised code.
Third-party scripts are full trust grants
A script executing in the page origin may read:
- DOM content;
- accessible application state;
- non-HttpOnly storage;
- user input;
- API responses visible to the page.
Load third-party code only when its capability and risk are justified.
Clickjacking
Clickjacking tricks a user into interacting with a framed page or disguised control.
Protect sensitive pages with framing policy and deliberate embedding rules.
Do not rely on visual design alone to prevent deceptive framing.
Prevent unauthorized framing
Relevant controls can include:
and appropriate legacy-compatible headers where needed.
Allow only known embedding origins when framing is an actual product requirement.
Browser isolation
Isolation policies help control interactions among browsing contexts and cross-origin resources.
They are useful for high-risk capabilities, cross-origin data, and protection from opener or embedding relationships.
They can also break integrations, so test before enforcement.
COOP
Cross-Origin-Opener-Policy controls whether a document shares a browsing context group with cross-origin documents.
It can reduce window.opener relationships and support stronger isolation.
window.opener
Opening a new window can create an opener relationship.
That relationship can enable unexpected navigation or cross-window interaction.
Use safe link behavior and appropriate opener policy for untrusted destinations.
COEP
Cross-Origin-Embedder-Policy controls whether cross-origin resources can be embedded under the document’s isolation requirements.
It may be needed for cross-origin isolated capabilities.
It can also require every dependency and integration to provide compatible headers.
CORP
Cross-Origin-Resource-Policy lets a resource express which origins may load it in certain cross-origin contexts.
It is a resource-side policy, not the same as CORS or COEP.
COOP, COEP, and cross-origin isolation
Together, appropriate COOP and COEP policies can establish a cross-origin isolated context for specific browser capabilities.
Check:
- workers;
- images and fonts;
- analytics;
- iframes;
- third-party libraries;
- CDN headers.
CORS, CORP, and COEP are different
| Policy | Main question |
|---|---|
| CORS | may browser script read this response? |
| CORP | may this resource be loaded cross-origin in this context? |
| COEP | which embedded resources may this document accept? |
Use the policy that matches the boundary being controlled.
Security headers need testing
Test headers in:
- production-like origins;
- authenticated and unauthenticated flows;
- embedded and popup scenarios;
- worker and asset loading;
- third-party integrations;
- error and redirect paths.
A header that “looks secure” but breaks recovery or silently disables a feature is not a finished design.
Security architecture for a typical SPA
flowchart LR
A["Browser UI Client"] -->|1. Same-Origin Cookie Session| B["BFF / Gateway Server"]
B -->|2. Server-Enforced Role & Scope Authorization| C["Core Business API"]
C -->|3. Validated Database Queries| D[("Municipal PostgreSQL")]
A -.->|NEVER trust client claims for authorization!| CThe front end presents permissions and handles UX.
The server owns enforcement, secrets, and trusted token validation.
Cookie-session SPA
Review:
- HttpOnly, Secure, SameSite;
- CSRF tokens or equivalent controls;
- session rotation;
- logout and cache clearing;
- CORS avoidance through same-origin deployment where practical.
OAuth SPA
Document:
- redirect URIs;
- state and nonce;
- token audience and scope;
- storage and refresh policy;
- logout and revocation;
- server-side validation.
BFF architecture
The BFF can reduce browser token exposure and normalize multiple APIs.
It adds a service to deploy, observe, scale, and secure.
Security UX matters
Good security behavior should tell users:
- what happened;
- whether their data was saved;
- whether they need to sign in;
- whether they lack permission;
- how to recover;
- whether an action is pending or rejected.
Do not reveal sensitive details merely to make an error sound precise.
Error messages must not leak sensitive detail
Avoid exposing:
- whether a private account exists;
- stack traces;
- database structure;
- tokens;
- internal URLs;
- authorization details that aid enumeration.
Log diagnostic context securely and show a useful, bounded user message.
Security and logging
Logs should support investigation without becoming a data leak.
Define:
- what identifiers are safe;
- what must be redacted;
- retention period;
- access controls;
- correlation IDs;
- incident response ownership.
Never log secrets simply because a request failed.
Sensitive data in URLs
URLs can appear in:
- browser history;
- referrer headers;
- server logs;
- analytics;
- screenshots;
- copied links.
Do not place passwords, access tokens, private records, or sensitive form values in query strings or fragments without a very deliberate design.
postMessage is an explicit cross-origin channel
Use an exact target origin where possible.
Treat both outgoing and incoming messages as untrusted protocol data.
Validate postMessage origin and shape
Check origin, source window when relevant, message type, and payload schema.
Iframes are security boundaries
For an iframe, decide:
- which origin it uses;
- whether it needs sandboxing;
- which capabilities it receives;
- whether it may navigate or submit forms;
- how it communicates;
- who may frame your application.
Embedding is an architecture decision, not only a layout choice.
Framework escaping is good - but not sufficient
React, Vue, and other frameworks make common text rendering safer by default.
They cannot decide:
- whether a URL is allowed;
- whether rich HTML should be sanitized;
- whether a third-party script is trustworthy;
- whether an API call is authorized;
- whether a token belongs in the browser.
Framework safety is one layer in a larger boundary model.
Security review checklist: inputs and rendering
Trace every untrusted source to its eventual sink.
Security review checklist: requests and credentials
Ask what happens when the request is replayed, delayed, cross-origin, or sent by a non-browser client.
Security review checklist: browser policies
Each policy should have:
- a threat it addresses;
- a scope;
- an owner;
- tested integrations;
- a failure and rollout plan.
Security review checklist: authentication and authorization
Verify:
- identity is established through a supported flow;
- tokens are validated by the correct server;
- permissions are enforced server-side;
- client UI reflects but does not enforce authority;
- logout clears or revokes relevant state;
- redirects and callbacks are constrained.
A layered security model
No layer is perfect.
The value comes from reducing the impact when one assumption fails.
Practical lab: Secure a Front-End Application Boundary
Review and harden a small application boundary involving untrusted input, cross-origin requests, cookies, OAuth-style redirects, and dangerous sinks.
The practical turns each security concept into an observable browser behavior and documented design decision.
Practical stages 1–4: origins and CORS
- Map the origins.
- Observe a CORS failure.
- Add explicit CORS permission.
- Trigger a preflight.
Distinguish browser read permission from server authentication and authorization.
Practical stages 5–10: rendering defenses
- Demonstrate safe text rendering.
- Identify a dangerous sink.
- Add rich-text sanitization.
- Add CSP in report-only mode.
- Enforce a basic CSP.
- Explore Trusted Types.
Document which content is text, which is approved rich HTML, and which sinks remain intentionally available.
Practical stages 11–14: cookies and permissions
- Implement a cookie session.
- Add CSRF protection.
- Experiment with SameSite.
- Separate authentication and authorization.
Verification: hiding a control does not grant permission, and cookie-authenticated mutations have an explicit CSRF decision.
Practical stages 15–18: identity architecture
- Draw a BFF architecture.
- Model OAuth authorization code plus PKCE.
- Compare ID and access tokens.
- Search the bundle for secrets.
Keep token validation and confidential credentials on the appropriate server boundary.
Practical stages 19–24: browser isolation
- Add SRI to a fixed external resource.
- Inventory third-party trust.
- Secure
postMessage. - Explore framing policy.
- Diagram COOP, COEP, and CORP.
- Build the security architecture map.
Document affected integrations and failure behavior for every security change.
Practical extension: compare session and bearer designs
Compare:
Evaluate XSS, CSRF, token exposure, refresh, logout, cross-origin APIs, and operational complexity.
Avoid declaring one architecture universally correct.
Try this yourself
Trace one value from:
Then repeat for:
Mark where validation, encoding, authentication, authorization, and logging occur.
Troubleshooting guide (Part 1)
| Symptom | Likely cause |
|---|---|
| CORS fails in production only | Development proxy hid real origins |
| Browser blocks a request but curl succeeds | CORS governs browser reads, not server access |
| Escaped text still creates a dangerous link | URL context was not validated |
| CSP breaks analytics or workers | Dependencies were not inventoried before enforcement |
| CSRF token is missing on a mutation | Cookie session and request protection were not designed together |
Troubleshooting guide (Part 2)
| Symptom | Likely cause |
|---|---|
| Hidden button is treated as authorization | Server enforcement is missing |
| Token payload looks valid | Decoding is not signature or audience validation |
| Logout leaves private data visible | Client caches, connections, or local storage were not cleared |
| iframe integration breaks after isolation headers | Policy was copied without testing dependencies |
Completion checklist
- origins and cross-origin relationships are documented;
- CORS is treated as browser read permission, not server security;
- untrusted inputs and dangerous sinks are mapped;
- text rendering is preferred over HTML;
- rich HTML has a reviewed sanitization policy;
- CSP and Trusted Types are defense in depth;
- cookie mutations have CSRF protection decisions;
- authentication and authorization are separate;
- tokens and secrets stay within appropriate boundaries;
- third-party and browser-isolation policies are tested.
Misconceptions to leave behind (Part 1)
| Misconception | Better mental model |
|---|---|
| CORS protects the API from attackers | It controls browser script read access |
| CORS failure means the server did not receive the request | Browser policy and server receipt differ |
XSS means only <script> tags | Any attacker-controlled executable context can matter |
| Framework escaping makes XSS impossible | Escape hatches and context-specific sinks remain |
| Sanitization and encoding are the same | They solve different content problems |
| CSP prevents XSS by itself | It is defense in depth |
| CSRF and XSS are the same | They exploit different trust paths |
Misconceptions to leave behind (Part 2)
| Misconception | Better mental model |
|---|---|
| HttpOnly prevents XSS | It limits direct cookie reads, not same-origin actions |
| Hidden controls enforce permission | Server authorization must enforce it |
| JWT means secure authentication | JWT is a format, not an architecture |
| ID and access tokens are interchangeable | They serve different audiences and purposes |
.env values are secret | Client-exposed build values are public |
| Lockfiles eliminate supply-chain risk | They improve reproducibility, not trust |
| COOP, COEP, and CORP are interchangeable | Each controls a different browser boundary |
The chapter in one sentence
Secure the browser application by tracing trust across origins, content, requests, identity, dependencies, and embedded contexts - and enforce each decision at the authoritative boundary.
Next: Chapter 14
The next chapter will build on secure architecture with:
- front-end testing strategy;
- confidence boundaries and test levels;
- behavior, integration, and browser verification;
- performance and resilience testing;
- quality workflows for production systems.
Questions
Which value in your application crosses the most trust boundaries, and which layer currently makes the security decision about it?
14 - Scaling Front-End Architecture: Design Systems, Monorepos & Micro-Frontends
Scaling Front-End Architecture: Design Systems, Monorepos & Micro-Frontends
Scale the boundaries, not the confusion
Chapter 14
Polla Fattah
Today’s goal
Understand how front-end architecture changes when products, packages, teams, and deployments grow.
We will connect:
- reusable components and design systems;
- tokens, themes, accessibility, documentation, and versioning;
- package APIs, ownership, workspaces, and monorepos;
- task graphs, affected builds, and caching;
- platform teams and golden paths;
- micro-frontend composition and runtime independence;
- Module Federation, contracts, failure containment, and observability;
- migration strategy and architecture decisions at scale.
By the end of today you can
- distinguish a component library from a design system;
- define raw and semantic design tokens;
- design stable package APIs and release policies;
- decide when a monorepo or separate repositories help;
- model dependency direction and affected builds;
- explain why micro-frontends are organizational and deployment architecture;
- choose server, client, build-time, or runtime composition;
- design contracts, failure boundaries, and observability;
- evaluate Module Federation trade-offs;
- recognize when a modular monolith is the better answer.
The central principle
Scale through explicit contracts, ownership, and failure boundaries - not through maximum sharing or maximum distribution.
Every shared abstraction creates a coordination obligation.
Every independently deployed unit creates an integration obligation.
Scale changes the nature of front-end problems
At small scale, a change may affect one application.
At larger scale, the same change may affect:
- several products;
- multiple teams;
- package consumers;
- release trains;
- shared tokens and accessibility;
- deployment and runtime contracts.
The architecture must make change impact visible.
Technical scale and organizational scale differ
A micro-frontend does not solve a team-ownership problem automatically.
A monorepo does not solve a dependency-design problem automatically.
Reuse is not automatically good
Reuse can reduce:
- duplicated fixes;
- inconsistent behavior;
- accessibility defects;
- visual drift.
Reuse can also create:
- shared release coupling;
- generic APIs;
- cross-domain assumptions;
- difficult migration;
- a global dependency nobody can change.
Reuse is a decision with a cost.
Prefer stable shared concepts
Good shared concepts often include:
Be careful sharing a component whose behavior is actually product-specific:
Share the stable concept, not the current visual coincidence.
Component library versus design system
A design system includes how people choose, use, test, document, version, and evolve the components.
A design system is a product
It has:
- users: product teams and customers;
- a roadmap;
- documentation;
- support expectations;
- release notes;
- migration paths;
- quality and accessibility goals;
- adoption feedback.
Treating it as “just a package” underestimates its operational work.
Design-system architecture
flowchart TD
subgraph DesignSystemPyramid["The Design System Architecture Pyramid"]
L1["1. Foundations: Color theory, typography scales, spacing grids"]
L2["2. Design Tokens: Global raw tokens & semantic intent tokens"]
L3["3. UI Primitives: Accessible Headless/Compound components (Button, Modal, Input)"]
L4["4. Composed Patterns: Form layouts, data tables, alert banners"]
L5["5. Documentation & Guidelines: Usage examples, do's/don'ts, accessibility notes"]
L6["6. Governance & Versioning: SemVer release policies, deprecation lifecycles"]
L1 --> L2 --> L3 --> L4 --> L5 --> L6
endThe layers should have intentional dependency direction.
Foundations define shared vocabulary
Foundations may include:
- color and contrast;
- typography;
- spacing;
- elevation;
- motion;
- layout;
- iconography;
- interaction states.
They should support product consistency without hiding meaningful product differences.
Design tokens are named decisions
Tokens make recurring design decisions inspectable and transformable across platforms.
Raw and semantic tokens
Raw tokens describe values.
Semantic tokens describe meaning and can change when a theme or brand changes.
Why semantic tokens matter
Components consume meaning:
The component need not know which palette value currently represents the primary action.
Token interoperability
Tokens may be consumed by:
- CSS variables;
- JavaScript or TypeScript;
- native mobile platforms;
- design tools;
- documentation and testing tools.
Define naming, units, themes, fallbacks, and transformation rules so the token system remains coherent across outputs.
A token pipeline
flowchart LR
A["Source Tokens (JSON / W3C Format)"] --> B["Token Transformer (Style Dictionary)"]
B --> C1["CSS Custom Properties (:root { --color-action: #... })"]
B --> C2["TypeScript Types & Constants"]
B --> C3["Figma / Design Tool Sync"]The pipeline should preserve semantic meaning, not only copy strings.
CSS variables provide runtime tokens
Runtime variables enable themes without rebuilding every component.
Tokens are not a complete design system
Tokens do not define:
- keyboard behavior;
- accessible names;
- component states;
- content guidance;
- composition rules;
- error behavior;
- ownership or release policy.
They are a foundation, not the whole system.
Component API stability matters more at scale
Every public prop, event, slot, CSS hook, and token name can have many consumers.
Before changing it, ask:
- who uses it?
- what behavior is relied upon?
- is the change additive?
- what migration path exists?
- can the old API be deprecated safely?
Public versus internal component APIs
Do not expose every internal component just because it exists in the repository.
Small public surfaces are easier to keep reliable.
Escape hatches need boundaries
Examples:
Escape hatches support real variation, but can undermine consistency and accessibility if unbounded.
Document what the consumer becomes responsible for when using one.
Headless components at scale
Headless components can centralize:
- keyboard behavior;
- focus management;
- ARIA relationships;
- selection state;
- interaction lifecycle.
The consuming product controls presentation while inheriting a tested behavior contract.
Accessibility is part of the component contract
For a dialog, tabs, menu, or combobox, the contract includes:
- roles;
- names;
- keyboard behavior;
- focus movement;
- disabled states;
- announcements;
- escape and dismissal rules.
Visual similarity is not enough for a shared component.
Design-system documentation is an interface
Documentation should show:
- when to use the component;
- when not to use it;
- accessible behavior;
- controlled and uncontrolled modes;
- composition examples;
- supported variants;
- migration and deprecation notes.
Examples are part of the package’s practical API.
Component explorers create feedback loops
A component explorer can provide:
- isolated rendering;
- visual examples;
- interaction states;
- accessibility checks;
- visual regression targets;
- documentation source.
Keep examples representative of supported contracts rather than every possible boolean combination.
Design-system versioning
Versioning must cover more than TypeScript signatures.
A release may change:
- visual output;
- keyboard behavior;
- ARIA structure;
- CSS variables;
- token meaning;
- generated assets;
- browser support.
Consumers need clear impact information.
Semantic versioning is a communication contract
Visual or accessibility changes can be breaking even when the TypeScript type still compiles.
Visual changes can be breaking changes
A “small” change can break:
- layout assumptions;
- screenshot baselines;
- contrast;
- click targets;
- responsive behavior;
- user workflows.
Describe visual and interaction impact in release notes.
Deprecation is a lifecycle
Leaving deprecated APIs forever increases package complexity and makes the preferred path unclear.
Codemods make migrations repeatable
A codemod can transform predictable usage:
Use codemods where the transformation is mechanically reliable, then review semantic edge cases.
Release notes and migration guides
Good notes explain:
- what changed;
- why it changed;
- who is affected;
- what to search for;
- how to migrate;
- how to verify behavior;
- when removal occurs.
Migration work is part of the design-system product.
Shared packages should represent responsibility
Do not make a shared package the default home for anything that might be reused.
Package boundaries are stronger than folders
A package boundary can enforce:
- public exports;
- dependency direction;
- versioning;
- ownership;
- build and test scope;
- consumer expectations.
Folders organize files. Packages communicate architecture.
Avoid the “shared everything” package
One giant shared package creates:
- unclear ownership;
- broad rebuilds;
- accidental imports;
- difficult releases;
- hidden domain coupling.
Split by stable responsibility, not by arbitrary file count.
Shared domain packages need caution
Sharing a domain package can reduce duplicated contracts.
It can also force products to adopt one domain model before their workflows are truly aligned.
Share stable concepts and schemas; keep product-specific orchestration local.
Shared types can create false confidence
Sharing a TypeScript type does not make two independently deployed systems compatible at runtime.
Add versioned contracts, validation, and compatibility tests where needed.
Internal package dependency direction
flowchart TD
Apps["Applications Layer (apps/citizen-portal, apps/inspector-app)"] --> Feat["Feature Packages (packages/feature-licensing)"]
Feat --> Domain["Domain Logic & Models (packages/domain-permits)"]
Domain --> UI["Design System UI (packages/ui-primitives)"]
UI --> Tokens["Foundations & Tokens (packages/design-tokens)"]If a shared package imports an application feature, the boundary has reversed.
Make direction enforceable with lint rules, package exports, or graph checks.
Workspaces coordinate packages
Workspaces can provide:
- local package linking;
- shared install and lockfile;
- cross-package scripts;
- consistent tooling;
- coordinated local development.
They do not decide whether packages are well-designed.
Monorepo benefits
These benefits are strongest when projects genuinely evolve together and ownership is clear.
Monorepo costs
Expect:
- larger graph and CI coordination;
- build ordering;
- ownership disputes;
- broad blast radius;
- package version and release decisions;
- accidental internal imports.
A monorepo makes coupling visible; it does not make coupling disappear.
Monorepo versus polyrepo
Neither is universally superior.
Choose based on release coordination, ownership, dependency contracts, and operational capacity.
Monorepo is not micro-frontend architecture
A monorepo can contain one modular application or many separately deployed front ends.
Micro-frontends can be built across one or many repositories.
Task graphs expose build relationships
flowchart LR
subgraph PipelineGraph["Monorepo Directed Acyclic Task Graph (DAG)"]
T["tokens:build"] --> U["ui:build"]
U --> S["storefront:build"]
D["domain:build"] --> A["admin:build"]
endA task graph can determine:
- build order;
- affected tests;
- cacheable work;
- parallel execution;
- what a change can impact.
Affected builds reduce unnecessary work
If only documentation changes, a full application rebuild may be unnecessary.
If tokens change, many consumers may be affected.
The graph should be accurate enough to make selective verification trustworthy.
Caching build tasks
Cache tasks when their outputs are determined by:
- source inputs;
- dependency versions;
- tool versions;
- environment inputs;
- configuration.
Incorrect cache keys are worse than no cache because they create false confidence.
Remote build cache
A remote cache can share verified task outputs across developers and CI.
Protect it with:
- correct input hashing;
- access control;
- artifact integrity;
- retention policy;
- invalidation rules.
Speed must not weaken reproducibility or trust.
Repository ownership
Ownership should answer:
- who maintains the package;
- who reviews changes;
- who handles releases;
- who responds to failures;
- who approves breaking changes.
Ownership is coordination infrastructure, not a reason to prevent useful contributions.
Federated contribution
A platform team can provide:
- foundations;
- paved workflows;
- documentation;
- release automation;
- support.
Product teams should be able to contribute improvements without routing every decision through one gatekeeper.
Platform teams
A platform team should reduce product team friction through:
- reliable defaults;
- clear extension points;
- observability;
- migration support;
- sensible constraints.
The platform is successful when teams can move safely, not when it controls every implementation.
Platform as a product
Treat internal developers as users.
Measure:
- adoption;
- time to first useful feature;
- migration cost;
- support volume;
- build reliability;
- accessibility and quality outcomes.
Build what reduces real product friction.
Paved road versus mandatory road
Use mandatory rules sparingly and explain the reason.
Too many constraints encourage teams to work around the platform.
Golden path
A golden path can provide:
- starter repository;
- standard scripts;
- secure deployment;
- observability;
- supported dependencies;
- examples and migration guides.
It should make the safe path convenient without making every product identical.
Shared infrastructure must remain replaceable
Avoid APIs that expose every implementation detail of:
- the build tool;
- the deployment platform;
- the state library;
- the rendering runtime.
Stable capabilities allow infrastructure to evolve without forcing every product to migrate at once.
Versioning internal packages
Possible models:
Choose according to release coordination and compatibility, not habit.
Lockstep versioning
All packages release together.
Benefits:
- coordinated compatibility;
- simple version alignment;
- atomic platform upgrades.
Costs:
- unrelated consumers receive synchronized changes;
- releases may become slower or noisier.
Independent versioning
Packages release on their own cadence.
Benefits:
- smaller consumer updates;
- local ownership;
- separate release urgency.
Costs:
- compatibility matrix;
- upgrade coordination;
- more release automation.
Source consumption in monorepos
Consuming internal source directly can speed local development and preserve one graph.
It can also blur package boundaries and make production behavior differ from published consumption.
Decide whether the package boundary is real and test the intended production mode.
Micro-frontends are not small components
Micro-frontends involve:
- team ownership;
- deployment lifecycle;
- runtime or route composition;
- integration contracts;
- failure and observability boundaries.
Why micro-frontends are considered
They may help when:
- teams need independent deployment;
- domains have different release cadence;
- a platform must integrate separately owned products;
- migration from a legacy application must be incremental;
- organizational boundaries are stable and meaningful.
They introduce coordination costs of their own.
Vertical slicing
Vertical slices can align technical ownership with user-facing capability better than splitting by framework layer.
Route-level composition
flowchart TD
Gateway["Edge Reverse Proxy / Gateway (Cloudflare / NGINX)"]
Gateway -->|/catalogue/*| App1["Catalogue Micro-App (Next.js / Team Commerce)"]
Gateway -->|/reports/*| App2["Audit & Reports Micro-App (Vite / Team Analytics)"]
Gateway -->|/settings/*| App3["Citizen Account Micro-App (Remix / Team Identity)"]Full-page or route-level composition is often simpler than embedding several runtimes into one screen.
It preserves clear navigation and failure boundaries.
Server-side composition
The server can assemble fragments or route outputs before delivery.
Benefits:
- clear browser payload;
- centralized navigation and auth;
- independent backend ownership.
Costs include server orchestration, shared layout contracts, and failure handling.
Client-side composition
Client composition can allow independent deployment and rich in-page integration.
It must manage loading, version compatibility, failure, routing, styling, and shared state.
Build-time composition
A shell can consume packages or generated artifacts during its build.
This gives strong integration testing and simple runtime behavior.
It reduces deployment independence because the shell must rebuild to receive changes.
Runtime independence is a spectrum
flowchart LR
A["1. Shared Git Monorepo"] --> B["2. Versioned npm Package"]
B --> C["3. Build-Time Module Split"]
C --> D["4. Runtime Module Federation"]
D --> E["5. Sandboxed <iframe>"]Moving right can increase deployment independence and isolation while increasing integration complexity.
The micro-frontend shell
The shell often owns:
- global navigation;
- authentication context;
- layout;
- routing;
- loading and failure boundaries;
- observability;
- shared design-system contract.
Avoid making the shell a hidden global store for every domain.
Shared state across micro-frontends
Prefer explicit contracts:
Avoid sharing mutable in-memory domain state unless the integration truly requires it and versioning is controlled.
Prefer server and URL contracts
Server APIs and URLs are durable boundaries.
They can survive:
- different framework versions;
- independent deployments;
- full-page navigation;
- process restarts;
- separate repositories.
In-memory object sharing is fast but fragile across runtime boundaries.
Custom events can cross local boundaries
Define:
- event names;
- payload schema;
- versioning;
- ownership;
- duplicate behavior;
- failure behavior.
Events should be protocol contracts, not accidental DOM messages.
Micro-frontend contract design
A contract may include:
- mount inputs;
- lifecycle methods;
- events;
- URL behavior;
- auth assumptions;
- CSS scope;
- data and error models;
- compatibility window.
Write the contract before choosing the integration technology.
CSS isolation
Options include:
No option is free. Test global resets, typography, overlays, z-index, and responsive behavior across boundaries.
The design system is the visual contract
Across micro-frontends, a shared design system can align:
- tokens;
- controls;
- accessibility;
- spacing and typography;
- interaction states.
It does not make domain workflows identical.
Design-system version drift
Independent applications may consume different versions.
Plan for:
- compatible token evolution;
- deprecation windows;
- visual regression;
- accessibility review;
- migration ownership;
- runtime and bundle duplication.
Independent deployment does not remove visual coordination.
Shared runtime dependencies
Sharing a framework or library can reduce duplicate downloads.
It can also create:
- version coupling;
- singleton assumptions;
- incompatible runtime copies;
- upgrade coordination.
Measure the cost of duplication against the cost of shared runtime coupling.
Module Federation: host and remote
The mechanism can enable runtime module loading.
It does not automatically solve contracts, security, routing, CSS, or failure.
Runtime module loading
sequenceDiagram
autonumber
actor User as Citizen / Inspector
participant Host as Shell Host Application (port 3000)
participant Remote as Remote Micro-App (port 3001)
User->>Host: Navigates to /licensing/audit
Note over Host: Host encounters dynamic import('licensingRemote/AuditPanel')
Host->>Remote: Fetch remoteEntry.js (Manifest of exposed modules & shared singletons)
Remote-->>Host: Returns module federation metadata
Note over Host: Host verifies shared React version (18.3.1 === 18.3.1)
Re-uses host React runtime without downloading second copy!
Host->>Remote: Fetch AuditPanel.[hash].js
Remote-->>Host: Emitted chunk code
Note over Host: Host mounts Remote Component into shell DOM tree!Every step can fail or become incompatible.
Represent loading, timeout, version, and fallback states in the host.
Runtime failure is normal
A remote can fail because of:
- deployment outage;
- network failure;
- incompatible API;
- missing asset;
- authentication issue;
- broken initialization.
The shell should preserve useful navigation and explain which capability is unavailable.
Shared dependencies and singleton risk
Sharing one runtime instance can reduce duplication.
It can also make independent remotes depend on:
- one framework version;
- one context implementation;
- one runtime initialization;
- one upgrade schedule.
Singletons are a coordination contract.
Deployment independence versus runtime coupling
The more a remote relies on host internals, the less independent it truly is.
Measure independence by the contracts a team can change without synchronized deployment.
Micro-frontend observability
Trace:
- host route and remote identity;
- remote version;
- load time and failure;
- mount and unmount;
- network and API errors;
- user journey across boundaries.
Without cross-boundary correlation, failures become “the page is broken” with no owner.
Micro-frontend testing
Use several levels:
No single level proves that independent deployment remains safe.
Micro-frontend deployment
Deployment should define:
- version publication;
- manifest or entry discovery;
- rollback;
- cache invalidation;
- compatibility support window;
- monitoring and alert ownership.
The deployment system is part of the runtime contract.
Rollback
Runtime composition should make it possible to:
- revert a remote version;
- disable a capability;
- serve a fallback;
- keep the host usable;
- identify affected users and routes.
Fast rollback is more valuable than theoretical independence.
Canary releases
Canaries can expose a new remote to:
- selected users;
- one region;
- one route;
- a small traffic percentage.
Monitor errors, load time, contract failures, and user outcomes before broad release.
Micro-frontend security
Ask whether a remote:
- shares the same origin;
- can read host DOM or storage;
- receives auth context;
- can navigate the shell;
- loads third-party code;
- needs iframe isolation;
- can affect the whole page.
Deployment independence is not security isolation.
Same-origin micro-frontends share trust
If multiple applications run in the same origin and page context, a compromised remote may have broad access to the host environment.
Use explicit trust review, CSP, dependency policy, and isolation where needed.
Iframe isolation
Iframes can provide stronger origin and DOM isolation.
They add costs:
- sizing and layout;
- navigation and focus;
- communication protocol;
- authentication;
- accessibility;
- performance;
- mobile behavior.
Use them when isolation is worth the integration cost.
Independent frameworks are not automatically better
Different teams may choose different frameworks.
That can support local autonomy but creates:
- duplicate runtimes;
- inconsistent accessibility;
- design drift;
- larger bundles;
- harder shared behavior.
Use multiple frameworks only when the boundary and trade-off are justified.
Conway’s Law effect
Systems tend to reflect the communication structures of the organization.
If teams own stable business capabilities, vertical architecture may align well.
If teams are split around temporary technical layers, the architecture may reproduce handoffs and integration friction.
Team topology and front-end architecture
Consider:
- stream-aligned product teams;
- platform teams;
- enabling teams;
- ownership of shared contracts;
- on-call and incident response.
Architecture should support how teams actually coordinate and change the product.
When a monolith is better
A modular monolith may be preferable when:
- one team can coordinate releases;
- shared state and navigation are central;
- deployment independence has little value;
- runtime composition would add more failure modes than it removes.
“Monolith” does not mean “unstructured”.
Modular monolith first
This can provide many architectural benefits before runtime distribution is necessary.
Signs micro-frontends may be justified
Consider them when:
- teams need independent release cadence;
- domains have stable boundaries;
- the shell can define durable contracts;
- failure containment has value;
- migration needs incremental replacement;
- observability and platform support exist.
Signs micro-frontends are probably unnecessary
Warning signs:
- the only reason is “the app is large”;
- teams need to share mutable state constantly;
- one team owns the whole product;
- independent deployment is not actually required;
- integration testing is weak;
- the same design system and framework are already tightly coupled.
Start with a modular monolith and re-evaluate.
The decision ladder
flowchart TD
subgraph ArchitectureScaleLadder["The Front-End Architectural Scale Ladder"]
L1["Level 1: Local Component (Start here: zero coordination overhead)"]
L2["Level 2: Workspace Package (Internal monorepo library with typed contracts)"]
L3["Level 3: Versioned Design System Package (Published npm library across repos)"]
L4["Level 4: Modular Monolith (Cohesive codebase with strict directory boundaries)"]
L5["Level 5: Route-Based Multi-Zone Apps (Independent deployments partitioned by URL path)"]
L6["Level 6: Runtime Micro-Frontends / Module Federation (High cost: earned only by organizational friction)"]
L1 --> L2 --> L3 --> L4 --> L5 --> L6
endMove right only when the lower level cannot satisfy a real requirement.
Architecture decision: design system
Ask:
- which concepts are stable?
- who are the consumers?
- what accessibility contract is shared?
- how are visual changes reviewed?
- what can remain product-specific?
A design system should reduce recurring decisions without flattening domain differences.
Architecture decision: monorepo
Ask:
- do projects change together?
- are atomic changes valuable?
- can the team operate the graph and CI?
- is ownership clear?
- are package APIs enforceable?
Choose a monorepo for coordination value, not only shared code.
Architecture decision: separate repositories
Separate repositories can provide:
- independent lifecycle;
- stronger ownership boundaries;
- smaller local graphs;
- separate release and access policy.
They require deliberate contract publication, compatibility testing, and cross-repository coordination.
Architecture decision: runtime federation
Ask:
- is independent deployment truly required?
- what happens when a remote is unavailable?
- how are versions compatible?
- who owns the shell contract?
- how are security and observability handled?
Runtime federation should be a response to a specific operational need.
Migration strategy
Large changes are safer when incremental:
Migration architecture should support rollback and coexistence.
The strangler pattern
flowchart TD
Proxy["API Gateway / Edge Router"]
Legacy["Legacy Monolith (JSP / ASP.NET / AngularJS)"]
NewApp["Modern Micro-App (Next.js / Vite)"]
Proxy -->|90% Legacy Traffic| Legacy
Proxy -->|10% Migrated Route: /permits/apply| NewApp
Note over Proxy: Gradually shift routes from Legacy to NewApp until Legacy is completely retired!Use stable URLs, data contracts, and ownership rules to keep the old and new systems coherent during transition.
Shared navigation during migration
Navigation can be a durable contract across applications:
- common route vocabulary;
- active-state behavior;
- permissions;
- full-page versus client navigation;
- analytics and focus.
Do not force all applications into one runtime merely to share a menu.
Full-page navigation can be a feature
Full navigation can provide:
- clean application boundaries;
- independent failure recovery;
- simpler deployment;
- less shared runtime coupling;
- natural browser history.
It is not automatically an outdated experience.
Design-system governance at enterprise scale
Governance should clarify:
- what belongs in foundations;
- what belongs in shared patterns;
- what remains product-specific;
- who approves breaking changes;
- how contributions are accepted;
- how accessibility and visual quality are measured.
Contribution tiers
Promote upward only when usage, stability, and ownership justify it.
Measure design-system value
Possible signals:
- adoption of supported components;
- accessibility defect reduction;
- time to build common workflows;
- migration effort;
- visual consistency;
- consumer satisfaction;
- release and support cost.
Count outcomes, not only number of components.
Measure monorepo health
Useful measures include:
- affected-build accuracy;
- CI duration;
- dependency cycle count;
- package boundary violations;
- ownership response time;
- cache hit rate;
- cross-app change reliability.
Repository size alone says little about health.
Measure micro-frontend health
Watch:
- remote load success;
- fallback frequency;
- contract failures;
- duplicate runtime cost;
- cross-boundary latency;
- deployment rollback time;
- team independence in practice.
If every change still requires synchronized coordination, the distribution may be mostly ceremonial.
Avoid architecture by organization chart alone
Team boundaries are evidence, not the entire design.
Also examine:
- business capability boundaries;
- data ownership;
- user journeys;
- release frequency;
- failure containment;
- security and performance requirements.
Architecture should reflect durable responsibilities, not temporary reporting lines.
Stable business capabilities
Good boundaries often align with capabilities such as:
They can be understood by users, teams, APIs, and deployment owners.
Contracts over implementation sharing
Two teams do not need to share implementation to coordinate.
They can share:
- URL contracts;
- API schemas;
- events;
- design tokens;
- accessibility expectations;
- versioned interfaces.
Implementation sharing is one option, not the default definition of collaboration.
Build-time safety versus runtime safety
Micro-frontends and independently deployed packages need both.
Compilation cannot prove that a remote URL will be available tomorrow.
Failure containment
A boundary is valuable when a failure can remain local:
flowchart TD
Shell["App Shell Navigation & Header (Healthy)"]
subgraph ViewContainer["Page Content Layout"]
Catalog["Permit Catalogue Widget (Healthy)"]
subgraph ErrorBoundary["React / Vue Error Boundary"]
FailedRemote["Remote Analytics Widget (Crash / 500 Network Drop)"]
Fallback["Fallback UI: 'Analytics temporarily unavailable. Retry'"]
end
end
Shell --- Catalog
Shell --- ErrorBoundary
FailedRemote -.->|Error caught by Boundary| Fallback
Note over Shell: Entire application survives! User continues browsing catalogue.If every remote failure breaks the shell, distribution has not created meaningful containment.
Global failure dependencies
Beware sharing:
- one global runtime initialization;
- one mutable store;
- one required remote;
- one unversioned design token package;
- one authentication callback that blocks every route.
These can make many “independent” units fail together.
Versioned contracts
Version the things that cross boundaries:
- package APIs;
- event payloads;
- route parameters;
- remote manifests;
- tokens;
- server schemas.
Compatibility should be testable and documented.
Consumer-driven contract thinking
Consumers should express the behavior they rely on.
Providers can verify that they remain compatible before publishing.
This catches integration breakage earlier than waiting for a full end-to-end environment.
Security boundaries versus team boundaries
Two teams owning separate modules does not isolate their code from each other at runtime.
If security isolation is required, use actual origin, process, sandbox, or iframe boundaries and design the communication protocol accordingly.
Performance budgets across micro-frontends
Each remote may appear small alone but expensive together.
Track:
- initial JavaScript;
- duplicate frameworks;
- remote request count;
- hydration or mount cost;
- CSS and font duplication;
- runtime memory.
Measure the assembled user experience.
Shared performance budget
Define budgets for:
Every team should know how its boundary contributes to the total.
Practical lab: Design-System Package and Ownership Map
Design a small shared UI package without turning every product-specific component into a design-system primitive.
Map tokens, public APIs, ownership, versioning, consumers, and release responsibility.
Practical stages 1–6: foundations and contracts
- Identify repeated foundations, primitives, and domain components.
- Define token ownership and semantic naming.
- Create a small package boundary and public API.
- Document compatibility, versioning, and affected consumers.
- Map team ownership.
- Add an accessibility contract.
Keep domain-specific behavior with the product team.
Practical stages 7–12: package and repository architecture
- Simulate a breaking change.
- Create workspace packages.
- Enforce package public APIs.
- Draw the package dependency graph.
- Add ownership.
- Add shared tooling.
Verification: the shared package has a small stable surface and does not become a hidden global dependency.
Practical stages 13–16: monorepo operations
- Simulate an atomic cross-app change.
- Simulate affected builds.
- Add task caching conceptually.
- Design a platform golden path.
Use the graph to explain what should rebuild, what can be cached, and who reviews each change.
Practical stages 17–22: micro-frontend contracts
- Model route-level app separation.
- Build a micro-frontend decision record.
- Build a small federation demo.
- Add remote failure handling.
- Explore shared dependencies.
- Add an integration contract.
Make runtime loading, version compatibility, fallback, and observability explicit.
Practical stages 23–29: boundary and final architecture
- Avoid shared in-memory domain state.
- Add a custom event.
- Explore CSS collision.
- Use the design system across host and remote.
- Draw the runtime architecture.
- Compare three architectures.
- Create the final scaling ADR.
Compare a modular monolith, route-level separation, and runtime federation.
Practical extension: package versus independently deployed remote
Compare:
Explain why they solve different problems around source sharing, release cadence, runtime failure, ownership, and contracts.
Try this yourself
Take one repeated UI element and classify it:
Write the evidence for the classification: consumers, stability, ownership, accessibility contract, and expected change direction.
Troubleshooting guide (Part 1)
| Symptom | Likely cause |
|---|---|
| Shared package accepts every prop | The concept is not stable or API is too generic |
| Design system contains domain workflows | Product behavior was promoted too early |
| Every package rebuilds for one change | Dependency graph or package boundary is too broad |
| Teams still coordinate every release | Runtime independence is mostly nominal |
| Remote failure breaks the whole shell | Failure containment was not designed |
Troubleshooting guide (Part 2)
| Symptom | Likely cause |
|---|---|
| Different remotes look inconsistent | Token, component, or version governance is weak |
| Micro-frontends share a giant store | Durable contracts were replaced by hidden coupling |
| Monorepo feels slow and opaque | Task graph, affected builds, or ownership is unclear |
| Security assumption follows team ownership | Team boundary is not a runtime isolation boundary |
Completion checklist
- shared concepts are stable and intentionally owned;
- tokens distinguish raw values from semantic meaning;
- public package APIs are small and documented;
- accessibility is part of the shared component contract;
- versioning and deprecation have migration paths;
- package dependency direction is enforceable;
- monorepo tasks and affected builds are observable;
- platform defaults are supported but not needlessly mandatory;
- micro-frontends have durable contracts and failure fallbacks;
- scaling decisions are measured against user and team outcomes.
Misconceptions to leave behind (Part 1)
| Misconception | Better mental model |
|---|---|
| A design system is a component library | It is a product of foundations, behavior, docs, and governance |
| Tokens are just CSS variables | Tokens are semantic cross-platform decisions |
| More abstraction means more reuse | Shared abstractions create coordination cost |
| A monorepo means one application | It is a repository and package organization choice |
| Workspaces and monorepos are identical | Workspaces coordinate packages; architecture needs boundaries |
| Micro-frontends are tiny components | They are independently owned application capabilities |
| Micro-frontends are always more scalable | They trade local autonomy for integration complexity |
Misconceptions to leave behind (Part 2)
| Misconception | Better mental model |
|---|---|
| Independent deployment means no coordination | Contracts, versions, and runtime failures still coordinate |
| Module Federation prevents remotes from breaking hosts | Runtime loading requires compatibility and fallback |
| Shared dependencies are always better | They reduce duplication but increase coupling |
| Different teams require different frameworks | Team boundaries do not automatically justify runtime diversity |
| Micro-frontends provide security isolation | Same-origin code often shares trust |
| Full-page navigation is outdated | It can provide valuable isolation and recovery |
The chapter in one sentence
Scale front-end systems by making shared contracts, package ownership, deployment independence, and failure boundaries explicit - and distribute only when the benefit exceeds the coordination cost.
Next: Chapter 15
The next chapter will build on scaling architecture with:
- performance engineering and budgets;
- measurement and profiling;
- loading, rendering, and interaction cost;
- production observability;
- performance decisions grounded in user journeys.
Questions
Which boundary in your current system is shared because it is genuinely stable, and which one is shared only because it was convenient to put in one package?
15 - Core Web Vitals & Performance Engineering
Core Web Vitals & Performance Engineering
Measure the experience before changing the code
Chapter 15
Polla Fattah
Today’s goal
Treat performance as a user-experience property and an engineering discipline.
We will connect:
- field data, lab data, and representative measurement;
- LCP, INP, and CLS;
- network, JavaScript, rendering, memory, and layout cost;
- images, fonts, lists, workers, caching, and third parties;
- rendering topology and route-transition performance;
- DevTools traces, User Timing, PerformanceObserver, and RUM;
- budgets, regression detection, ownership, and performance culture.
By the end of today you can
- explain what Core Web Vitals measure and what they do not;
- use p75 and segmentation instead of hiding slow users in averages;
- diagnose LCP resource discovery and server delay;
- trace INP through input, main-thread work, rendering, and paint;
- prevent CLS through reserved space and stable layout;
- identify request waterfalls and unnecessary JavaScript;
- optimize large lists and CPU work only when evidence supports it;
- use browser tools and production signals together;
- define route and user-journey budgets;
- turn an optimization into a hypothesis, measurement, and regression guardrail.
The central principle
Performance work begins with a measured user problem, forms a hypothesis about its cause, and ends with a verified improvement under representative conditions.
An optimization is not successful because it sounds sophisticated.
It is successful when a user journey improves without unacceptable trade-offs.
The performance journey
flowchart TD
A["1. User Request (Navigation / URL enter)"] --> B["2. Time to First Byte (TTFB - Server Response)"]
B --> C["3. First Contentful Paint (FCP - Initial Typography/DOM)"]
C --> D["4. Largest Contentful Paint (LCP - Hero Image/Header Rendered)"]
D --> E["5. Interaction to Next Paint (INP - Responsive Main Thread)"]
E --> F["6. Cumulative Layout Shift (CLS - Visual Stability Maintained)"]Performance is a sequence of experiences, not one score.
Core Web Vitals
flowchart LR
subgraph CoreWebVitals["The Three Core Web Vitals (Google Web Standards)"]
LCP["Largest Contentful Paint (LCP)
Target: <= 2.5s (p75)
Measures: Perceived Loading Speed"]
INP["Interaction to Next Paint (INP)
Target: <= 200ms (p75)
Measures: Overall Page Responsiveness"]
CLS["Cumulative Layout Shift (CLS)
Target: <= 0.1 (p75)
Measures: Visual Stability & Jitter"]
endTogether they cover important parts of the user journey.
They are not the complete product experience or business metric.
Why the 75th percentile matters
The worst quarter of experiences still matters.
An excellent average can hide slow devices, difficult networks, or specific routes with real users.
Field data versus lab data
| Field data | Lab data |
|---|---|
| real users and devices | controlled environment |
| shows what happened | helps explain why it may happen |
| segmented by route and conditions | repeatable experiments |
| noisy but representative | limited but diagnostic |
Use both; do not treat them as interchangeable.
Field data answers “what is happening?”
Field data can reveal:
- a slow region or device class;
- a route with poor p75;
- a regression after release;
- real-user impact of a third-party script;
- differences by network, locale, or interaction.
It rarely identifies the exact code line by itself.
Lab data answers “why might it be happening?”
Lab tools can reveal:
- request waterfalls;
- long tasks;
- layout calculations;
- LCP resource delay;
- JavaScript execution;
- memory and rendering traces.
Use controlled experiments to test a specific hypothesis.
Synthetic testing needs representative conditions
Vary:
- CPU speed;
- network latency and bandwidth;
- device memory;
- viewport;
- locale and direction;
- route state;
- cache state.
A single laptop and fast connection do not represent the user population.
Real User Monitoring
RUM can collect:
- Web Vitals;
- route and release;
- device and connection class;
- user journey markers;
- errors and long tasks;
- interaction context.
Collect only what is needed, protect privacy, and make the data actionable.
Segment performance data
Useful segments include:
A global p75 can improve while an important route or region regresses.
Largest Contentful Paint
LCP measures when the largest relevant content element becomes visible in the viewport.
It can be:
- an image;
- a text block;
- a poster or background-like visual;
- another large content element selected by the browser.
The metric is about the critical content experience, not only one file.
LCP is not just image download time
An image can download quickly but still become LCP late because the page discovers or renders it late.
LCP subparts
Diagnose:
flowchart LR
A["1. Time to First Byte
(Server & Network TTFB)"] --> B["2. Resource Load Delay
(Time until browser discovers LCP image)"]
B --> C["3. Resource Load Duration
(Time to download image asset)"]
C --> D["4. Element Render Delay
(Time spent decoding, layout & paint)"]Each subpart suggests a different fix.
Optimizing the wrong subpart can change little.
Time to First Byte
TTFB includes:
- connection setup;
- redirects;
- server processing;
- time until the first response bytes.
Slow TTFB can delay every downstream rendering milestone.
Investigate server, deployment, cache, and network paths before optimizing browser code.
Redirects add delay
Remove unnecessary redirects from critical navigation and asset paths.
Check protocol, host, trailing slash, authentication, and locale redirects.
LCP resource discovery
The browser cannot request what it does not know exists.
An LCP image discovered only after:
- JavaScript execution;
- CSS evaluation;
- a late component mount;
may arrive too late even if its file is optimized.
Prioritize the LCP resource selectively
Use:
- semantic HTML;
- early discoverability;
- correct fetch priority;
- selective preload;
- appropriate caching.
Priority hints are not a substitute for a correct document and should not be applied to every image.
Do not lazy-load above-the-fold LCP images
Lazy loading tells the browser that a resource is not immediately important.
If that resource is the primary visible content, the hint contradicts the user experience.
Lazy-load content that is genuinely below the initial view.
Preload selectively
Preload can help when:
- the resource is certain to be needed;
- discovery is otherwise delayed;
- it is part of the critical path.
It can hurt when it competes with CSS, fonts, scripts, or a different actual LCP resource.
Optimize LCP images
Consider:
- correct dimensions;
- modern format;
- responsive sources;
- compression quality;
- CDN transformation;
- caching;
- avoiding unnecessary oversized files.
Image optimization must preserve visual quality and correct art direction.
Responsive images reduce waste
Send an image appropriate to the viewport and layout rather than the largest available file to every device.
Image dimensions help CLS too
Reserve the rendered aspect ratio:
Known dimensions let the browser allocate space before the image arrives.
LCP text can be delayed by fonts
A text LCP may wait for:
- font discovery;
- font download;
- font blocking behavior;
- style calculation;
- fallback-to-webfont swap.
Treat fonts as part of the critical rendering path when they affect the largest content.
Interaction to Next Paint
INP reflects how quickly the page produces the next visual update after user interactions across a session.
It considers more than the event handler’s own duration.
INP is a lifecycle metric
The slowest meaningful interaction can influence the session metric.
Optimize the user journey, not only one fast demo click.
Anatomy of an interaction
flowchart LR
subgraph INP_Anatomy["Anatomy of an Interaction (INP Breakdown)"]
I1["1. Input Delay
(Queued behind prior main-thread long tasks)"] --> I2["2. Processing Time
(Execution duration of event handlers)"]
I2 --> I3["3. Presentation Delay
(Style recalc, layout, compositing & GPU paint)"]
endAny phase can dominate the perceived response.
Input delay
Input delay occurs when the main thread is busy before the event can be handled.
Possible causes:
- startup JavaScript;
- long tasks;
- synchronous storage;
- parsing or layout work;
- another interaction’s work.
The handler code may be fast while the user still waits.
Main-thread contention
flowchart TD
HTML["HTML Parsing"] & JS["Heavy JavaScript Execution"] & Style["Style Recalculations"] & Input["User Click / Keystroke"] --> MT["The Single Browser Main Thread"]
MT --> Blocked["Long Task (>50ms)
Main thread frozen! User input delayed = High INP!"]Work competes for responsiveness.
Reduce unnecessary work before moving necessary work to more complex mechanisms.
Long tasks
Long tasks block the main thread long enough to delay input and visual updates.
Find:
- the task’s initiator;
- script evaluation;
- event handler work;
- rendering and layout;
- third-party contribution.
Then optimize the cause, not only the symptom.
Yielding
Break long work into opportunities for the browser to process input and paint.
Yielding improves responsiveness only when work can be safely divided and scheduled.
Avoid work before optimizing work
First ask:
- is this calculation needed?
- is it repeated?
- is its input larger than necessary?
- can it happen later?
- can it be moved off the critical path?
Making unnecessary work faster is usually less valuable than removing it.
Event handler work
Keep urgent interaction work small:
Avoid synchronous parsing, large filtering, layout reads, and unrelated analytics in the immediate path.
Rendering can dominate INP
The handler may finish quickly, but a large update can cause:
- broad component rendering;
- large DOM changes;
- style recalculation;
- layout;
- paint.
Trace from the input through the committed UI.
Reduce render scope
Keep state close to the consumers that need it.
Use stable identities and explicit boundaries to avoid updating unrelated parts of the interface.
Do not split components blindly; make the dependency graph narrower.
Virtualize large lists
Virtualization renders only the visible window plus a buffer.
flowchart LR
Data["Dataset in Memory
(10,000 municipal records)"] --> Virtual["Virtual Scroller Window
(Calculates scroll offset)"]
Virtual --> DOM["Lightweight DOM
(Only 35 active rendered <tr> elements!)"]
DOM --> Smooth["60fps Butter-Smooth Scrolling
Zero memory bloat!"]It can reduce rendering and layout cost for genuinely large collections.
Virtualization has UX costs
Consider:
- keyboard navigation;
- find-in-page;
- screen readers;
- variable row height;
- scroll position;
- focus preservation;
- copy and selection behavior.
Virtualize when the trace shows a need and test the full interaction model.
Web Workers move CPU work off the main thread
Workers can help with:
- large parsing;
- data transformation;
- computation;
- search indexing;
- image or file processing.
They do not make an algorithm cheaper and do not remove communication or serialization cost.
Worker communication has cost
sequenceDiagram
autonumber
participant Main as Browser Main Thread (UI at 60fps)
participant Worker as Dedicated Web Worker Thread
Main->>Worker: postMessage({ type: 'CALCULATE_AUDIT', data: largeMatrix })
Note over Worker: Worker executes heavy 400ms calculation in background!
Main thread remains 100% responsive to user clicks.
Worker-->>Main: postMessage({ type: 'AUDIT_COMPLETE', result: summary })
Note over Main: Main thread updates UI badge with zero frame dropsMove enough work to justify the boundary.
Use transferable data or shared strategies deliberately and safely.
Interaction feedback first
The user needs an immediate response such as:
- pressed state;
- focus change;
- optimistic visual update;
- progress indication;
- disabled or pending state.
Defer nonessential work so feedback is not behind logging, formatting, or secondary requests.
Cumulative Layout Shift
CLS measures unexpected layout movement during page use.
Common causes:
- images without dimensions;
- late banners;
- font swaps;
- inserted content;
- changing ads or embeds;
- transitions that move unrelated content.
Layout shift sources
Trace each shift to:
The fix depends on the cause, not only the visual symptom.
Reserve space
Reserve space for:
- media;
- ads or sponsored regions;
- async controls;
- embeds;
- validation messages;
- banners and notifications.
Stable geometry improves reading, focus, and interaction even beyond the metric.
User-initiated changes are different
A layout change directly caused by a user action may not count the same way as an unexpected shift.
It can still be a poor experience if it moves the user away from the control they are using.
Metrics do not replace design judgment.
Avoid inserting content above the view
If a late response inserts a banner above the user’s reading position, the page can jump.
Reserve the region or place updates where they do not displace current content unexpectedly.
Fonts and layout shift
Font changes can alter:
- text width;
- line wrapping;
- block height;
- button size;
- layout position.
Choose fallback metrics and loading behavior that preserve useful geometry.
Font loading strategy
Options include:
- system fonts;
- self-hosted subsets;
- preload for truly critical fonts;
font-displaypolicy;- metric-compatible fallbacks;
- delayed noncritical families.
The fastest font is often the one the critical path does not need.
font-display
Font display controls how text behaves while a webfont loads.
Choose based on:
- content importance;
- brand requirements;
- fallback compatibility;
- readability;
- layout stability.
There is no single setting that is correct for every text region.
System fonts can be excellent
System fonts can provide:
- immediate text;
- no font transfer;
- good platform integration;
- stable performance.
Brand typography should justify its loading and layout cost.
Network performance
The network path includes:
Optimize the path that is actually slow.
Request waterfalls
flowchart LR
A["1. HTML Document
(index.html)"] -->|Downloads & Parses| B["2. Client JS Bundle
(app.js)"]
B -->|Executes & Fetches| C["3. API JSON Data
(/api/permits/104)"]
C -->|Reads Image URL| D["4. LCP Hero Image
(hero.webp - Loaded at last!)"]Sequential discovery delays the final content.
Find dependencies that could be:
- discovered earlier;
- requested in parallel;
- embedded in the initial response;
- cached or prefetched.
Flatten unnecessary waterfalls
Parallelize independent work.
Do not parallelize operations that depend on one another or that would overwhelm the server.
Connection reuse
Reuse can reduce setup cost for:
- HTTP connections;
- TLS;
- pooled server connections;
- persistent sessions.
Avoid unnecessary origins that prevent effective reuse and increase discovery work.
Compression helps transfer, not execution
Compression reduces bytes over the network.
It does not remove:
- JavaScript parsing;
- compilation;
- execution;
- hydration;
- memory;
- DOM work.
Optimize both transfer and device work.
JavaScript is expensive in several ways
A small compressed bundle can still be expensive on a slower device if it contains heavy startup work.
Reduce unnecessary JavaScript
Possible strategies:
- remove unused dependencies;
- delay noncritical features;
- keep static content server-rendered;
- use route and component boundaries;
- avoid shipping server-only logic;
- replace a library with a platform capability when appropriate.
Measure the user journey after each change.
Route-level code splitting
Route splitting aligns code delivery with navigation and often gives a strong return.
Define loading, error, and prefetch behavior for each route.
Component-level lazy loading
Use it for:
- rarely opened dialogs;
- expensive editors;
- below-the-fold visualizations;
- optional reports;
- feature-specific integrations.
Avoid splitting tiny, always-used components into network overhead.
Third-party JavaScript is shared product cost
Third-party code can add:
- transfer;
- CPU;
- layout;
- network requests;
- privacy and security review;
- failure dependencies.
The product owns the cost even when another company wrote the script.
Lazy-load third-party features
Delay analytics, chat, experiments, maps, and widgets when they are not needed for the first useful interaction.
Load them after consent, idle time, visibility, or explicit interaction where appropriate.
CSS performance
CSS affects:
- style calculation;
- layout;
- paint;
- render-blocking behavior;
- asset discovery;
- layout stability.
Keep critical styles available and avoid shipping large unused style systems to every route.
Critical CSS
Critical CSS covers styles needed for the initial visible content.
Possible approaches include:
- inline critical rules;
- route-specific CSS;
- efficient stylesheet delivery;
- deferred noncritical styles.
Avoid making the critical path more complex than the benefit justifies.
Avoid flashing unstyled content
Coordinate:
- stylesheet discovery;
- font behavior;
- class application;
- theme initialization;
- server and client markup.
A fast first paint that changes dramatically is not a stable experience.
DOM size and depth
A large or deeply nested DOM can increase:
- style calculation;
- layout;
- accessibility tree work;
- memory;
- query and update scope.
Do not reduce markup mechanically; identify the subtree causing measured work.
Batch reads and writes
Avoid alternating layout reads and writes:
flowchart TD
subgraph LayoutThrashing["The Layout Thrashing Anti-Pattern"]
R1["Read: elem1.offsetWidth"] --> W1["Write: elem1.style.width = '...'"]
W1 -->|Forces immediate sync layout recalc!| R2["Read: elem2.offsetWidth"]
R2 --> W2["Write: elem2.style.width = '...'"]
W2 -->|Forces second layout recalc!| Bad["Result: 30 dropped frames / massive jank"]
end
subgraph BatchedSolution["The Batched Solution (Fast)"]
BR1["Batch All Reads First:
r1 = elem1.offsetWidth;
r2 = elem2.offsetWidth;"] --> BW1["Batch All Writes in next frame:
elem1.style.width = ...;
elem2.style.width = ...;"]
BW1 --> Good["Result: Exactly ONE clean layout recalculation!"]
endGroup reads, then writes where possible to reduce forced layout and synchronization.
Paint and composite
Visual updates can cost through:
- large paint regions;
- expensive shadows or filters;
- image decoding;
- transparency and layers;
- layout-triggering properties.
Use DevTools traces to identify actual paint and composite costs.
will-change is not a speed button
will-change can reserve resources or create layers.
Use it only for a measured, anticipated change and remove it when the change ends.
Overuse increases memory and can make rendering worse.
Animation frame budget
At a common 60Hz target, a frame arrives roughly every 16.7ms.
That time includes:
- JavaScript;
- style;
- layout;
- paint;
- composite;
- browser overhead.
Animation and interaction work must share the frame budget.
requestAnimationFrame
Use it to coordinate visual writes with the browser’s rendering cycle.
It does not make expensive work free; it schedules it at a meaningful time.
Caching is a performance policy
Define:
- what is immutable;
- what can be stale;
- what is user-specific;
- what invalidates it;
- how long it lives;
- how it behaves offline.
Caching can reduce latency and increase correctness risk if the policy is unclear.
Fingerprinted static assets
Content-addressed assets can be cached for a long time because a new content version receives a new URL.
HTML cache policy differs from hashed assets
HTML points to the current asset names and may need more frequent revalidation.
Caching HTML as aggressively as fingerprinted assets can serve old entry points or stale deployment manifests.
Match cache lifetime to artifact meaning.
API caching
API caching must consider:
- freshness;
- user identity;
- permissions;
- query inputs;
- invalidation after mutation;
- privacy;
- offline usefulness.
The fastest response is not useful if it is the wrong user’s data.
Cache hit rate is not the whole story
Also measure:
- stale response rate;
- revalidation cost;
- memory and storage;
- cache eviction;
- incorrect reuse;
- time to useful content.
A high hit rate with stale or mismatched data is not a success.
Prefetching
Prefetch when a next action is likely and the cost is bounded.
Good signals include:
- visible navigation intent;
- hover or focus with caution;
- route prediction;
- idle time;
- sufficient connection and storage.
Speculation can waste resources
Prefetching and prerendering can consume:
- bandwidth;
- battery;
- memory;
- server capacity;
- privacy budget.
Do not optimize one user’s transition by imposing invisible cost on every user.
Preconnect and DNS prefetch
These hints can reduce connection setup for origins known to be needed.
Use them selectively:
preconnectfor a critical known origin;- DNS prefetch when a connection may be useful later.
Unnecessary origins still add work and complexity.
Performance and rendering topology
The rendering choice changes where performance work appears.
Hydration cost
Measure:
- JavaScript needed for the route;
- components that hydrate;
- event setup;
- client data handoff;
- main-thread time before interaction.
HTML already visible does not mean the page is interactive.
Streaming performance
Streaming can improve:
- first useful output;
- perceived progress;
- independent slow regions.
It can also add:
- boundary complexity;
- layout shifts;
- coordination and error states;
- client handoff work.
Measure the user journey, not only the first chunk.
SPA and soft-navigation performance
After initial load, measure:
- route request time;
- code chunk loading;
- data loading;
- transition feedback;
- rendering and layout;
- scroll and focus restoration.
A fast initial page can still have slow navigation.
Route-transition metrics
Define markers:
Use them to locate the phase that users are waiting through.
User Timing API
Custom marks connect browser traces to product journeys.
Name them consistently and avoid collecting sensitive data.
PerformanceObserver
Observers can collect browser performance entries such as:
- largest contentful paint;
- layout shifts;
- long tasks;
- navigation timing;
- resource timing.
Use them to build useful signals, not to collect every possible event without a question.
DevTools Performance panel
Start with a user action:
The trace should answer a hypothesis about the delay.
Flame charts
Flame charts reveal:
- which functions consume time;
- call depth;
- repeated work;
- long tasks;
- layout and paint boundaries.
Read from the interaction or milestone rather than scanning for the largest colorful block without context.
Bottom-up analysis
Bottom-up views aggregate cost by function or category.
They can reveal:
- a dependency’s shared cost;
- repeated parsing;
- expensive event handlers;
- a framework operation called from many paths.
Use call stacks and user journeys to interpret the total.
Network panel
Inspect:
- request order;
- queueing;
- connection setup;
- response time;
- transfer size;
- cache status;
- priority;
- initiator chain.
Network evidence often explains why an apparently optimized asset still arrives late.
Coverage and bundle analysis
Coverage can show unused code during one journey.
Bundle analysis can show:
- large dependencies;
- duplicate packages;
- shared chunk cost;
- route chunk composition;
- accidental client inclusion.
Unused in one journey is a clue, not automatic proof that code should be removed.
Lighthouse is one lab view
Lighthouse can provide useful repeatable diagnostics.
It is not:
- real-user data;
- a complete accessibility audit;
- a business metric;
- proof that every route is healthy;
- a substitute for traces and production monitoring.
Use it as one instrument in the measurement stack.
Repeat measurements
Control:
- cache state;
- CPU and network;
- viewport;
- route data;
- browser version;
- device conditions.
Compare enough runs to distinguish signal from noise.
CPU and network throttling
Representative throttling helps reveal:
- startup execution;
- input delay;
- slow resource discovery;
- dependency waterfalls;
- layout and paint pressure.
Do not mistake one artificial throttle for the entire user population.
Memory performance
Watch for:
- detached DOM nodes;
- unbounded caches;
- retained event listeners;
- subscriptions without cleanup;
- large closures;
- worker and image memory.
Memory problems can become interaction and crash problems later.
Memory leaks
A leak is retained data that should no longer be reachable.
Reproduce a lifecycle repeatedly:
Compare heap snapshots and retained references rather than relying on one memory reading.
Detached DOM nodes
Detached nodes may remain because:
- a listener retains them;
- a closure stores a reference;
- a library cache never releases them;
- a subscription outlives the component.
Clean up at the owner boundary.
Unbounded caches
Every cache needs:
- maximum size or expiry;
- invalidation;
- identity scope;
- persistence policy;
- behavior under storage pressure.
An optimization that grows without bound is a future outage.
Subscription cleanup
Clean up:
- event listeners;
- observers;
- timers;
- sockets;
- workers;
- requests;
- media resources.
The cleanup should belong to the lifecycle that created the subscription.
Performance budgets
A budget turns an aspiration into a reviewable constraint.
Possible budgets include:
- initial transfer;
- JavaScript execution;
- route transition;
- LCP, INP, and CLS;
- memory;
- third-party cost.
Budgets need context
A single global number hides route and user differences.
Define budgets by:
Keep the budget small enough to guide decisions and flexible enough to reflect product reality.
Bundle budgets
Track:
- initial compressed transfer;
- initial uncompressed parse volume;
- route chunk sizes;
- third-party additions;
- duplicate dependencies.
Bundle size is a useful constraint, not the complete performance outcome.
Metric budgets
Metrics can be reviewed by:
- route;
- release;
- device segment;
- user journey;
- p75 or other percentile.
Define what action occurs when the budget is exceeded.
User-journey budgets
The journey budget can include:
- useful content;
- interaction readiness;
- route transition;
- server confirmation;
- recovery after failure.
It connects performance to product value.
Performance regression
A regression is meaningful when:
- a measured signal worsens;
- under comparable conditions;
- for a meaningful segment;
- beyond expected noise;
- with an owner and response path.
Do not fail a build over a noisy metric without understanding variance.
Performance ownership
Assign responsibility for:
- budgets;
- monitoring;
- dependency review;
- route regressions;
- third-party approval;
- incident response;
- optimization follow-up.
Performance that belongs to nobody becomes cleanup work after users complain.
Performance review in pull requests
Ask when a change affects:
- initial route imports;
- large lists;
- images or fonts;
- third-party scripts;
- rendering scope;
- network waterfalls;
- cache behavior;
- memory lifecycle.
Not every PR needs a benchmark, but every costly change needs a reasoning path.
Performance and accessibility
Do not optimize by removing:
- labels;
- focus behavior;
- semantic structure;
- keyboard support;
- announcements;
- readable loading states.
A fast interface that cannot be used is not a performance success.
Performance and internationalization
Localization can change:
- text length;
- line wrapping;
- font coverage;
- direction;
- date and number formatting;
- layout and interaction size.
Measure representative locales and directions, not only the default language.
Font subsetting and multilingual products
Subset fonts by language or route when appropriate.
Verify:
- fallback behavior;
- glyph coverage;
- layout stability;
- preload scope;
- cache reuse.
Reducing font bytes should not create missing characters or unstable text.
Performance and security
Security controls can affect performance:
- CSP and third-party loading;
- integrity checks;
- authentication redirects;
- encryption and headers;
- sanitization;
- logging and monitoring.
Do not remove a security boundary for a benchmark win. Find the right design trade-off.
Performance and offline architecture
Offline systems can improve perceived speed with local reads.
They also add:
- storage work;
- synchronization;
- reconciliation;
- memory and disk cost;
- stale-data decisions.
Measure both online first load and offline/restore journeys.
Performance and micro-frontends
Micro-frontends can add:
- remote requests;
- duplicate runtimes;
- mount work;
- CSS and font duplication;
- integration waterfalls.
They can also isolate route loading and reduce one giant application bundle.
Measure the assembled page and route transition.
Diagnose LCP
Ask in order:
- Is TTFB slow?
- Is the LCP resource discovered late?
- Is the resource too large or poorly prioritized?
- Is rendering blocked by fonts or JavaScript?
- Is the final element delayed by layout or hydration?
Match the fix to the slow subpart.
Diagnose INP
Ask:
- Was the input delayed by a long task?
- Is the event handler doing unnecessary work?
- Does state update too broad a region?
- Is rendering or layout expensive?
- Can work be deferred, batched, or moved?
The answer may be outside the handler itself.
Diagnose CLS
Ask:
- Which element moved?
- What changed its geometry?
- Was space reserved?
- Did a font, image, ad, or async message arrive late?
- Was the change necessary and user-initiated?
Fix the layout model, not only the visible jump.
Optimization should match cause
Avoid applying a favorite optimization to every symptom.
Performance triage
Prioritize by:
- user impact;
- affected population;
- severity;
- confidence in the cause;
- cost and risk of the fix;
- reversibility;
- business importance of the journey.
The biggest measured cause is often more valuable than a long list of minor improvements.
Performance hypotheses
Use a structured statement:
Then measure before and after under comparable conditions.
Before and after measurement
Record:
- baseline;
- environment;
- route and data;
- change;
- expected signal;
- observed result;
- side effects;
- rollback or follow-up.
An optimization without a baseline is a story, not evidence.
Avoid benchmark theater
Do not:
- optimize a synthetic score disconnected from users;
- report one best run;
- hide slow segments;
- compare different route states;
- claim universal frame rates;
- remove accessible behavior to win a metric.
Performance evidence should improve decisions, not decorate a release.
Production verification
After deployment, verify:
- field metrics;
- route segments;
- release comparison;
- error and abandonment signals;
- cache behavior;
- device and network impact.
Lab success does not prove production success.
Core Web Vitals are not business metrics
They describe aspects of experience.
Business outcomes may include:
- task completion;
- search success;
- conversion;
- retention;
- support contacts;
- revenue or cost.
Connect performance improvements to meaningful user outcomes.
Core Web Vitals are not the whole experience
Also consider:
- accessibility;
- correctness;
- content clarity;
- responsiveness outside measured interactions;
- resilience and offline behavior;
- perceived progress;
- privacy and security.
Optimize the experience, not only the dashboard.
A performance measurement stack
flowchart TD
A["1. Browser Native Observers
(PerformanceObserver: LCP, INP, CLS)"] --> B["2. Custom User Timing API
(performance.mark / performance.measure)"]
B --> C["3. Real User Monitoring (RUM) Beacon
(navigator.sendBeacon to Telemetry API)"]
C --> D["4. Metric Aggregation & p75 Analysis
(Segmented by device tier, connection, country)"]
D --> E["5. Performance Budget Alerts & Regression CI Gates"]Each layer answers different questions.
Performance dashboard design
A useful dashboard shows:
- route and release;
- p75 and distribution;
- device/network segment;
- sample size;
- metric subparts where available;
- business journey;
- regression threshold;
- owner and next action.
Avoid a single score that hides all context.
Percentiles beyond p75
p75 is useful for common experience reporting.
p90, p95, and tail behavior can reveal severe experiences for a smaller group.
Choose percentiles based on the decision and sample size; do not compare tiny samples as if they were stable populations.
Sample size matters
A metric from a handful of sessions can move dramatically by chance.
Use:
- confidence intervals or uncertainty awareness;
- minimum sample thresholds;
- longer windows for sparse routes;
- segmentation that does not destroy statistical usefulness.
Performance budgets by route
Different routes may need different budgets:
Route budgets align architecture with user journeys.
Component performance contracts
A shared component can define expectations for:
- DOM size;
- interaction work;
- image behavior;
- hydration or client code;
- accessibility;
- cleanup;
- bundle contribution.
Contracts should guide design without making every component artificially identical.
Performance and architecture decisions
Before choosing:
- CSR or SSR;
- monolith or micro-frontend;
- client or server component;
- eager or lazy feature;
- cache or refetch;
- virtualized or full list;
state the user journey, bottleneck, measurement, and trade-off.
A performance culture
Healthy teams:
- measure before optimizing;
- share traces and context;
- review performance impact early;
- protect accessibility and security;
- assign ownership;
- learn from production data;
- keep budgets actionable.
Performance is a continuous engineering property, not a final cleanup phase.
Practical lab: Measure and Improve a Slow Interface
Diagnose a slow catalogue using field-like and lab measurements, then test whether virtualization, scheduling, asset changes, or caching address the measured cause.
Every change must begin with a hypothesis and end with a representative re-measurement.
Practical stages 1–5: loading diagnosis
- Record LCP, INP, and CLS baselines.
- Capture a trace for a slow interaction.
- Identify long tasks, layout work, network delays, and rendering cost.
- Inspect LCP.
- Optimize hero delivery and test a resource-priority experiment.
Do not preload or optimize blindly.
Practical stages 6–9: interaction work
- Diagnose INP.
- Remove unnecessary work.
- Virtualize a large list only if the trace supports it.
- Move CPU work to a worker.
Compare responsiveness, memory, communication cost, accessibility, and implementation complexity.
Practical stages 10–13: layout and network
- Diagnose CLS.
- Test font strategy.
- Inspect network waterfalls.
- Remove an artificial waterfall.
Re-measure under representative CPU, network, locale, and direction settings.
Practical stages 14–18: code and navigation
- Inspect JavaScript.
- Audit third-party scripts.
- Add caching.
- Measure an SPA route transition.
- Add User Timing.
Connect each change to a specific journey and signal.
Practical stages 19–24: production discipline
- Investigate memory.
- Create performance budgets.
- Add a CI guardrail.
- Build a RUM design.
- Build a performance regression report.
- Draw the final performance architecture.
Verification: the optimization improves a user journey or validated signal; no universal “zero INP” promise is made.
Practical extension: budget and regression test
Add a performance budget and regression test.
Document why the budget is useful and why it is not a replacement for Core Web Vitals, field data, or user-journey measurement.
Try this yourself
Choose one slow interaction and write:
If you cannot fill in the evidence, measure before changing code.
Troubleshooting guide (Part 1)
| Symptom | Likely cause |
|---|---|
| Lighthouse is good but users are slow | Lab conditions or route differ from field data |
| LCP image is compressed but still late | Discovery, TTFB, priority, or render delay |
| Click handler is short but INP is poor | Input delay or expensive rendering follows |
| CLS occurs after font load | Metrics and fallback geometry differ |
| Virtualization improves speed but breaks keyboard use | UX and accessibility contract was not tested |
Troubleshooting guide (Part 2)
| Symptom | Likely cause |
|---|---|
| Worker adds no benefit | Transfer and setup cost exceed CPU savings |
| Bundle shrinks but interaction is unchanged | Execution, rendering, or network is the actual bottleneck |
| Cache hit rate rises but data is wrong | Freshness, identity, or invalidation policy is weak |
| Budget fails randomly | Sample, environment, or threshold is not controlled |
Completion checklist
- field and lab data are used for different questions;
- p75 and segments are considered;
- LCP is diagnosed by subparts;
- INP includes input, handler, rendering, and paint;
- CLS sources have reserved geometry;
- network and JavaScript waterfalls are inspected;
- list and worker optimizations are evidence-driven;
- memory, cache, and subscription lifecycles are bounded;
- route budgets and performance ownership are explicit;
- production verification follows lab experiments.
Misconceptions to leave behind (Part 1)
| Misconception | Better mental model |
|---|---|
| Performance means Lighthouse score | It is a measured user experience |
| Fast on my machine means fast | Device, network, and route segments differ |
| Average performance is enough | Percentiles reveal slow experiences |
| LCP is image download time | Discovery, server, download, and render all matter |
| Every above-fold image should be preloaded | Prioritize the actual critical resource |
| Lazy loading always helps | It can delay content the user needs now |
| INP is only handler duration | Input, processing, render, and paint form the interaction |
| Memoization always improves speed | Caches have cost and may not target the cause |
Misconceptions to leave behind (Part 2)
| Misconception | Better mental model |
|---|---|
| Virtualize every list | Virtualization has UX and accessibility trade-offs |
| Workers make code faster | They trade main-thread work for communication cost |
| Compression solves JavaScript bloat | Parsing and execution remain |
| More code splitting is always better | Chunks have request and coordination cost |
| Prefetch everything | Speculation consumes user and server resources |
| SSR guarantees good Web Vitals | Topology moves work; it does not remove it |
| Optimization is final polish | Performance is an architectural constraint |
The chapter in one sentence
Measure the real journey, diagnose the actual bottleneck, and make the smallest evidence-backed change that improves users without sacrificing accessibility, security, or correctness.
Next: Chapter 16
The next chapter will build on performance engineering with:
- observability and production diagnostics;
- logging, tracing, and error reporting;
- reliability signals and incident response;
- operational feedback for front-end systems;
- architecture that remains explainable in production.
Questions
Which user journey is slow for a real segment of your users, and what evidence would distinguish network, JavaScript, rendering, layout, and server causes?
16 - Testing Strategies for Resilient Interfaces
Testing Strategies for Resilient Interfaces
Build confidence around real behavior
Chapter 16
Polla Fattah
Today’s goal
Design a test strategy that finds important failures without coupling every test to implementation details.
We will connect:
- risk, evidence, and the testing pyramid;
- static analysis, unit, component, integration, and E2E tests;
- semantic queries, accessible names, keyboard, and focus behavior;
- async UI, network mocking, cancellation, and races;
- Playwright journeys, visual regression, contracts, and browser matrices;
- offline, cache, routing, permissions, security, and performance tests;
- flakiness, CI selection, coverage, maintenance, and test architecture.
By the end of today you can
- choose a test level from the risk and boundary involved;
- test user-visible behavior without inspecting private state unnecessarily;
- use semantic queries and accessible names effectively;
- distinguish simulated DOM tests from real-browser tests;
- test loading, empty, error, retry, cancellation, and optimistic rollback;
- mock at stable boundaries rather than mocking every internal module;
- design realistic browser journeys and failure artifacts;
- detect flakiness instead of normalizing it;
- use coverage as a map rather than a grade;
- maintain a layered suite that remains useful as the UI evolves.
The central principle
A test is valuable when it provides trustworthy evidence about a risk the product actually has.
Fast tests, realistic tests, and broad tests each provide different evidence.
Quality comes from a balanced set of boundaries - not from a single test type or coverage percentage.
Testing is risk management
Ask:
- what can fail?
- who is affected?
- how likely is it?
- how costly is detection after release?
- which test level observes the risk most directly?
The answer determines where to invest test effort.
Confidence comes from different evidence
flowchart TD
SA["Static Checks<br/>(Types, Linters, Schema)"] --> UT["Unit Tests<br/>(Pure Logic & Reducers)"]
UT --> CT["Component Tests<br/>(DOM Semantics & Events)"]
CT --> IT["Integration Tests<br/>(Boundaries & State Flow)"]
IT --> E2E["E2E Tests<br/>(Real Browser Journeys)"]
E2E --> FS["Field Signals<br/>(RUM & Telemetry)"]No one layer can prove the others.
Testing pyramid, trophy, and reality
The shape is less important than the reasoning:
- fast checks should catch cheap mistakes early;
- integration tests should cover important boundaries;
- E2E tests should protect critical journeys;
- production signals should reveal what test environments missed.
Choose the mix from risk, not from a diagram’s proportions.
The cost-confidence spectrum
flowchart LR
A["Static Checks<br/>Lowest Cost / Narrow"] --> B["Unit Tests"]
B --> C["Component Tests"]
C --> D["Integration Tests"]
D --> E["Browser Journeys<br/>Highest Cost / Broad"]Use the narrowest test that provides enough confidence for the risk.
Put a failure at the closest useful boundary.
Static analysis is part of testing strategy
Static checks can catch:
- type inconsistencies;
- unreachable branches;
- unsafe imports;
- dependency direction violations;
- accessibility lint issues;
- unused or suspicious code.
They are fast evidence, but they do not observe runtime behavior.
Static analysis has limits
It cannot fully prove:
- server responses;
- browser timing;
- focus behavior;
- visual layout;
- network failures;
- authentication flows;
- user comprehension;
- correct business outcomes.
Use runtime tests where the risk exists.
Unit tests target deterministic logic
Good unit targets include:
- parsers;
- reducers;
- formatters;
- validation;
- URL serialization;
- cache-key construction;
- state-transition logic.
Avoid unit testing language syntax
Do not write tests to prove that:
Array.prototype.mapmaps;- an
ifstatement branches; - a framework renders a known element;
- a constant equals itself.
Test your decision logic and product behavior around the language feature.
Test behavior, not line count
This is more valuable than asserting that every internal line executed.
Coverage can identify untested areas, but behavior determines confidence.
Component tests should resemble user interaction
Prefer interactions and visible outcomes over calls to private helpers or internal state snapshots.
Why role-based queries are valuable
Role-based queries:
- reflect the accessibility tree;
- survive many visual refactors;
- encourage meaningful semantics;
- match how assistive technologies identify controls.
They are a quality signal, not complete accessibility certification.
Accessible name is part of the contract
The role identifies the control type.
The accessible name identifies what it does.
Test both when the behavior depends on them.
Accessible name computation is a web concept
Accessible names can come from:
- visible text;
- associated labels;
aria-label;aria-labelledby;- other standard mechanisms.
Testing-library queries reflect platform semantics; they do not invent a private testing concept.
Preferred query order
Often prefer:
- role and accessible name;
- label;
- placeholder or visible text where appropriate;
- semantic state or value;
- test ID as a fallback.
Choose the query that represents the user-facing contract.
getByLabelText for forms
This tests that the label and control are associated in a meaningful way.
It can reveal form accessibility problems that a class selector would miss.
Role queries do not replace accessibility audits
A control can be findable by role and still have:
- poor focus behavior;
- incorrect state announcements;
- bad color contrast;
- keyboard traps;
- confusing reading order;
- missing error association.
Automated checks supplement human and assistive-technology evaluation.
Test IDs are legitimate fallbacks
Use a test ID when:
- no meaningful user-facing semantic exists;
- a structural target is required;
- a complex visualization needs a stable anchor;
- the stronger query would be misleading.
The issue is not the attribute; it is using it to avoid designing a semantic interface.
Avoid overfitting to CSS selectors
This couples the test to styling and class names.
Use CSS selectors when CSS structure is genuinely the behavior under test, such as a visual or layout integration.
Avoid testing internal state
Prefer:
Over:
Internal state is an implementation choice; visible behavior is the contract.
Avoid testing private methods
Private methods can change during a refactor while the product behavior remains correct.
Test them indirectly through the public behavior unless the logic is extracted into a meaningful pure unit with its own contract.
Component test example
The test follows interaction and outcome rather than implementation.
Integration tests connect boundaries
Integration may include:
- component plus reducer;
- form plus validation;
- query layer plus cache;
- route plus URL state;
- API adapter plus parser;
- shell plus remote contract.
The scope should reflect a real interaction boundary.
Integration is a spectrum
flowchart LR
M["Two Collaborating Modules"] --> FB["Feature Boundary"]
FB --> RT["Route & URL State"]
RT --> AW["Application Workflow"]Name the scope clearly.
The goal is not to make every test “large”; it is to verify the collaboration that matters.
Browser simulation versus a real browser
| Simulated environment | Real browser |
|---|---|
| fast and focused | realistic layout and platform behavior |
| easy unit/integration loop | network, focus, CSS, storage, workers |
| incomplete browser APIs | higher cost and setup |
| good for most logic | needed for critical browser behavior |
Use both intentionally.
Vitest and test-runner responsibilities
A runner typically provides:
- test discovery;
- assertions;
- isolation;
- mocks and timers;
- watch mode;
- coverage;
- reporting.
It does not decide whether the tested behavior is valuable.
Watch mode supports rapid feedback
Watch mode should:
- rerun affected tests;
- preserve readable failure output;
- make focused testing easy;
- encourage small feedback loops.
Do not let a slow or noisy watch workflow push developers toward skipping tests.
Browser mode
Browser-based component tests can reveal:
- real DOM behavior;
- CSS and layout interaction;
- focus and selection;
- browser API differences.
Use them where simulated DOM limitations are relevant rather than moving every unit test into a browser.
Simulated DOM still has value
It is often sufficient for:
- roles and labels;
- event flows;
- state transitions;
- loading and error rendering;
- form validation;
- network boundary behavior.
Know what the environment does not implement before trusting it for a browser-specific claim.
User event simulation
Prefer realistic sequences over synthetic dispatch:
flowchart LR
F["focus"] --> KD["keydown"]
KD --> IN["input"]
IN --> CH["change"]
CH --> BL["blur"]A single direct property assignment may skip behavior your product depends on.
Use a user-event tool when the interaction sequence matters.
Keyboard testing
Test:
- Tab order;
- Enter and Space activation;
- arrow-key navigation;
- Escape behavior;
- focus after open and close;
- disabled controls;
- keyboard traps.
Mouse-only tests miss a major part of interaction architecture.
Focus testing
Focus is observable user state.
Test where focus goes, what happens on dismissal, and whether the trigger is restored.
Testing async UI
Cover asynchronous state transitions:
stateDiagram-v2
[*] --> Idle
Idle --> Loading: User Trigger / Fetch
Loading --> Success: 200 OK Response
Loading --> Empty: Zero Results Found
Loading --> Error: Network / 500 Fail
Error --> Loading: Retry Action
Success --> [*]
Empty --> [*]Assert after the relevant state settles, not immediately after starting an asynchronous operation.
Avoid arbitrary sleeps
Sleeps make tests slow and still do not prove the condition is ready.
Wait for a meaningful outcome, event, network response, or visible state.
Assertions should match user outcomes
Prefer:
Over asserting:
Unless the timer itself is the contract, test what the user experiences.
Mock at stable boundaries
Good boundaries include:
- network API;
- clock;
- storage adapter;
- repository;
- browser capability;
- feature flag provider.
Mocking at a stable boundary keeps tests aligned with architecture.
Mocking internal modules can over-couple tests
If every internal import is mocked, a refactor changes hundreds of tests without changing behavior.
Mock only where the boundary is expensive, nondeterministic, or outside the test’s responsibility.
Network mocking
Network mocks should model:
- status codes;
- response shapes;
- latency where relevant;
- malformed payloads;
- cancellation;
- retries;
- partial failure.
Mock the protocol the feature depends on, not a convenient internal helper.
Mock Service Worker style
An interception layer can let application code use real request paths while tests control responses.
This preserves more of the transport boundary than mocking fetch() in every module.
Use realistic response contracts and reset handlers between tests.
Test success, error, and empty states
For remote features, cover:
The happy path is one state among several that users will encounter.
Test cancellation and races where relevant
sequenceDiagram
autonumber
participant UI as Search Input
participant Net as Network Boundary
participant Srv as Server
UI->>Net: Query "erb" (Request 1)
UI->>Net: Query "erbil" (Request 2)
Note over Net: Request 1 delayed (800ms)
Net->>Srv: Process Request 2
Srv-->>UI: Results for "erbil" (Arrives at 200ms)
Srv-->>UI: Results for "erb" (Arrives at 800ms - LATE!)
Note over UI: Race Bug: Stale results clobber newest search!The test should prove that Query 2 remains authoritative and Request 1 was aborted.
Demonstration: Failing race condition
A test that does not control network timing will never detect this intermittent race.
Demonstration: Setting up the race test
Simulate network latency on the first query using MSW:
Demonstration: Asserting race resilience
Trigger rapid typing and verify late responses are ignored:
The test controls network timing at the transport boundary.
What the passing race test proves
Passing the mocked component race test validates:
- Order-independence: delayed earlier requests do not overwrite newer state;
- DOM accuracy: rendered search results reflect the active query;
- Cancellation:
AbortControllerdispatches signal on new keystrokes.
The contract between input and view is verified.
What the passing test still cannot prove
Even with the component test green, it cannot prove:
- Backend capacity: server rate-limiting under concurrent queries;
- Screen-reader cadence:
aria-livespeech queue congestion; - Device rendering: frame drops on low-tier mobile hardware;
- Input methods: IME Arabic/CJK composition events.
Confidence requires complementary evidence across layers.
Fake timers
Fake timers help test:
- debounce;
- retry backoff;
- polling;
- expiration;
- scheduled UI work.
Advance time deliberately and restore real timers after each test.
Do not use fake time to hide a missing synchronization condition.
Test data builders
Builders provide valid defaults and make the relevant variation visible.
They reduce repetitive fixture setup without hiding important input differences.
Avoid magic fixtures
A giant fixture can make every test depend on irrelevant fields.
Prefer small, intention-revealing builders and explicit edge-case data.
The test should show why the data matters.
End-to-end testing
E2E tests verify a deployed-like journey through:
- browser;
- route;
- application code;
- server or controlled backend;
- network and storage;
- visible user outcomes.
They provide broad confidence at higher cost.
Playwright locators
Prefer locators that reflect user semantics:
Use stable test IDs when semantic locators do not describe the target.
Auto-waiting
A browser tool can wait for:
- visibility;
- enabled state;
- attachment;
- navigation;
- assertions.
Use built-in waiting and web-first assertions rather than adding arbitrary sleeps.
Web-first assertions
Assertions should wait for the user-visible condition and produce useful failure output.
Test real user journeys
Good E2E candidates include:
- sign-in and redirect;
- search and filter;
- edit and save;
- error and retry;
- offline draft recovery;
- back and forward navigation;
- critical purchase or submission.
Do not use E2E to cover every small rendering branch.
The critical-path suite
Keep a small, reliable suite for:
It should run quickly enough to block unsafe releases and produce actionable artifacts on failure.
E2E data isolation
Tests should use:
- isolated accounts or tenants;
- unique records;
- resettable data;
- explicit cleanup;
- deterministic server state.
Shared mutable test data creates order dependence and flakiness.
Test setup through APIs when appropriate
Create test state through a supported API or fixture layer when the behavior under test is not the setup flow.
Then use the UI for the actual journey.
Do not bypass the behavior you are trying to verify.
E2E authentication
Choose between:
- UI login for the login journey;
- API or storage setup for unrelated tests;
- reusable authenticated contexts with isolated identities.
Document which boundary each test actually verifies.
Multi-browser and mobile testing
Test a matrix based on risk:
- supported browser engines;
- viewport classes;
- touch and keyboard behavior;
- locale and direction;
- critical layout and input paths.
Do not run every test on every browser without a reason; do not test only one browser when compatibility matters.
Accessibility-oriented testing
Combine:
Each catches different failures.
Automated accessibility checks
Automated checks can find:
- missing labels;
- invalid ARIA relationships;
- contrast issues in some cases;
- landmark and structural problems.
They cannot fully evaluate meaning, workflow, focus strategy, or real assistive-technology experience.
Visual regression testing
Visual tests are useful for:
- design-system components;
- themes;
- responsive layouts;
- typography;
- overlays;
- cross-route visual contracts.
They should complement behavior tests, not replace them.
Visual regression noise
Control:
- browser version;
- fonts;
- viewport;
- animations;
- network content;
- time and locale;
- screenshot masking.
Review visual changes as product changes, not blindly approve every diff.
Snapshot testing
Snapshots can reveal broad output changes quickly.
They become low-value when:
- output is huge;
- reviewers approve without reading;
- implementation details dominate;
- snapshots replace focused assertions.
Keep snapshots small and meaningful.
Contract tests
Contract tests verify an agreement at a boundary:
- API request and response schema;
- event payload;
- package public API;
- micro-frontend mount contract;
- generated client assumptions.
Test the boundary the consumer actually depends on.
Schema-driven testing
A runtime schema can support:
- response validation;
- generated examples;
- malformed-input tests;
- compatibility checks;
- property-based variations.
Shared schemas improve agreement but do not remove the need to test real delivery and failure.
Test network conditions
Include scenarios such as:
- slow response;
- offline;
- timeout;
- aborted request;
- malformed response;
- 401/403;
- 404;
- 422;
- 500;
- partial dashboard failure.
Resilient UI is defined by recovery behavior.
Test offline behavior where promised
Verify:
- draft survives reload;
- queued work is visible;
- retry has a limit;
- permanent failure is recoverable;
- reconnection does not duplicate work;
- stale data is labeled appropriately.
Do not claim offline support from an offline screen alone.
Flaky tests are signals
Flakiness can indicate:
- uncontrolled time;
- shared state;
- races;
- missing awaits;
- unstable selectors;
- order dependence;
- real product nondeterminism.
Treat it as a defect in the test or system until understood.
Never normalize flakiness
“It passes on retry” hides:
- release risk;
- missing confidence;
- slow CI;
- developer distrust;
- real timing bugs.
Quarantine only with an owner, issue, scope, and removal plan.
Retries are not a fix
Retries may help identify intermittent infrastructure failures.
They can also:
- hide races;
- double mutations;
- make failures slower;
- produce false green builds.
Use retries as a diagnostic or bounded operational policy, not as proof of reliability.
Trace artifacts improve failure diagnosis
Keep useful artifacts such as:
- screenshots;
- video;
- browser trace;
- console output;
- network logs;
- server correlation IDs;
- DOM snapshots where safe.
Artifacts should help reconstruct the user journey without leaking sensitive data.
Failure messages are part of test quality
A useful failure says:
- which journey failed;
- what state was expected;
- what state appeared;
- which request or boundary was involved;
- where artifacts are stored.
“Expected true to be false” is rarely enough for a team to act quickly.
Test names are documentation
is more useful than:
Name the behavior and the important condition.
Arrange–Act–Assert
flowchart LR
ARR["Arrange<br/>Establish state & mock boundaries"] --> ACT["Act<br/>Perform meaningful user interaction"]
ACT --> AST["Assert<br/>Verify user-visible outcomes & semantics"]Keep the structure readable, even when a test uses several assertions.
Given–When–Then
This style can clarify product behavior for technical and nontechnical reviewers.
One test can have several assertions
Several assertions are appropriate when they describe one outcome:
Split tests when assertions represent independent behaviors with different setup or failure meaning.
Avoid mega tests
A mega test that covers an entire application can be:
- slow;
- hard to diagnose;
- stateful;
- impossible to run in parallel;
- fragile under small changes.
Keep critical journeys focused and compose confidence across layers.
Test independence
Each test should control or reset:
- data;
- clock;
- network;
- storage;
- authentication;
- feature flags;
- browser state.
Order should not decide whether a test passes.
Parallel execution
Parallel tests need:
- isolated data;
- independent ports or contexts;
- deterministic seeds;
- no shared mutable files;
- bounded server resources.
Parallelism improves speed only when the environment preserves independence.
Determinism
Control or model:
- time;
- randomness;
- IDs;
- network order;
- locale;
- timezone;
- animation;
- async scheduling.
Do not remove realistic variability from the product merely to make tests easy.
Locale, RTL, and date testing
Test representative:
- long translations;
- plural forms;
- right-to-left layout;
- localized dates and numbers;
- timezone boundaries;
- daylight-saving transitions.
The default locale hides real layout and logic bugs.
Browser APIs need realistic tests
If a feature uses:
- storage;
- service workers;
- workers;
- notifications;
- permissions;
- clipboard;
- media;
- WebSockets or SSE;
use a test environment that models the relevant behavior or include a real-browser test.
Testing service workers and live connections
Test:
- install and update;
- cache strategy;
- offline fallback;
- message validation;
- reconnect;
- duplicate events;
- missed-event recovery;
- cleanup.
Do not test only the connected happy path.
Race testing
Make the race controllable at the boundary:
sequenceDiagram
participant Test as Test Runner
participant MSW as Mock Boundary
participant App as Client UI
Test->>MSW: Hold Response A
App->>MSW: Dispatch Request A
App->>MSW: Dispatch Request B
Test->>MSW: Resolve Response B
MSW-->>App: Render B
Test->>MSW: Resolve Response A (Delayed)
MSW-->>App: Discard or Ignore A
Test->>App: Assert UI displays BAssert that the current identity or version wins according to the product policy.
Error and loading boundary testing
Verify that a failure is contained at the intended boundary:
Also test loading fallbacks for focus, layout, and accessibility.
Cache behavior testing
Test:
- key separation;
- fresh versus stale;
- deduplication;
- invalidation;
- mutation reconciliation;
- cache eviction;
- user identity scope.
A cache test should prove what the user sees, not merely that a map received a value.
Optimistic update test
Cover the optimistic state lifecycle:
flowchart TD
Init["User Submits Action"] --> Prov["Apply Provisional Value to UI"]
Prov --> Net{"Server Response"}
Net -->|200 Confirmed| Keep["Preserve or Reconcile"]
Net -->|500 Rejected| Roll["Rollback to Previous State & Show Error Alert"]
Net -->|Newer Update Exists| Guard["Avoid Overwriting Newer Truth"]Optimism without rollback testing is only a happy-path demo.
Testing forms
Cover:
- labels and names;
- touched and dirty behavior;
- field and cross-field validation;
- disabled and pending states;
- server validation mapping;
- draft preservation on failure;
- successful reset or navigation.
Forms are workflows, not just input snapshots.
Testing routing
Verify:
- route parsing and invalid parameters;
- redirects;
- URL state serialization;
- back and forward behavior;
- scroll and focus;
- route-level loading and errors;
- code-split failure and recovery.
The browser history is part of the product contract.
Permission and security testing boundaries
Test:
- unauthorized UI behavior;
- server rejection despite hidden controls;
- CSRF decisions;
- CORS integration;
- token and session expiration;
- safe rendering;
- open redirect prevention.
Do not treat a client-side permission branch as security enforcement.
Coverage is a map, not a grade
Coverage can show:
- unexecuted branches;
- missing error paths;
- dead code candidates;
- areas needing review.
High coverage can still miss wrong assertions, incorrect fixtures, inaccessible behavior, and integration failures.
Mutation testing
Mutation testing changes code deliberately to see whether tests fail.
It can reveal weak assertions and untested logic.
Use it selectively on high-risk pure logic; applying it everywhere may cost more than the confidence gained.
Property-based testing
Property-based tests generate many inputs and verify invariants:
They are useful for parsers, state transitions, URLs, and domain rules with broad input space.
Themes and responsive visual testing
Visual coverage should include:
- light and dark themes;
- narrow and wide viewports;
- high text scale where relevant;
- RTL;
- long content;
- loading and error states.
A component that looks correct in one theme and viewport is not fully verified.
Test environment strategy
Define which environment supports each layer:
flowchart LR
U["Unit<br/>Fast Node / Bun"] --> C["Component<br/>jsdom / Browser Runner"]
C --> I["Integration<br/>Controlled Mock Services"]
I --> E["E2E<br/>Real Headless Browsers"]
E --> P["Production<br/>Canary / RUM Telemetry"]Keep environment differences documented and intentional.
Smoke tests
Smoke tests provide a fast release signal:
- app loads;
- critical route works;
- one key action succeeds;
- health and assets are available.
They should be small, reliable, and safe to run against production-like systems.
Production tests must be safe
Use:
- read-only checks;
- isolated test accounts;
- non-destructive identifiers;
- explicit cleanup;
- rate limits;
- privacy-aware logging.
Do not test a production payment or deletion path by accident.
Feature flags and testing
Flags multiply states:
Define flag ownership, expiry, test coverage, rollout, and removal.
An old flag is permanent untested architecture.
Test matrix explosion
You cannot test every combination of:
- browsers;
- locales;
- roles;
- feature flags;
- network conditions;
- route states;
- data sizes.
Prioritize combinations by risk and use representative sampling plus targeted coverage.
CI test parallelism
Parallel CI can reduce feedback time when:
- tests are independent;
- data is isolated;
- workers have resources;
- artifacts remain attributable;
- test selection is reliable.
Fast, nondeterministic CI is not a quality improvement.
Test selection
Use changed files, dependency graphs, and risk labels to select fast feedback.
Still run a broader suite at release boundaries or on a reliable schedule.
Selection is safe only when the graph and ownership assumptions are accurate.
Quarantine flaky tests carefully
A quarantine policy needs:
- owner;
- issue;
- reason;
- expiry or review date;
- alternative coverage;
- removal plan.
Quarantine is a temporary containment measure, not a second home for broken tests.
Delete low-value tests
Remove tests that:
- assert implementation with no product risk;
- duplicate stronger coverage;
- fail noisily without useful evidence;
- protect behavior intentionally being removed;
- cost more to maintain than the confidence they provide.
Test suites need maintenance like production code.
Testing refactors
Good tests allow internal refactoring while preserving contracts.
If a visual extraction breaks dozens of tests, inspect whether those tests depend on private structure rather than user behavior.
Testing should support architecture evolution, not freeze every implementation detail.
Test architecture smells
Watch for:
- every test queries a test ID;
- every internal function is mocked;
- refactoring markup breaks hundreds of tests;
- E2E suite takes hours;
- failures disappear on retry;
- 100% coverage but critical bugs escape;
- snapshots are approved without review;
- production bugs cannot be reproduced.
These are signals to redesign the evidence strategy.
A balanced strategy by layer
flowchart TD
S["Static Checks: Contracts & unsafe patterns"]
U["Unit Tests: Pure rules & transformations"]
C["Component Tests: Visible interaction & semantics"]
I["Integration Tests: Boundaries & failure recovery"]
E["E2E Tests: Critical user revenue journeys"]
P["Production: Safe smoke, RUM & telemetry"]
S --> U --> C --> I --> E --> PThe layers should reinforce rather than duplicate one another.
Testing a dialog
Verify:
- accessible name;
- initial focus;
- Escape and close button;
- focus restoration;
- background interaction policy;
- error and loading states;
- keyboard navigation.
One user-facing behavior can need several complementary tests.
Testing a pricing function
Unit-test:
- currency and rounding;
- zero and negative handling;
- locale formatting;
- discounts and boundaries;
- invalid input policy.
Then integration-test where the formatted value appears in a real product journey.
Testing search
Cover the complete search user journey:
flowchart LR
Type["User Types Query"] --> Debounce["Debounce Timer"]
Debounce --> Req["Network Request"]
Req --> Load["Loading Skeleton"]
Load --> Res["Results Display"]
Res --> URL["Sync URL Search Params"]Search is a small feature with many asynchronous contracts.
Testing offline drafts
Verify:
- draft persistence;
- reload recovery;
- outbox entry;
- sync status;
- retry limit;
- duplicate prevention;
- conflict behavior.
Do not test only the initial save click.
Testing micro-frontends
Use:
- remote unit and component tests;
- contract tests for mount and events;
- host integration tests;
- remote failure and fallback tests;
- end-to-end cross-route journeys;
- version compatibility checks.
The runtime boundary needs evidence at both sides.
Testing rendering topologies
Verify:
- server and client output agreement;
- hydration behavior;
- static freshness;
- stream loading and failure;
- client-only boundaries;
- route navigation and handoff;
- secrets staying server-side.
Rendering architecture changes what the test environment must observe.
Testing performance contracts
Test or monitor:
- route budgets;
- asset and chunk limits;
- key interaction timing;
- list behavior at representative sizes;
- memory cleanup;
- loading and transition markers.
Use performance tests as signals with context, not brittle universal promises.
Testing security contracts
Verify:
- safe text rendering;
- sanitized rich content;
- CORS and credential behavior;
- CSRF protection;
- server authorization;
- redirect validation;
- security headers;
- no secrets in browser artifacts.
Security tests should target trust boundaries explicitly.
The test strategy document
Document:
- risks;
- test layers;
- supported browsers;
- data and environment strategy;
- ownership;
- CI selection;
- flaky-test policy;
- production verification;
- review and deletion rules.
The document keeps the suite intentional as the system evolves.
Testing philosophy
flowchart TD
C["Test the Contract, Not Implementation"]
B["Choose the Nearest Useful Boundary"]
A["Make Every Failure Actionable"]
J["Protect Critical Revenue Journeys"]
T["Keep the Suite Fast & Trustworthy"]
C --- B --- A --- J --- TMore tests do not automatically mean higher quality.
Better evidence does.
Practical lab: Resilient UI Integration Suite
Test a catalogue workflow through user-visible behavior, accessibility semantics, network boundaries, and recovery paths.
The practical builds a layered strategy rather than one giant end-to-end test.
Practical stages 1–2: pure logic & accessible semantics
- Stage 1 (Pure Logic & Parser Unit Testing):
- Test price formatting, pagination math, and schema parsers against valid, boundary, and corrupt input.
- Pure fast Node runtime without DOM overhead.
- Stage 2 (Component Semantics & Accessible Names):
- Query by
getByRoleandgetByLabelText. - Test keyboard navigation (
Tab,Escape) and focus retention.
- Query by
Verification: Tests depend on platform accessibility contracts, not CSS classes or private component state.
Practical stages 3–4: boundary mocking & fault injection
- Stage 3 (Network Interception with MSW):
- Intercept requests at the HTTP transport boundary.
- Simulate network delays, 500 server crashes, and offline states.
- Stage 4 (Asynchronous Resilience & Fault Injection):
- Inject deliberate faults (dropped
AbortController, broken optimistic rollback, missingaria-invalid). - Confirm tests fail immediately and pinpoint the exact failure mechanism.
- Inject deliberate faults (dropped
Verification: Tests prove loading, empty, error, retry, and cancellation behaviors without arbitrary sleep() delays.
Practical stage 5: Playwright critical-path journey
- Stage 5 (Playwright Critical-Path Browser Journey):
- Run a real headless Chromium/Firefox/WebKit test.
- Exercise the full user journey: search, filter, optimistic order drafting, and error recovery.
- Capture trace artifacts, network waterfalls, and screenshots on failure.
Verification: Fast feedback in CI with zero flakiness; high confidence across real browser layout engines.
Practical extension: contract and mutation tests
Add:
- a contract test for the runtime-validated API response;
- a deliberate mutation test that changes an important pricing or validation rule.
Confirm that the suite fails for the right reason and reports an actionable difference.
Try this yourself
Choose one critical user journey and map:
flowchart LR
R["Risk Identified"] --> B["Test Boundary"]
B --> S["Setup / Seed"]
S --> A["User Action"]
A --> AS["Semantic Assertion"]
AS --> AR["Diagnosable Artifact"]Then remove one test that duplicates stronger evidence and explain why confidence remains adequate.
Troubleshooting guide (Part 1)
| Symptom | Likely cause |
|---|---|
| Tests break after harmless markup refactor | Assertions depend on private structure |
| Everything uses test IDs | User-facing semantics are missing or ignored |
| Async tests need sleeps | Tests wait for time instead of conditions |
| Mocks hide integration failures | Mock boundary is too deep |
| E2E suite is slow and flaky | Too much setup and shared mutable data |
Troubleshooting guide (Part 2)
| Symptom | Likely cause |
|---|---|
| Retry makes CI green | Flakiness is being normalized |
| Coverage is high but bugs escape | Assertions and risk mapping are weak |
| Visual diffs are always approved | Review policy and baseline ownership are weak |
| Offline test passes only once | Storage and cleanup are not isolated |
Completion checklist
- risks determine the test layers;
- static, unit, component, integration, and E2E roles are clear;
- tests use semantic user-facing queries where appropriate;
- keyboard and focus behavior are covered;
- async states and races are deterministic;
- network mocks sit at stable boundaries;
- critical browser journeys have failure artifacts;
- accessibility and visual checks supplement behavior tests;
- flakiness has ownership and a removal policy;
- coverage and CI selection support, rather than replace, judgment.
Misconceptions to leave behind (Part 1)
| Misconception | Better mental model |
|---|---|
| 100% coverage means well tested | Coverage is a map, not a confidence grade |
| Unit tests are always better | Use the nearest boundary that proves the risk |
| Everything should be E2E | Broad tests are costly and should protect journeys |
| Component tests should inspect state | Test visible behavior and semantics |
| CSS selectors are forbidden | Use the selector that matches the contract |
| Test IDs are bad | They are a legitimate fallback when semantics are unsuitable |
Misconceptions to leave behind (Part 2)
| Misconception | Better mental model |
|---|---|
getByRole certifies accessibility | Queries are one quality signal, not an audit |
| Simulated DOM equals a browser | Real browser behavior needs targeted tests |
| Mocks make tests reliable | Boundary choice and realistic failures matter |
| Sleeps fix async tests | Wait for meaningful conditions |
| Retries solve flaky tests | They can hide nondeterminism |
| More tests always mean higher quality | Better evidence and maintainability matter |
The chapter in one sentence
Build a layered, user-centered test strategy that protects real risks, observes meaningful boundaries, and remains trustworthy under failure and change.
Next: Chapter 17
The next chapter will build on resilient testing with:
- maintainability and refactoring;
- technical debt and architectural evolution;
- documenting decisions;
- sustainable quality over a system’s lifetime.
Questions
Which important user journey currently has the most confidence from implementation details - and the least evidence from the behavior users actually experience?
17 - Continuous Delivery, Observability & Maintenance
Continuous Delivery, Observability & Maintenance
Close the loop from commit to production learning
Chapter 17
Polla Fattah
Today’s goal
Make delivery and production feedback part of front-end architecture.
We will connect:
- CI, reproducibility, artifacts, and environments;
- preview, staging, production, and deployment protection;
- release strategies, feature flags, canaries, and rollback;
- logs, metrics, traces, RUM, errors, and session context;
- dashboards, alerts, SLOs, error budgets, and synthetic monitoring;
- dependency, browser, API, service-worker, and technical-debt maintenance;
- ownership, runbooks, incident response, privacy, and production readiness.
By the end of today you can
- design a CI pipeline with fast and broad verification stages;
- produce and promote one versioned artifact;
- separate preview, staging, and production responsibilities;
- use feature flags without confusing them with authorization;
- plan canary, rollback, and kill-switch behavior;
- distinguish monitoring from observability;
- define useful frontend telemetry without oversharing data;
- create SLOs and alerts with owners and response actions;
- maintain dependencies, browser support, APIs, storage, and service workers;
- write a production runbook and rehearse the delivery loop.
The central principle
A production system is complete only when it can be built reproducibly, released safely, observed meaningfully, rolled back deliberately, and maintained continuously.
Delivery is not the final step after architecture.
It is where architecture meets users, operators, and reality.
The complete production loop
flowchart TD
Commit["1. Commit & Code Review"] --> Verify["2. Parallel CI Verification Gates"]
Verify --> Build["3. Build Immutable Hashed Artifact"]
Build --> Preview["4. Deploy Isolated Preview URL"]
Preview --> Canary["5. Progressive Canary Rollout (5% → 25% → 100%)"]
Canary --> Observe["6. Continuous Observability & SLO Monitoring"]
Observe --> Mitigate{"Health Check"}
Mitigate -->|Degraded / Spikes| Rollback["Automated Rollback / Kill Switch"]
Mitigate -->|Healthy| Maintain["7. Technical Debt & Dependency Maintenance"]Every arrow needs an owner, evidence, and a recovery path.
Delivery is part of architecture
Delivery decisions determine:
- which code reaches users;
- how quickly a fix can ship;
- whether a release can be identified;
- how a failure is contained;
- whether rollback is possible;
- how production evidence reaches developers.
A frontend that cannot be safely changed is not operationally complete.
Continuous integration
CI validates changes in a clean, shared environment.
It can verify:
- installation;
- linting and formatting;
- type checking;
- unit and integration tests;
- production build;
- artifact integrity;
- smoke behavior.
CI is a feedback system, not merely a hosted command runner.
CI is more than “run tests”
flowchart LR
Inputs["Source + Lockfile + Node Runtime"] --> Install["Clean 'npm ci' Install"]
Install --> Checks["Type Checks, Linters, Tests"]
Checks --> Build["Production Bundle with Content Hashes"]
Build --> Artifact["Immutable Artifact + Release Metadata"]The environment and produced artifact are part of what CI verifies.
CI events
Useful triggers include:
- pull request;
- push to a protected branch;
- release tag;
- scheduled dependency check;
- manual promotion;
- rollback or emergency fix.
Choose the checks and permissions appropriate to each event.
Fast feedback still matters
flowchart LR
Ed["Editor / IDE"] --> Loc["Pre-commit Hooks"]
Loc --> PR["PR Verification CI"]
PR --> Merge["Merge to Main"]
Merge --> Rel["Canary Release"]Fast checks should catch common mistakes near the change.
Broad checks should protect integration and delivery confidence.
Do not make every local edit wait for the entire release pipeline.
Parallel CI
Independent jobs can run concurrently:
flowchart LR
subgraph ParallelGates["Parallel CI Gates"]
L["Lint & Formatting"]
T["Type Checking (tsc)"]
U["Unit & Component Tests"]
S["Security Secrets Scan"]
end
L --> Decision["Build & Deploy Decision"]
T --> Decision
U --> Decision
S --> DecisionParallelism requires isolated environments, clear dependencies, and attributable artifacts.
CI caching
Cache:
- package downloads;
- transformed dependencies;
- test artifacts;
- build tasks;
- browser binaries when safe.
Cache keys must include every input that affects correctness.
An incorrect cache can create false green builds.
Reproducibility
A reproducible build depends on:
- locked dependencies;
- known runtime version;
- explicit environment inputs;
- stable build configuration;
- deterministic generation;
- documented external services.
If two clean builds produce different artifacts, investigate why before relying on promotion.
Continuous integration versus delivery versus deployment
flowchart LR
CI["Continuous Integration<br/>Automated verification of every push"] --> CD["Continuous Delivery<br/>Always maintains a deployable release candidate"]
CD --> CDeploy["Continuous Deployment<br/>Automatic rollout to production upon passing gates"]An organization can practice delivery without automatically deploying every commit.
Automation level should match risk and release confidence.
Build artifact
An artifact is the output intended for deployment:
- static assets;
- HTML;
- server bundle;
- manifests;
- source-map references;
- release metadata.
It should be identifiable and traceable back to source and dependencies.
Immutable artifacts
Rebuilding separately for staging and production can produce different output.
Promotion should change environment configuration, not application bytes, whenever possible.
Static frontend artifact
For a static application, verify:
- asset fingerprints;
- manifest paths;
- base URL;
- cache headers;
- source-map policy;
- service-worker version;
- public environment values.
The static artifact is still an operational release unit.
Server-rendered artifact
A server-rendered release may include:
- server code;
- client chunks;
- route configuration;
- environment references;
- source maps;
- migration compatibility.
Test the server and browser parts as one release contract.
Environments
Different environments have distinct operational responsibilities:
flowchart LR
Dev["Development<br/>Local / Fast HMR"] --> Prev["Preview<br/>PR-Isolated Ephemeral URL"]
Prev --> Stg["Staging<br/>Shared Pre-Production Environment"]
Stg --> Prod["Production<br/>Global Edge CDN & Telemetry"]Each environment should have a purpose rather than existing only because a template created it.
Document data, secrets, access, observability, and deployment differences.
Development environment
Optimize for:
- fast feedback;
- local debugging;
- safe test data;
- representative boundaries;
- developer control.
Development convenience must not hide production-critical behavior.
Preview environments
A preview environment connects a change to a deployed, reviewable experience.
It can reveal:
- asset path issues;
- routing failures;
- environment assumptions;
- visual changes;
- integration behavior.
Preview environments improve communication
They let reviewers discuss:
- the actual UI;
- loading and error states;
- responsive behavior;
- accessibility;
- product intent.
Preview is evidence that complements code review; it does not replace it.
Preview environment data
Use data that is:
- safe to expose to reviewers;
- representative enough to reveal behavior;
- isolated from production;
- resettable or expiring;
- privacy-aware.
Never copy sensitive production data into preview casually.
Staging
Staging can validate:
- production-like infrastructure;
- integrations;
- deployment scripts;
- migration compatibility;
- smoke and end-to-end behavior.
It is not automatically identical to production and should not create false certainty.
Production
Production needs:
- controlled access;
- secrets management;
- monitoring;
- rollback;
- incident response;
- privacy controls;
- support ownership;
- documented change history.
The production environment is part of the product’s operating system.
Deployment protection
Protect production with:
- required reviews;
- environment approvals;
- scoped credentials;
- branch or tag rules;
- concurrency control;
- audit history;
- automated verification.
Protection should reduce dangerous mistakes without making recovery impossible.
Environment secrets
Secrets should be injected through the deployment environment, not committed or embedded in public frontend assets.
Separate:
Rotate and audit access.
Principle of least privilege
Give each job and environment only the access it needs.
Examples:
- preview cannot delete production data;
- build cannot deploy unrelated services;
- frontend runtime cannot read deployment credentials;
- telemetry cannot expose raw user content.
Least privilege limits blast radius.
Deployment concurrency
Decide what happens when two deployments target one environment:
- queue;
- cancel older pending deploy;
- serialize production;
- allow parallel previews;
- block promotion until verification.
Ambiguous concurrency creates releases that are difficult to identify and roll back.
Deployment history
Record:
- release version;
- commit;
- artifact digest;
- environment;
- time;
- actor or automation;
- feature flags;
- result and rollback.
History turns “something changed” into an actionable investigation.
Release version
Expose a release identity safely:
It helps correlate errors, performance, support reports, and rollbacks.
Release correlation
Telemetry should connect:
Do not collect identifying data unnecessarily. Use bounded, privacy-aware correlation IDs.
Deployment strategies
Common strategies include:
- rolling;
- blue-green;
- canary;
- feature-flagged release;
- full replacement with fast rollback.
Select based on compatibility, traffic, data migration, and operational capability.
Rolling deployment
Replace instances or assets gradually.
During the transition, old and new versions may coexist.
Ensure:
- API compatibility;
- shared asset availability;
- session behavior;
- migration safety;
- observability by version.
Backward compatibility during deployment
For a period, support:
Long-lived browser tabs, cached assets, service workers, and delayed requests make compatibility a real requirement.
Blue-green deployment
flowchart LR
Router["Global Edge Router / CDN"]
subgraph Blue["Blue Environment (Active v1.4)"]
B_App["Production App Servers"]
end
subgraph Green["Green Environment (Staged v1.5)"]
G_App["New Release Ready"]
end
Router -->|100% Active Traffic| Blue
Router -. Instant Cutover .-> GreenIt can enable fast switching and rollback.
Costs include duplicate capacity, data compatibility, and confidence that the inactive environment is genuinely ready.
Blue-green trade-offs
Consider:
- database and API state;
- cache warmth;
- background jobs;
- live connections;
- asset references;
- traffic switching and rollback timing.
Switching traffic does not undo side effects already performed by the new version.
Canary release
Expose a release to a small cohort and compare:
- errors;
- performance;
- task completion;
- support signals;
- conversion or domain outcomes.
Canaries require meaningful cohort identity and enough traffic to observe signal.
Architectural timeline of a production release
sequenceDiagram
autonumber
actor Dev as Engineer
participant CI as CI Pipeline
participant Reg as Immutable Registry
participant Edge as Edge CDN Router
participant Telemetry as RUM / Telemetry
Dev->>CI: Push Git commit (Tag v2.4.0)
CI->>CI: Run lint, types, unit, E2E gates
CI->>Reg: Publish immutable assets (hash: a9f3c1)
CI->>Edge: Deploy v2.4.0 to 10% Canary cohort
Edge-->>Telemetry: Stream real user Core Web Vitals & errors
Note over Telemetry: 15-Minute Observation Window
Telemetry-->>Edge: SLO Healthy (error rate < 0.05%)
Edge->>Edge: Promote v2.4.0 to 100% trafficConcrete failure & recovery scenario: Safari crash
sequenceDiagram
autonumber
participant Users as Safari 16 Users (10% Canary)
participant Edge as Edge CDN Router
participant RUM as Error Telemetry
participant OnCall as On-Call Engineer
Edge->>Users: Serve v2.4.0 bundle
Note over Users: Syntax Error: Unsupported Regex Lookbehind
Users-->>RUM: Spikes unhandled exceptions (5.8% error rate!)
RUM->>OnCall: PagerDuty Alert: Critical SLO breach on Safari
OnCall->>Edge: Trigger Instant Rollback to v2.3.0
Edge->>Users: Serve previous known-good bundle v2.3.0
Note over Users: Errors cease immediately (Recovery Time: 90s)
OnCall->>OnCall: Triage syntax target in tsconfig for v2.4.1Release strategy versus feature strategy
A release can ship code behind a flag.
A canary can expose a release to a cohort.
Do not treat the mechanisms as interchangeable.
Feature flags are operational controls
Flags can support:
- gradual rollout;
- experiments;
- kill switches;
- migration paths;
- permission-aware presentation.
Each flag needs an owner, purpose, lifecycle, and removal plan.
Feature flags are not just booleans
A flag may depend on:
- user or account;
- percentage cohort;
- environment;
- route;
- experiment assignment;
- release version;
- server-side policy.
Use a clear evaluation model and keep the result observable.
Vendor-neutral flag APIs
Keep product code dependent on a small capability interface rather than one vendor’s SDK throughout the application.
Flag evaluation location
Flags may evaluate:
Choose based on security, consistency, latency, and rollout needs.
Feature flags are not authorization
This controls exposure or experience.
The server must still authorize the delete operation.
Flag categories
Different categories need different owners, lifetimes, and audit rules.
Flag lifecycle
If a flag remains after its decision is complete, it becomes permanent branching complexity.
Flag debt
Old flags create:
- dead paths;
- extra tests;
- confusing support behavior;
- inconsistent analytics;
- migration risk.
Track flag removal like technical debt with an owner and due condition.
Kill switches
A kill switch disables a harmful or expensive capability quickly.
It should have:
- safe fallback;
- owner;
- access control;
- audit trail;
- test coverage;
- clear recovery process.
A kill switch is useful only if operators can trust it under pressure.
Rollback
Rollback restores a previous known release or behavior.
It requires:
- an identifiable artifact;
- compatible data and APIs;
- deployment access;
- tested procedure;
- observability to confirm recovery.
Rollback is a normal safety mechanism, not an admission of failure.
Rollback must be practiced
A document is not evidence that rollback works.
Rehearse:
- who decides;
- what command or action occurs;
- which version returns;
- how flags change;
- what users experience;
- how verification confirms recovery.
Practice reveals missing access, compatibility, and monitoring assumptions.
Database and API compatibility
A frontend rollback may meet:
- a changed API;
- a migrated schema;
- a new required field;
- an invalidated cache;
- an incompatible service worker.
Use expand-and-contract changes and backward-compatible periods where rollback matters.
Rollback versus forward fix
flowchart TD
Incident["Production Incident Detected"] --> Q1{"Is root cause obvious<br/>and fix trivial (<5 min)?"}
Q1 -->|Yes, zero risk| Forward["Forward Hotfix via CI Pipeline"]
Q1 -->|No / High Risk| Q2{"Was client storage or<br/>DB schema mutated?"}
Q2 -->|Backward Compatible| Rollback["Instant Artifact / CDN Rollback"]
Q2 -->|Schema Broken| Kill["Activate Feature Kill Switch & Triage"]Choose based on blast radius, data effects, detection certainty, and fix confidence.
Observability
Observability asks whether internal system state can be understood from emitted evidence.
For frontend systems, evidence can include:
- errors;
- performance events;
- network failures;
- route transitions;
- release identity;
- user journey markers;
- feature-flag context.
Monitoring versus observability
You need both:
- dashboards and alerts for known failure;
- logs, traces, context, and exploration for unfamiliar failure.
Three traditional telemetry signals
flowchart TD
Logs["Logs<br/>Structured contextual events & breadcrumbs"]
Metrics["Metrics<br/>Aggregated counters, rates, and Web Vitals percentiles"]
Traces["Traces<br/>End-to-end distributed execution paths across client and API"]
Logs --- Metrics --- TracesFrontend observability adapts these signals to browser privacy, lifecycle, and network constraints.
Logs
Useful frontend logs are:
- structured;
- bounded;
- correlated to release and journey;
- privacy-aware;
- actionable.
Do not turn the browser console into a production data lake.
Browser console is not production observability
Console output can disappear because:
- the user closes the page;
- the browser filters it;
- no one can access the user’s console;
- context is incomplete;
- sensitive data was logged.
Send carefully designed events to a controlled telemetry system when investigation requires it.
Metrics
Frontend metrics can include:
- error rate;
- route load time;
- Web Vitals;
- interaction timing;
- request failure rate;
- feature adoption;
- queue depth or sync status.
Define units, aggregation, segment, and action before collecting a metric.
High cardinality
Values such as raw URLs, user IDs, or arbitrary error messages can create too many metric dimensions.
Prefer bounded labels:
Keep detailed context in traces or sampled events with privacy controls.
Traces and spans
flowchart TD
Root["User Journey: Checkout Submission"]
Root --> S1["Span 1: Form Validation & Client State Update"]
Root --> S2["Span 2: HTTP POST /api/v1/checkout (traceparent)"]
S2 --> S3["Span 3: Backend Gateway & Payment Provider"]
Root --> S4["Span 4: DOM Paint & Confirmation View Render"]Traces connect frontend work with backend and network events.
They are especially useful for slow or distributed journeys.
Frontend tracing has special challenges
The browser has:
- intermittent sessions;
- sampling constraints;
- privacy boundaries;
- offline periods;
- multiple tabs;
- limited background time;
- user-controlled execution.
Design telemetry that remains useful without assuming server-like process lifetime.
OpenTelemetry
OpenTelemetry provides a vocabulary and instrumentation approach for traces, metrics, and logs.
Browser support and semantic coverage continue to evolve.
Use it as an interoperability tool, not a promise that every frontend detail is automatically standardized.
Real User Monitoring revisited
RUM should connect:
Collect enough context to act while avoiding unnecessary personal or form data.
Error monitoring
Error monitoring should capture:
- error type and message;
- stack with private source maps;
- release identity;
- route and operation;
- browser context;
- breadcrumbs;
- safe user and feature context.
Group similar failures so teams can prioritize rather than count noise.
Global error handlers
Global handlers can catch unexpected failures and report them.
They should not:
- hide the failure without a recovery UI;
- send secrets;
- duplicate every error repeatedly;
- replace local error boundaries;
- create a false sense that all errors are observable.
Source maps in production monitoring
Source maps make minified errors actionable.
Prefer private upload to the error-monitoring system rather than public deployment when exposure is not needed.
Tie maps to exact release artifacts.
Error grouping
Group by stable cause signals:
- error type;
- normalized message;
- stack location;
- operation;
- release.
Avoid grouping all failures into one generic event or splitting one bug into thousands of dynamic messages.
Error rate needs context
Define:
An error count alone does not tell whether users are affected more or fewer.
Network error monitoring
Capture categories such as:
- offline;
- timeout;
- DNS or connection;
- CORS;
- 401/403;
- 404;
- 429;
- 5xx;
- schema failure;
- cancellation.
Different categories need different owners and actions.
Client versus server error
The same user-visible failure may require evidence from both sides.
Correlate requests and releases safely.
User context without oversharing
Useful context might be:
- anonymous session ID;
- account tier or role category;
- route template;
- release;
- feature flag assignment.
Avoid raw form values, tokens, passwords, private content, and unnecessary identity data.
Breadcrumbs
Breadcrumbs can record a bounded sequence such as:
They should explain the journey without recording sensitive payloads or growing without limit.
Session replay
Session replay has strong privacy and compliance implications.
Define:
- consent;
- masking;
- field exclusion;
- retention;
- access;
- sensitive-route policy;
- sampling.
Debugging value does not override user privacy.
Observability sampling
Sampling controls:
- cost;
- storage;
- user impact;
- signal volume.
Sample ordinary success heavily, but retain enough rare failures and high-severity journeys to investigate them.
Head-based versus tail-based sampling
Tail sampling can retain slow or failed traces more effectively but requires infrastructure and buffering.
Telemetry has performance cost
Telemetry can add:
- network requests;
- serialization;
- CPU;
- memory;
- storage;
- privacy review.
Instrument critical signals and batch or sample responsibly.
Telemetry buffering and sendBeacon
When a page is closing, a beacon can send small analytics or diagnostic payloads without blocking navigation.
It is not a guarantee of delivery and should not carry sensitive or large data.
Observability naming
Use stable names:
Consistent naming makes dashboards and searches usable across releases.
Operational dashboards
A dashboard should show:
- current release;
- errors and affected sessions;
- performance signals;
- traffic and sample size;
- top routes and operations;
- recent deployments;
- feature flags and cohorts;
- known incidents.
Dashboards should support a decision, not display every metric.
Alerting
An alert should specify:
- condition;
- severity;
- owner;
- response time;
- runbook;
- suppression or grouping;
- recovery signal.
If nobody knows what to do after an alert, it is likely not ready to page.
Alert fatigue
Too many low-value alerts cause operators to ignore high-value ones.
Reduce fatigue with:
- actionable thresholds;
- grouping;
- maintenance windows;
- ownership;
- severity tiers;
- periodic review.
Static thresholds versus anomaly detection
Static thresholds are understandable for known limits.
Anomaly detection can reveal unusual changes relative to baseline.
Both need context, sample size, and a response plan. Sophisticated detection without action is noise.
Service-level indicators
An SLI is a measured aspect of service behavior:
Define the numerator, denominator, population, and measurement boundary.
Service-level objectives
An SLO sets a target for an SLI over a time window.
Examples:
- 99.9% of save attempts receive a usable result;
- 95% of catalogue routes meet a journey threshold;
- 99% of releases pass smoke verification.
The objective should be meaningful to users and operators.
Error budgets
The budget can inform release pace, reliability investment, and risk acceptance.
It should not be used to justify harming a small but important user group.
Frontend SLOs need care
Frontend metrics are affected by:
- user devices;
- networks;
- browser extensions;
- third-party scripts;
- sampling;
- route and feature mix.
Define scope and segments so the objective reflects what the team can influence and what users experience.
Health checks
A frontend health check may verify:
- static asset availability;
- route response;
- API reachability;
- configuration;
- basic rendering;
- deployment identity.
It should be safe, bounded, and not confuse “server responds” with “user journey works”.
Synthetic production monitoring
Synthetic checks can periodically perform a safe journey:
They can detect outages before enough real users generate field data.
Synthetic monitoring trade-offs
Synthetic traffic can:
- miss real device diversity;
- create false confidence;
- add load;
- require test identities;
- fail because of test data rather than product behavior.
Use it alongside RUM and clear ownership.
Deployment verification
After release, check:
- artifact and version;
- critical route;
- authentication;
- API requests;
- assets and service worker;
- errors and performance;
- feature flag exposure.
Deployment is not complete until the user-facing path is verified.
Automatic rollback
Automatic rollback can be appropriate when:
- the signal is reliable;
- the failure threshold is clear;
- the previous artifact is compatible;
- rollback is safe;
- operators are notified;
- a forward fix path exists.
Do not automate a destructive or ambiguous recovery.
Progressive delivery
At each step, observe release health and decide whether to continue, pause, or revert.
Release cohorts
Cohorts should be:
- stable enough to compare;
- representative enough to reveal risk;
- privacy-aware;
- consistently evaluated;
- identifiable in telemetry.
Random percentage alone is not enough if users move between cohorts unpredictably.
Percentage rollout consistency
Hash a stable identity or account key so a user does not switch variants on every request.
Handle anonymous users and identity changes deliberately.
Document how rollout interacts with caching and server/client evaluation.
Experimentation
Experiments need:
- hypothesis;
- assignment policy;
- primary metric;
- guardrail metrics;
- sample and duration;
- analysis plan;
- stop condition;
- cleanup plan.
An experiment is not a permanent feature flag.
Guardrail metrics
Alongside product outcomes, monitor:
- errors;
- latency;
- accessibility failures;
- support contacts;
- abandonment;
- memory or resource use;
- security signals.
Stop an experiment quickly when it harms users even if the primary metric looks promising.
Maintenance is architecture
The system changes after launch:
- dependencies update;
- browsers change;
- APIs evolve;
- storage persists;
- service workers remain installed;
- teams and ownership change.
Maintenance paths must be designed, not improvised after years of drift.
Dependency maintenance
Maintain dependencies with:
- regular review;
- security triage;
- compatibility testing;
- controlled updates;
- rollback or pinning path;
- removal of unused packages.
Never updating is also a risk strategy, usually a poor one.
Update continuously, not once every three years
Small frequent updates reduce:
- migration distance;
- surprise incompatibilities;
- debugging scope;
- security exposure;
- ownership uncertainty.
They still require review and release discipline.
Dependency update automation
Automation can open updates and run checks.
It should not merge every update blindly.
Use grouping, ownership, security priority, release notes, and artifact/performance review.
Grouping updates
Group compatible low-risk updates to reduce review noise.
Keep risky or architectural updates separate so failures and migration decisions remain understandable.
Security advisories need triage
A vulnerability report needs:
- affected package and path;
- whether the code ships or executes;
- exposure and exploitability;
- available fix;
- mitigation;
- owner and deadline.
Severity score alone does not decide urgency.
Transitive dependencies
You may not import a package directly and still ship it through a dependency chain.
Track:
- why it exists;
- who owns the direct dependency;
- whether it is reachable in production;
- how updates propagate;
- whether it can be removed.
Removing dependencies
Removing a dependency can improve:
- bundle size;
- security surface;
- build time;
- maintenance;
- licensing clarity.
Verify that a platform capability or small local implementation truly meets the requirement before replacing it.
Browser support maintenance
Browser policy affects:
- transformation targets;
- polyfills;
- testing matrix;
- CSS behavior;
- support cost;
- user access.
Do not drop support from global statistics alone if your product serves a meaningful affected population.
Deprecating browser support
Use a planned lifecycle:
Provide a clear reason and verify that the release does not strand critical users unexpectedly.
API maintenance
Maintain APIs with:
- backward-compatible additions;
- versioned breaking changes;
- schema monitoring;
- deprecation windows;
- client compatibility;
- migration communication.
Frontend and backend releases often overlap in time.
Schema monitoring
Monitor:
- unknown fields;
- missing required fields;
- invalid types;
- deprecated fields still used;
- response version distribution.
Runtime validation turns silent drift into observable evidence.
Data migration and frontend compatibility
During migration, old clients may still send old shapes.
Use additive changes, compatibility periods, and server-side normalization where rollback and long-lived clients matter.
Local storage migration
Persisted client state needs:
- version;
- schema validation;
- migration path;
- reset behavior;
- identity scope;
- failure handling.
Never assume all users have the current local schema.
Service-worker maintenance
Service workers can outlive deployments.
Maintain:
- cache versioning;
- activation policy;
- update notification;
- old client compatibility;
- cache cleanup;
- offline data migration.
The installed worker is part of the deployed client population.
The stale-client problem
Users may keep a tab open while a new release ships.
The old client may call:
- a changed API;
- a removed route;
- an incompatible event;
- an old asset path.
Design compatibility and update behavior for long-lived sessions.
Update notification
An update message should explain:
- what changed;
- whether current work is safe;
- when reload is appropriate;
- whether the user can defer;
- what happens to drafts and connections.
Do not reload unexpectedly while the user is editing important data.
Technical debt
Debt is the future cost of a current shortcut, including:
- missing tests;
- unsupported dependencies;
- unclear ownership;
- stale flags;
- undocumented platform behavior;
- brittle deployment assumptions.
Debt is not automatically bad; unmanaged debt is.
Debt register
Track:
- debt item;
- user or team impact;
- trigger for action;
- owner;
- risk;
- estimated effort;
- expiry or review date.
An explicit register turns vague concern into prioritizable work.
Maintenance budget
Reserve capacity for:
- dependency updates;
- browser support;
- observability changes;
- flag cleanup;
- API migration;
- service-worker maintenance;
- refactoring and test improvement.
If maintenance has no budget, it will compete with emergencies.
End-of-life policy
Define how the system retires:
- old browser versions;
- APIs;
- packages;
- feature flags;
- service-worker caches;
- preview environments;
- telemetry schemas.
Retirement is part of lifecycle architecture.
Ownership and runbooks
A runbook should state:
- what signal indicates the problem;
- how to confirm it;
- immediate mitigation;
- rollback or kill switch;
- communication path;
- recovery verification;
- follow-up owner.
Runbooks reduce decision load during incidents.
Incident response lifecycle
flowchart LR
Det["1. Detect<br/>SLO Alert"] --> Tri["2. Triage<br/>Assess Blast Radius"]
Tri --> Mit["3. Mitigate<br/>Rollback / Kill Switch"]
Mit --> Com["4. Communicate<br/>Status Page Update"]
Com --> Rec["5. Recover<br/>Verify Telemetry Normal"]
Rec --> Lrn["6. Learn<br/>Blameless Post-Mortem"]Keep the user impact and current system state visible throughout the incident.
Blameless learning
Post-incident review should ask:
- what conditions made the failure possible?
- which signal was missing or late?
- which recovery step worked or failed?
- what system change reduces recurrence?
The goal is stronger systems, not individual blame.
Post-incident actions
Actions should be:
- specific;
- owned;
- prioritized;
- measurable;
- connected to the failure;
- reviewed for completion.
“Be more careful” is not a system improvement.
Maintenance metrics
Useful signals include:
- mean time to recovery;
- deployment failure rate;
- rollback time;
- dependency age;
- stale flag count;
- unresolved vulnerability age;
- test and build feedback time;
- observability coverage.
Do not turn one metric into a target that damages the system.
Deployment frequency is not quality by itself
Frequent deployment can indicate healthy delivery or uncontrolled churn.
Pair it with:
- change failure rate;
- recovery time;
- user impact;
- reliability and performance;
- maintenance health.
Mean time to recovery
MTTR asks how quickly the system returns to an acceptable state after failure.
Improve it through:
- detection;
- clear ownership;
- safe rollback;
- kill switches;
- runbooks;
- practiced communication.
Prevention matters, but recovery is part of reliability.
Observability completes testing
Tests cover known scenarios.
Production observability reveals:
- unknown combinations;
- real devices and networks;
- integration drift;
- deployment-specific failures;
- user behavior outside assumptions.
Use production evidence to improve the next test strategy.
The production feedback loop
Production is not the end of engineering; it is a source of evidence.
Privacy and compliance
Telemetry and debugging must respect:
- data minimization;
- consent;
- retention;
- access control;
- regional requirements;
- deletion and subject rights;
- sensitive fields.
Operational visibility cannot be purchased by recording everything.
Do not record form contents by default
Forms can contain:
- passwords;
- financial data;
- health information;
- personal identifiers;
- private business content.
Mask or exclude inputs and inspect telemetry payloads before production use.
Session replay needs stronger governance
Replay can show real failures, but may capture private content and interactions.
Define:
- masking defaults;
- consent and regional policy;
- sensitive route exclusions;
- retention and access;
- sampling and incident use.
Analytics versus observability
They may share infrastructure, but have different purpose, privacy, retention, and ownership.
Product versus operational metrics
| Product | Operational |
|---|---|
| conversion | error rate |
| task completion | latency |
| retention | availability |
| feature adoption | deployment health |
Connect them carefully without assuming one causes the other.
Business-aware observability
A useful release dashboard can connect:
This helps teams prioritize an outage that affects a critical workflow over a noisy low-impact error.
Release dashboard example
Show:
- current and previous release;
- deployment status;
- error and performance deltas;
- critical journey success;
- flag cohorts;
- open incidents;
- rollback readiness.
Keep the dashboard decision-oriented.
Deployment health window
After release, observe a defined window with:
- expected traffic;
- known baseline;
- key route metrics;
- error and support signals;
- feature exposure;
- rollback threshold.
Do not declare health from one successful smoke test alone.
Low-traffic products
Sparse traffic makes automatic detection harder.
Use:
- longer observation windows;
- synthetic checks;
- targeted smoke tests;
- manual review;
- confidence-aware thresholds;
- support and user reports.
Low volume does not mean low importance.
Maintenance and architecture decisions
Prefer architectures with:
- understandable upgrade paths;
- supported tooling;
- observable boundaries;
- reversible deployment;
- clear ownership;
- proportionate operational cost.
The most elegant architecture is not useful if nobody can maintain or recover it.
Boring technology has value
Mature, documented, well-understood tools can reduce:
- operational surprise;
- onboarding cost;
- incident ambiguity;
- upgrade risk;
- dependence on one expert.
Novelty should solve a real constraint, not create one.
Upgrade path is a selection criterion
Before adopting a tool, ask:
- how is it upgraded?
- how are breaking changes announced?
- can output be inspected?
- is ownership clear?
- can the system roll back?
- what happens when the maintainer changes?
Current features are only one part of technology choice.
Avoid undocumented internal platforms
An internal platform without:
- public contracts;
- documentation;
- release notes;
- support ownership;
- migration path;
- observability;
becomes a hidden dependency and a bottleneck for every product team.
Operational readiness review
Review:
The checklist should be proportional to user impact and system complexity.
Production readiness is proportional
A small static page and a financial workflow need different readiness depth.
Scale review according to:
- user harm;
- data sensitivity;
- change frequency;
- dependency count;
- availability expectations;
- recovery cost.
Proportionate does not mean careless.
Practical lab: Delivery, Observability, and Rollback Loop
Design and rehearse a safe frontend release from commit to deployment, observation, rollback, and cleanup.
The practical makes the delivery artifact, release identity, flags, telemetry, alerts, and recovery path explicit.
Practical stages 1–6: build and deploy safely
- Create the CI pipeline.
- Add parallel jobs.
- Produce a versioned artifact.
- Create a preview deployment.
- Protect production.
- Add deployment concurrency.
Verification: the tested artifact is the deployed artifact.
Practical stages 7–13: release control
- Expose the release version.
- Add a preview smoke test.
- Design a canary rollout.
- Add a release flag.
- Give the flag an owner.
- Create a kill switch.
- Remove a completed flag.
Feature flags do not act as authorization.
Practical stages 14–20: observe production
- Add error monitoring.
- Upload source maps privately.
- Add RUM.
- Add a custom journey metric.
- Create an operational dashboard.
- Define alerts.
- Add synthetic production monitoring.
Every alert needs an owner and an operational response.
Practical stages 1–3: reproducible artifact & preview verification
- Stage 1 (Immutable Reproducible Artifact):
- Produce a production build with deterministic content hashes.
- Generate
release-manifest.jsonwith commit SHA, timestamp, and metadata.
- Stage 2 (CI Verification Gates & Secrets Scanning):
- Execute parallel linting, type checks, unit/integration suites, and bundle budgets.
- Run automated secret scanning to prevent token leaks into client bundles.
- Stage 3 (Preview Environments & Release Identity):
- Deploy isolated PR preview environments.
- Inject
window.__RELEASE_INFO__for runtime telemetry attribution.
Practical stages 4–5: observability & rehearsed rollback
- Stage 4 (Front-End Observability & Breadcrumbs):
- Implement zero-dependency client telemetry for unhandled errors and RUM metrics.
- Capture user interaction breadcrumbs with strict PII masking.
- Stage 5 (Simulated Disaster Rehearsal & Safe Rollback):
- Inject a deliberate production failure into v2.4 (breaking Safari form submissions).
- Observe automated SLO breach and trigger instant CDN rollback to v2.3.
- Verify client data compatibility (
localStorage) and draft a blameless post-mortem.
Practical extension: frontend runbook
Write a short runbook covering:
Include commands, owners, dashboards, thresholds, and the evidence that confirms recovery.
Try this yourself
For the next release, write:
If any field is blank, the release loop has an unexamined assumption.
Troubleshooting guide (Part 1)
| Symptom | Likely cause |
|---|---|
| CI passes but deployed app fails | Tested artifact differs from deployed artifact |
| Rollback restores old code but breaks API | Compatibility window was not designed |
| Alerts fire constantly | Thresholds lack context or ownership |
| Errors cannot be debugged | Release identity or private source maps missing |
| Feature flag remains forever | Lifecycle and owner were never defined |
Troubleshooting guide (Part 2)
| Symptom | Likely cause |
|---|---|
| Telemetry is expensive or unsafe | Payload, sampling, and privacy policy are weak |
| Users keep stale behavior | Long-lived clients and service-worker updates ignored |
| Dependency update is frightening | Updates were deferred instead of maintained continuously |
| Incident response is slow | Runbook and rollback were not practiced |
Completion checklist
- CI produces a reproducible, identifiable artifact;
- the same artifact is promoted across environments;
- production access and secrets use least privilege;
- releases can be canaried, flagged, killed, or rolled back;
- telemetry connects release, route, journey, and failure;
- alerts have owners and runbooks;
- SLOs and error budgets are meaningful and segmented;
- dependencies, browsers, APIs, storage, and workers have maintenance paths;
- telemetry respects privacy and data minimization;
- rollback and recovery have been rehearsed.
Misconceptions to leave behind (Part 1)
| Misconception | Better mental model |
|---|---|
| CI means one hosted test command | CI is clean verification and artifact production |
| Delivery means every commit deploys | Delivery keeps a safe release ready |
| Passing tests makes deployment safe | Artifact, environment, integration, and recovery also matter |
| Staging is production without users | Environments have different purposes and gaps |
| Preview replaces code review | It provides deployed evidence, not design judgment |
| Feature flags are authorization | The server still enforces permissions |
| Flags can stay forever | Flags need ownership and removal |
Misconceptions to leave behind (Part 2)
| Misconception | Better mental model |
|---|---|
| Rollback is failure | Rollback is a normal safety mechanism |
| Observability means logs | Logs, metrics, traces, errors, and context work together |
| More telemetry is always better | Telemetry has cost, privacy, and signal limits |
| Every error should page someone | Alerts require actionable owners and thresholds |
| Dependency updates should be automatic | Automation needs review, grouping, and risk triage |
| Users refresh after deployment | Long-lived clients require compatibility and update policy |
| Maintenance is separate from architecture | Lifecycle and recovery are architectural properties |
The chapter in one sentence
Ship reproducible artifacts through a controlled release loop, observe real user impact, recover deliberately, and budget continuously for maintenance.
Next: Chapter 18
The next chapter will build on production architecture with:
- final system integration;
- architectural decision records;
- capstone planning;
- cross-cutting quality and delivery review;
- a complete front-end platform blueprint.
Questions
If the next release harms users, how will you identify the exact artifact, detect the impact, reduce exposure, roll back safely, and learn what the system failed to tell you?
18 - Front-End Architecture & Technical Decision-Making
Front-End Architecture & Technical Decision-Making
Choose deliberately, learn continuously
Chapter 18
Polla Fattah
Today’s goal
Turn architecture from a collection of technologies into a transparent set of decisions.
We will connect:
- requirements, quality attributes, constraints, and trade-offs;
- cohesion, coupling, dependency direction, and blast radius;
- complexity budgets, reversibility, locality, and explicit data flow;
- progressive enhancement, platform-first design, and dependency evaluation;
- fitness functions, ADRs, spikes, risk reduction, and failure modes;
- performance, security, accessibility, testing, observability, and deployability;
- team topology, governance, migration, and architecture outcomes.
By the end of today you can
- define architecture as decisions made under constraints;
- distinguish functional requirements from quality attributes;
- compare alternatives without fake precision;
- reduce coupling and blast radius through boundaries;
- prefer reversible decisions when information is weak;
- use ADRs to record decisions and revisit triggers;
- turn important architecture rules into fitness functions;
- choose a proportionate architecture for different scenarios;
- design migration instead of assuming rewrites are cleaner;
- defend the smallest coherent solution with evidence.
The central principle
Architecture is a set of explicit, contextual decisions about trade-offs, boundaries, and future change - not a technology stack or a prediction of everything that might happen.
Good architecture makes important change cheaper, safer, and easier to reason about.
The decision progression
flowchart TD
Prob["1. Define Problem & Context"] --> Qual["2. Identify Quality Attributes"]
Qual --> Const["3. Record Constraints"]
Const --> Alt["4. Generate Alternatives"]
Alt --> Spike["5. Reduce Unknowns via Spikes"]
Spike --> ADR["6. Decide & Document in ADR"]
ADR --> Guard["7. Enforce via CI Guardrails"]
Guard --> Rev["8. Observe & Revisit with Evidence"]
Rev -. Feeds into new decisions .-> ProbArchitecture remains a loop rather than a one-time ceremony.
Architecture is a set of decisions
Examples:
The framework is one input to those decisions, not the decision itself.
Not everything can be optimized at once
Trade-offs may exist between:
- autonomy and consistency;
- speed and flexibility;
- freshness and cacheability;
- simplicity and isolation;
- performance and feature richness;
- delivery independence and runtime coupling.
Architecture makes the trade-off visible so the team can choose intentionally.
Maximum team autonomy has costs
Autonomy can increase:
- release speed;
- local decision quality;
- ownership;
- experimentation.
It can also increase:
- duplication;
- inconsistent behavior;
- platform cost;
- integration burden;
- fragmented user experience.
The goal is useful autonomy within coherent contracts.
Quality attributes
Quality attributes describe how the system behaves:
They are architecture inputs because they shape boundaries and decisions.
Functional requirements versus quality attributes
Functional requirements describe capability.
Quality attributes describe the conditions under which that capability remains useful.
Not every quality attribute has equal priority
A government service may prioritize accessibility and reliability.
An internal analytics tool may prioritize iteration speed and data accuracy.
A marketing page may prioritize cacheability and loading performance.
Priorities should be explicit rather than assumed.
Make quality attributes measurable
An attribute becomes architectural guidance when the team can observe it.
Constraints
Constraints can come from:
- budget;
- deadline;
- existing APIs;
- browser support;
- team skills;
- regulation;
- deployment platform;
- organization;
- data sensitivity.
Constraints are not annoyances to ignore; they define the decision space.
Requirements, constraints, and decisions
flowchart TD
subgraph DecisionForces["The Forces Shaping Architecture"]
Req["Functional Requirements<br/>(What the system must achieve)"]
Qual["Quality Attributes<br/>(How well the system must perform: LCP, a11y, MTTR)"]
Const["Constraints<br/>(Budget, team size, legacy APIs, legal compliance)"]
Req --- Dec["Architecture Decision<br/>(Selected structure & accepted trade-offs)"]
Qual --- Dec
Const --- Dec
endConfusing a preference with a constraint produces unnecessary architecture.
Architecture is contextual
The same question can have different answers:
Context is part of correctness.
Start with the problem, not the tool
Weak question:
Stronger question:
The tool choice follows the problem model.
Technology selection is downstream
flowchart LR
Forces["Requirements + Quality + Constraints"] --> Boundaries["Boundaries & Operating Model"]
Boundaries --> Candidates["Evaluate 3+ Technology Candidates"]
Candidates --> Spike["Empirical Spike & Evidence"]
Spike --> Selection["Committed Technology Selection"]Choosing a library before defining the need turns a decision into a justification exercise.
Architecture decisions are often about boundaries
Decide where to place:
- state;
- trust;
- rendering;
- ownership;
- caching;
- deployment;
- failure;
- team responsibility.
Boundaries determine what can change independently.
Good boundaries reduce change cost
The boundary should expose a stable contract and keep volatile implementation private.
Cohesion and coupling
flowchart LR
subgraph HighCohesion["High Cohesion (Desirable)"]
direction TB
C1["Route Logic"] <--> C2["UI View"]
C2 <--> C3["Local State"]
Note1["Changes together for one feature"]
end
subgraph LowCoupling["Low Coupling (Desirable)"]
direction LR
FeatureA["Catalogue Feature"] <-- Narrow API Contract --> FeatureB["Checkout Feature"]
Note2["Changes in A do not break B"]
endGood architecture seeks high cohesion inside meaningful units and intentional coupling between them.
Cohesion example
These concerns may belong together because they change around the same user capability.
Coupling example
The card now knows more than its responsibility requires.
Pass a product and an intent-oriented capability instead.
Dependency direction
Depend toward stability:
flowchart TD
App["Application Entrypoint / Shell<br/>(Most Volatile)"] --> Feat["Feature Modules (Catalogue, Permits)"]
Feat --> Domain["Domain Rules & State Reducers"]
Domain --> Found["Shared Foundation & Design Tokens<br/>(Most Stable)"]Dependencies should flow toward stable, reusable concepts.
When a foundation imports application behavior, the architecture becomes difficult to reuse and change.
Stable dependencies
Depend on:
- small interfaces;
- domain contracts;
- platform capabilities;
- stable shared primitives.
Avoid depending on:
- private files;
- temporary implementations;
- broad containers;
- undocumented framework internals.
Stable dependencies reduce future blast radius.
Blast radius
Blast radius asks:
Reduce blast radius through:
- narrow APIs;
- isolated data and runtime boundaries;
- feature flags;
- fallback behavior;
- independent tests and deployment;
- clear ownership.
Centralization versus local autonomy
Centralize when consistency, security, or shared lifecycle matters.
Keep local when:
- behavior is domain-specific;
- requirements differ;
- coordination cost exceeds reuse value;
- independent evolution is important.
The right location is a trade-off, not a moral position.
The complexity budget
Every system spends complexity on:
- code;
- tools;
- runtime;
- deployment;
- operations;
- team coordination;
- cognitive load.
Spend complexity only where it buys a required quality attribute.
Accidental versus essential complexity
Architecture work should remove accidental complexity without pretending essential complexity can disappear.
Complexity has several forms
A “simpler” client architecture may move complexity into the server or CI system.
Count the whole system.
Prefer reversible decisions
When information is weak, prefer choices that can be changed without rewriting the entire platform.
Examples:
- explicit adapter around a vendor;
- route-level boundary before full micro-frontends;
- local state before global store;
- package API before runtime federation.
Reversibility buys learning time.
One-way and two-way doors
Categorize decisions by reversibility:
flowchart LR
subgraph TwoWay["Two-Way Door (Reversible)"]
direction TB
D1["Decision: UI styling library / Local state tool"]
D1 --> E1["Easy to migrate or revert"]
D1 --> A1["Rule: Decide quickly, test in production"]
end
subgraph OneWay["One-Way Door (Hard to Reverse)"]
direction TB
D2["Decision: Micro-frontends / Core database schema"]
D2 --> E2["Costly, multi-month migration to undo"]
D2 --> A2["Rule: Require spikes, ADRs & executive sign-off"]
endDo not apply heavyweight governance to every reversible choice.
Do not rush a decision that creates irreversible migration or data cost.
Delay irreversible decisions when information is weak
Use:
- a spike;
- a prototype;
- a compatibility layer;
- a small route;
- a feature flag;
- an explicit revisit trigger.
Learning before commitment is architecture work.
YAGNI does not mean ignore the future
Avoid building speculative systems for unknown requirements.
Do preserve:
- clear boundaries;
- migration paths;
- explicit ownership;
- replaceable dependencies;
- observable assumptions.
The future is supported by flexibility, not by implementing every possibility now.
Premature abstraction
An abstraction created before variation is understood may encode the wrong concept.
Wait for evidence about:
- repeated behavior;
- stable vocabulary;
- change direction;
- real consumers;
- shared accessibility and lifecycle.
The rule of three
One use may be local.
Two uses reveal similarity.
Three uses can provide enough evidence to decide whether the concept is genuinely shared.
It is a heuristic, not a command to duplicate exactly three times.
Duplication can be cheaper than coupling
Small duplicated code may be safer than:
- a generic API with many modes;
- a shared release dependency;
- cross-domain assumptions;
- a package nobody can evolve.
Compare maintenance and change cost rather than counting lines.
Locality
Locality keeps related code, data, tests, and ownership near each other.
It reduces:
- navigation cost;
- hidden dependencies;
- coordination overhead;
- accidental reuse.
Move code outward only when the new boundary creates real value.
Explicit data flow
Explicit inputs, outputs, events, URLs, and requests are easier to test and change than invisible reads from global context.
Hidden coupling
Hidden coupling appears through:
- global mutable state;
- import-time side effects;
- shared storage keys;
- undocumented events;
- CSS selectors across packages;
- framework-specific assumptions;
- environment variables used everywhere.
Make coupling visible or remove it.
Global state is an architectural decision
Use global state when:
- lifetime is application-wide;
- consumers are genuinely distributed;
- synchronization is needed;
- ownership is explicit.
Do not use it to avoid passing one value through a well-defined local boundary.
Server state is not ordinary client state
Server state has:
- remote authority;
- freshness;
- errors;
- caching;
- invalidation;
- concurrency.
Use a server-aware boundary rather than copying it into every client store.
Derived state should remain derived
Store the inputs and calculate the result.
Duplicated derived state creates another source of truth and another synchronization path.
URL as architecture
URL state provides:
- shareability;
- reload persistence;
- browser history;
- deep links;
- route ownership.
It should contain meaningful public view state, not secrets or every ephemeral interaction.
Progressive enhancement as a layered model
Layering can improve resilience and reduce the amount of functionality that depends on one runtime.
When progressive enhancement is valuable
It is especially useful for:
- public content;
- forms and navigation;
- poor connectivity;
- accessibility;
- long-lived pages;
- critical tasks.
Not every application needs full no-JavaScript operation, but every app should understand its essential failure path.
Native platform first
Before adding a dependency, ask whether the platform already provides:
- links;
- forms;
- dialog;
- details/summary;
- URL and history;
- storage;
- fetch;
- observers;
- workers.
Platform capabilities often have strong accessibility, performance, and compatibility foundations.
Native does not automatically mean better
Evaluate:
- browser support;
- accessibility behavior;
- required customization;
- interaction consistency;
- testing;
- polyfills and fallbacks;
- team familiarity.
Use the platform deliberately, not ideologically.
Dependency evaluation
Consider:
- capability fit;
- bundle and runtime cost;
- security and maintenance;
- accessibility;
- license and ecosystem;
- upgrade path;
- lock-in;
- team skill.
The smallest dependency is not always the cheapest total solution.
Build, buy, or adopt
Compare total lifecycle cost, not only initial implementation time.
Core versus commodity
Protect and understand capabilities that differentiate the product.
Avoid spending strategic team capacity reinventing commodity infrastructure unless the trade-off is intentional.
Also avoid outsourcing a core capability that determines security, user trust, or domain advantage without sufficient control.
Framework selection is contextual
Evaluate frameworks through:
- rendering and data needs;
- team familiarity;
- ecosystem and support;
- accessibility;
- testing;
- deployment;
- upgrade path;
- performance on target users.
Benchmark claims without context are weak evidence.
Framework benchmarks are contextual
Results vary by:
- application shape;
- route size;
- data and interaction;
- device and network;
- build configuration;
- developer usage.
Run a spike with representative behavior instead of selecting a framework from a leaderboard.
React versus Vue is often not the main decision
The larger questions may be:
- where server and client boundaries lie;
- how state is owned;
- how routes are delivered;
- how APIs are validated;
- how teams release;
- how failures recover.
Framework syntax rarely determines all architecture.
Framework lock-in
Lock-in can be acceptable when the framework provides strong value and the team accepts its lifecycle.
Reduce unnecessary lock-in at boundaries with:
- domain models;
- adapters;
- platform APIs;
- stable package contracts;
- route and API boundaries.
Avoid turning lock-in avoidance into architecture by itself.
Architecture fitness
Fitness asks whether the system continues to support important properties as it evolves.
Examples:
- no dependency cycles;
- route budget stays within threshold;
- public package imports only;
- all remote responses are validated;
- critical journey remains keyboard-usable.
Evolutionary architecture
Instead of assuming the initial architecture is perfect:
Architecture can evolve safely when important properties are observable and protected.
Fitness functions
Automate checks that preserve properties continuously:
flowchart LR
Code["Pull Request Code Push"] --> Linter["ESLint Boundary Rules<br/>(No feature-to-feature imports)"]
Code --> Budget["Bundle Budget Checker<br/>(Initial route < 180 KB)"]
Code --> Schema["Contract Validator<br/>(Validates Zod schemas)"]
Linter --> Gate{"CI Gate"}
Budget --> Gate
Schema --> Gate
Gate -->|All Pass| Merge["Merge Approved"]
Gate -->|Any Fail| Block["Block PR Deployment"]Automate rules that matter enough to protect continuously.
Architecture rule as code
Rules can be enforced through:
- lint boundaries;
- package exports;
- dependency graph checks;
- CI budgets;
- contract tests;
- runtime monitoring.
Written policy that cannot be observed often becomes forgotten policy.
Fitness functions should protect important properties
Do not create checks for every preference.
Protect properties whose regression would cause meaningful harm:
- security;
- accessibility;
- performance;
- deployability;
- dependency direction;
- public contract compatibility.
Architecture Decision Records
An ADR records:
- context;
- decision;
- alternatives;
- consequences;
- validation;
- status;
- revisit trigger.
It preserves reasoning after the original meeting is forgotten.
ADRs are not meeting minutes
An ADR should not capture every discussion detail.
It should explain:
Example ADR: URL state for catalogue filters
flowchart LR
URL["URL Query Params<br/>(?category=permits&page=2)"] --> Router["Router State Hook"]
Router --> Search["Search Input Component"]
Search --> API["Fetch API Client"]
API --> Results["Display Results Table"]Worked decision: Civic Platform Architecture
Context & Forces (Erbil Citizen Portal):
- 4 autonomous product squads (Health, Transport, Commerce, Education).
- Target users: 70% mobile browsers on congested 3G/4G networks; WCAG 2.1 AA legal mandate.
- High SEO requirement for public municipal announcements and legal circulars.
flowchart TD
Req["Citizen Portal Forces"] --> F1["Autonomy for 4 Squads"]
Req --> F2["Fast Mobile LCP < 2.0s over 3G"]
Req --> F3["Zero Accessible Regressions"]
Req --> F4["SEO for Legal Circulars"]Candidate architectures: Rejected options
| Candidate Option | Architecture Model | Reason for Rejection |
|---|---|---|
| Option A: Micro-Frontends | Webpack Module Federation; multi-repo independent deploys | 4.2s mobile LCP on 3G: duplicate React vendor runtimes violate performance budget |
| Option B: Client-Side SPA | Pure CSR single-page app; static CDN hosting | SEO & blank screen: fails public legal circular indexing and initial 3G render |
Candidate architectures: Accepted decision
Option C: Modular Monolith with Edge SSR (ACCEPTED)
| Architectural Dimension | Strategy & Evaluation |
|---|---|
| Mobile LCP | 1.4s (p75): cached semantic HTML at CDN edge |
| Search Engine Indexing | 100% crawlable: complete server-rendered document |
| Squad Autonomy | pnpm monorepo with strict package.json "exports" |
| Operational Overhead | Single container pipeline; zero distributed federation complexity |
Worked decision: Evidence that would reverse it
The decision to adopt a Modular Monolith with Edge SSR will be formally revisited if:
flowchart TD
Rev["Reversal Triggers (ADR-018)"]
Rev --> T1["Team Scale: Engineering squads grow from 4 to >12 squads<br/>and CI queue times exceed 30 minutes"]
Rev --> T2["Regulatory Mandate: Ministry of Education mandates<br/>hosting in an independent sovereign data center"]
Rev --> T3["Traffic Spikes: SSR compute costs exceed budget<br/>requiring static HTML export for catalog routes"]The ADR makes ownership and trade-offs explicit.
ADR status
stateDiagram-v2
[*] --> Proposed: Drafted by Engineer
Proposed --> Accepted: Team Architectural Consensus
Proposed --> Rejected: Fails Constraints / Trade-offs
Accepted --> Superseded: Replaced by Newer ADR
Accepted --> Deprecated: Capability Retired
Superseded --> [*]
Deprecated --> [*]
Rejected --> [*]Status tells readers whether the decision is active and whether another document replaces it.
Decision scope
Record what the decision does and does not cover.
Clear scope prevents a local decision from becoming accidental universal policy.
Decision matrices
A matrix can expose trade-offs:
| Option | Freshness | Complexity | Reversibility | Team fit |
|---|---|---|---|---|
| A | high | medium | high | high |
| B | medium | high | low | medium |
Use it to structure reasoning, not to pretend qualitative judgment is exact arithmetic.
Avoid weighted-score theater
Numbers can create false certainty when:
- criteria are subjective;
- weights are arbitrary;
- unknowns are hidden;
- important risks are averaged away.
Show the assumptions and discuss the decisive trade-offs directly.
Proof of concept and spike
flowchart LR
subgraph Spike["Technical Spike"]
Q1["Question: Does Rolldown bundle this in <500ms?"] --> Code1["Write 30-line test script"]
Code1 --> Ans1["Answer: Yes (420ms). Spike discarded."]
end
subgraph PoC["Proof of Concept"]
Q2["Question: Can Edge SSR integrate with legacy auth?"] --> Code2["Build skeleton prototype with 2 routes"]
Code2 --> Ans2["Validate feasibility across systems"]
endKeep the purpose narrow and record what was learned.
Do not mistake experimental code for a production architecture.
Architecture review
Review asks:
- does the decision fit the requirements?
- are quality attributes explicit?
- are failure modes understood?
- is the boundary appropriate?
- can it be operated and changed?
- what evidence remains weak?
It should improve the decision, not only approve a document.
Decision confidence
State confidence and uncertainty:
Low confidence should trigger a spike, guardrail, or revisit date.
Assumption register
Track assumptions such as:
- expected traffic;
- team capacity;
- API stability;
- browser support;
- deployment independence;
- data sensitivity;
- user behavior.
An assumption that changes can invalidate the decision.
Architecture risk and risk reduction
Reduce uncertainty early when it could make the chosen architecture expensive to reverse.
Failure mode analysis
For a feature, ask:
- what can fail?
- how will the user see it?
- what remains usable?
- can the operation retry?
- can the release roll back?
- who is alerted?
Failure behavior is part of the architecture, not a later polish step.
Graceful degradation
Examples:
- cached content when the network fails;
- static navigation when enhancement fails;
- panel fallback when one remote fails;
- queued draft when submission is offline.
Critical path
Identify the work required for the user’s first valuable outcome.
Keep optional features, telemetry, personalization, and secondary data off the critical path when possible.
Critical-path design affects rendering, performance, reliability, and architecture together.
Availability architecture
Availability includes:
- resilient requests;
- cache and fallback;
- independent failure boundaries;
- deployment and rollback;
- monitoring;
- recovery ownership.
“The server is up” does not prove that the user can complete the task.
Security architecture
Define:
- trust boundaries;
- identity and authorization;
- secret location;
- content and request validation;
- browser policies;
- third-party trust;
- logging and privacy.
Security belongs in the decision, not in a final header checklist.
Performance architecture
Make choices about:
- rendering topology;
- JavaScript responsibility;
- cacheability;
- code splitting;
- data waterfalls;
- device cost;
- route and journey budgets.
Performance constraints should influence boundaries before slow code accumulates.
Accessibility architecture
Accessibility is affected by:
- semantic HTML;
- component contracts;
- focus ownership;
- routing and navigation;
- progressive enhancement;
- design-system governance;
- testing and release policy.
Treat it as a system property, not only a component task.
Internationalization architecture
Plan for:
- locale and direction;
- text expansion;
- date and number formatting;
- routing and content negotiation;
- font coverage;
- translation loading;
- persistence and URL state.
Adding i18n late often reveals hidden assumptions across every layer.
Testability as an architecture attribute
Testability improves when the system has:
- explicit inputs and outputs;
- replaceable boundaries;
- deterministic transitions;
- isolated ownership;
- observable failures;
- stable contracts.
If architecture makes the important behavior impossible to isolate, testing cost is design feedback.
Observability as an architecture attribute
Make it possible to answer:
- which release failed?
- which route and user journey?
- which dependency or boundary?
- which users are affected?
- can the issue be mitigated or rolled back?
Observability should be designed into boundaries and identifiers.
Deployability as an architecture attribute
A deployable system has:
- build reproducibility;
- artifact identity;
- safe configuration;
- progressive delivery;
- compatibility strategy;
- rollback;
- operational ownership.
An architecture that cannot be released safely limits product change.
Maintainability, reliability, scalability
These properties interact and should be evaluated together.
Different kinds of scale
A system may scale users while failing teams, or scale features while becoming impossible to reason about.
Identify which scale is actually driving the decision.
Cognitive load
Architecture should reduce the amount a developer must understand at once.
Use:
- local reasoning;
- clear names;
- stable boundaries;
- focused modules;
- useful documentation;
- predictable workflows.
The fastest system to change is often the one with the clearest mental model.
Architecture documentation
Useful documentation includes:
- context and system diagrams;
- ADRs;
- package and dependency maps;
- ownership;
- runbooks;
- contracts;
- failure and recovery behavior.
Document decisions that future maintainers need to preserve or revisit.
C4-style thinking
Move through levels of architectural abstraction:
flowchart TD
L1["1. System Context<br/>Citizens, Portal, External Auth, Ministry APIs"] --> L2["2. Containers<br/>Web App SPA, Edge SSR Gateway, Redis Cache"]
L2 --> L3["3. Components<br/>Permit Catalogue, Routing Shell, Auth Provider"]
L3 --> L4["4. Code<br/>TypeScript Classes, Functions, React/Vue Components"]Use the level that answers the current question.
Do not create diagrams so detailed that nobody can use them.
Architecture diagram purpose
A diagram should help someone:
- understand ownership;
- find a trust boundary;
- trace data flow;
- identify failure containment;
- evaluate a change;
- operate the system.
If it only decorates a presentation, it is not architecture documentation.
Avoid architecture astronautics
Warning signs:
- abstractions with no current consumer;
- patterns chosen for prestige;
- distributed systems without distributed requirements;
- infrastructure whose operation exceeds its value;
- diagrams detached from code and ownership.
Complexity is not evidence of maturity.
Simplicity is a feature
Simple architecture can provide:
- faster onboarding;
- easier debugging;
- fewer failure modes;
- lower operating cost;
- more reversible change.
Simple does not mean primitive or careless.
Standards before custom infrastructure
Prefer platform and established protocols when they meet the requirement:
- HTTP;
- URL and history;
- HTML forms;
- browser storage;
- standard security headers;
- common observability formats.
Custom infrastructure should solve a demonstrated gap.
Libraries should amplify the platform
A library is valuable when it:
- reduces repeated risk;
- provides missing capability;
- preserves accessibility;
- improves productivity;
- remains replaceable enough.
It is risky when it hides basic browser behavior the team needs to understand.
Architecture and team skill
Choose a design the team can:
- explain;
- debug;
- test;
- deploy;
- operate;
- migrate.
An architecture that requires one specialist for every incident has a bus-factor problem.
Hiring and onboarding
Architecture affects:
- how quickly new engineers understand the system;
- how safely they make changes;
- how much tribal knowledge is required;
- how ownership is transferred.
Good boundaries and documentation are productivity tools.
Bus factor and ownership
Reduce dependence on one person through:
- shared runbooks;
- pair or review practice;
- explicit package ownership;
- reproducible workflows;
- decision records;
- operational rehearsal.
Ownership should create accountability, not private territory.
Guardrails versus gates
Use gates for security, compatibility, or release-critical requirements.
Use guardrails for recommended patterns where local judgment remains useful.
Paved roads
A paved road provides:
- safe defaults;
- templates;
- supported tooling;
- examples;
- deployment and observability;
- extension points.
It should reduce cognitive load without pretending every product has the same needs.
Architecture debt
Architecture debt is accumulated cost from decisions that no longer fit:
- obsolete boundaries;
- excessive coupling;
- stale platform assumptions;
- missing migrations;
- unsupported dependencies;
- unclear ownership.
Track it and prioritize by change cost and user impact.
Refactoring architecture
Refactor toward boundaries through small steps:
Preserve behavior while changing structure.
Strangler strategy revisited
Replace one capability at a time while the old system continues to serve the rest.
The migration needs:
- routing or proxy boundary;
- data compatibility;
- shared authentication;
- observability;
- rollback;
- ownership during coexistence.
Big-bang rewrite risk
A rewrite can:
- delay user value;
- lose hidden behavior;
- recreate old mistakes;
- require long coexistence anyway;
- concentrate technical and organizational risk.
Rewrite only when the expected value and migration control justify the discontinuity.
When a rewrite is justified
Possible evidence:
- the current platform cannot meet a required constraint;
- incremental migration is more expensive than replacement;
- ownership and product scope are stable;
- behavior is well understood;
- rollout and rollback are credible;
- the organization can sustain the work.
“The code feels old” is not enough evidence by itself.
Technical-debt prioritization
Prioritize by:
- frequency of change;
- blast radius;
- user harm;
- security or compliance;
- delivery friction;
- probability of failure;
- reversibility.
Debt that is stable and harmless may be lower priority than a small coupling point that blocks every release.
Cost of change
Measure:
- how many files or packages change;
- how many teams coordinate;
- how many environments verify;
- how many tests break;
- how much release work is needed;
- how hard rollback is.
Change amplification is architecture evidence.
Scenario: public documentation
Likely priorities:
- cacheability;
- discoverability;
- accessibility;
- content publishing;
- low client cost.
Static or revalidated output with focused enhancement may be a strong fit.
Scenario: internal admin
Likely priorities:
- authentication;
- dense interaction;
- productivity;
- resilient forms;
- domain workflows;
- observability.
A modular client application or hybrid route may be simpler than a public-content topology.
Scenario: e-commerce
Different routes can need different strategies:
One application does not require one rendering strategy.
Scenario: collaborative editor
Priorities may include:
- real-time updates;
- conflict handling;
- offline drafts;
- presence;
- low interaction latency;
- strong domain state.
The architecture should begin with synchronization and failure requirements, not a component library choice.
Scenario: government service
Likely priorities:
- accessibility;
- reliability;
- security;
- auditability;
- long support life;
- progressive enhancement;
- clear recovery.
Operational simplicity and standards may matter more than fashionable distribution.
Scenario: large enterprise platform
Possible pressures include:
- team autonomy;
- shared foundations;
- independent release cadence;
- multiple domains;
- legacy migration;
- governance and security.
Use explicit contracts and incremental separation before reaching for runtime federation.
No architecture is context-free
A decision without context is a slogan.
Record:
- product;
- users;
- team;
- constraints;
- quality priorities;
- deployment model;
- expected change;
- evidence and unknowns.
Resume-driven architecture
This occurs when a technology is chosen mainly because it is interesting or marketable.
Counter it with:
- explicit requirements;
- measurable constraints;
- a representative spike;
- total lifecycle cost;
- a revisit trigger.
Technology can be valuable without being the reason for the decision.
Cargo-cult architecture
Copying a pattern from another company without its context can import:
- unnecessary services;
- incompatible team assumptions;
- different scale costs;
- hidden operational requirements.
Learn the principle, then re-evaluate it against your constraints.
Framework as architecture
The framework participates in architecture but does not replace system decisions.
Pattern collection
Using every pattern creates:
- duplicated concepts;
- conflicting state paths;
- more documentation;
- harder onboarding;
- ambiguous ownership.
Patterns should solve a problem that exists in the system.
Hidden architecture
Architecture is hidden when behavior depends on:
- undocumented globals;
- convention nobody can find;
- private package imports;
- manual deployment steps;
- tribal knowledge;
- unrecorded exceptions.
Make important decisions visible in code, checks, docs, and ownership.
Architecture freeze
Architecture should be stable enough to guide work but revisable when evidence changes.
Use:
- ADR status;
- review triggers;
- fitness functions;
- migration paths;
- measured outcomes.
“We decided once” is not a reason to ignore new constraints.
Architecture review cadence
Review architecture when:
- product scope changes;
- team topology changes;
- a quality attribute regresses;
- an integration becomes a bottleneck;
- a dependency reaches end of life;
- a migration creates new boundaries.
Review should be triggered by change, not only by calendar ceremony.
Measure architecture outcomes
Possible outcomes include:
- change lead time;
- deployment failure and recovery;
- accessibility defects;
- performance budgets;
- dependency cycle count;
- onboarding time;
- support burden;
- team autonomy in practice.
Architecture quality is visible through system behavior.
Architecture is sociotechnical
Technology and organization interact through:
- ownership;
- incentives;
- communication;
- release authority;
- expertise;
- support expectations.
A design that ignores people will be changed by people in ways the diagram did not predict.
Architecture and product lifecycle
Prototype, growth, maturity, and retirement may need different priorities.
Do not optimize a prototype for the operational needs of a global platform without evidence.
Architecture and deadlines
Deadlines do not eliminate architecture decisions.
Make the trade-off explicit:
- what is deferred;
- what risk is accepted;
- what boundary remains safe;
- what follow-up is required;
- what cannot be compromised.
Shortcuts are safer when named and bounded.
Architecture and future change
Do not attempt to predict every future feature.
Instead, preserve:
- clear ownership;
- replaceable boundaries;
- stable contracts;
- observable assumptions;
- reversible choices where possible.
Good architecture makes learning and change affordable.
The decision framework
This turns architecture into a repeatable practice.
Full ADR template
Keep the document concise enough to read and specific enough to preserve reasoning.
Revisit triggers
Examples:
- traffic exceeds the assumed range;
- team ownership changes;
- deployment becomes a bottleneck;
- performance budget fails;
- a dependency is retired;
- a security requirement changes;
- evidence from a spike contradicts an assumption.
Triggers turn an ADR into a living decision rather than a permanent verdict.
Example: catalogue state decision
The decision improves reload, sharing, caching, and testability without requiring a global store.
Example: rendering decision
The architecture is route-specific because requirements differ.
Example: design-system decision
Share:
- tokens;
- accessible primitives;
- stable interaction patterns.
Keep local:
- domain workflows;
- product-specific orchestration;
- experimental layouts.
Promote concepts only when stability and ownership justify it.
Example: micro-frontend decision
Choose runtime separation only when:
- independent deployment is a real bottleneck;
- capability boundaries are stable;
- contracts and observability exist;
- failure containment matters;
- the organization can operate the integration.
Otherwise start with packages or a modular monolith.
An architectural “no”
Saying no can protect the system:
An architectural no should include the reason and the condition that could change it.
An architectural “yes”
A good yes is bounded:
Boundaries make adoption safer.
The browser is still the foundation
Even sophisticated systems eventually interact with:
- HTML;
- CSS;
- JavaScript;
- URLs;
- HTTP;
- browser security;
- input, focus, and rendering.
Architecture should amplify platform capabilities rather than hide them completely.
HTML, CSS, and JavaScript are architecture
Choices at the platform layer shape accessibility, performance, testing, and resilience.
Testing and observability are architecture feedback
Tests reveal whether boundaries are behaviorally useful.
Observability reveals whether boundaries work under real users, networks, devices, and releases.
Use both to decide whether architecture is serving its purpose.
Architecture is continuous
The goal is not a final perfect diagram.
The goal is a system that can change without losing control.
Practical lab: Make and Defend an Architecture Decision
Create an evidence-backed architecture decision for a product platform without ranking frameworks or adopting complexity by default.
The practical ends with an ADR, fitness functions, failure modes, migration path, and decision report.
Practical stages 1–5: define the problem
- Define the product.
- List functional requirements.
- Rank quality attributes.
- Record constraints.
- Define state ownership.
State the user, team, deployment, security, performance, and organizational context.
Practical stages 6–12: design the boundaries
- Choose rendering topology.
- Choose component boundaries.
- Decide on shared state.
- Choose API boundary strategy.
- Define security boundaries.
- Define performance constraints.
- Define accessibility requirements.
Treat each as a decision with trade-offs, not a framework checkbox.
Practical stages 13–19: evaluate the platform
- Evaluate dependencies.
- Evaluate framework fit.
- Decide repository structure.
- Evaluate micro-frontends.
- Define testing layers.
- Define delivery.
- Define observability.
Prefer the smallest coherent solution that satisfies the constraints.
Practical stages 20–24: reduce uncertainty
- Create two alternatives.
- Run a technical spike.
- Write the ADR.
- Add fitness functions.
- Define failure modes.
The spike should answer the highest-risk unknown, not build the entire future system.
Practical stages 25–27: migration and defense
Practical stages 1–3: constraints, options & empirical spike
- Stage 1 (Explicit Problem & Constraint Mapping):
- Document functional goals and prioritize non-negotiable quality attributes.
- Stage 2 (Formulating Three Viable Candidate Architectures):
- Candidate A (Micro-Frontends), Candidate B (Client-Side SPA), Candidate C (Modular Monolith with Edge SSR).
- Stage 3 (The Investigative Technical Spike):
- Run a benchmark spike measuring bundle size and p75 LCP under 4x CPU throttling.
Practical stages 4–5: ADR & reversal plan
- Stage 4 (Drafting the Formal ADR):
- Record Title, Status, Context, Decision, Consequences, and Automated Fitness Functions.
- Stage 5 (Reversal Plan & Review Trigger Conditions):
- Define exact quantitative thresholds (team size >12 squads, CI queue >30m) that trigger architectural review.
Verification: Decisions reflect empirical evidence and trade-offs rather than technology fashion.
Try this yourself
Choose one architecture decision and write:
If the decision cannot name a problem, it may be technology fashion rather than architecture.
Troubleshooting guide (Part 1)
| Symptom | Likely cause |
|---|---|
| Architecture debate never ends | Requirements and decision criteria are unclear |
| Every solution is global | Ownership and locality were not evaluated |
| Micro-frontends are proposed immediately | Organizational pressure was mistaken for runtime need |
| ADRs are long and unread | They record meetings instead of decisions |
| Fitness checks are ignored | They protect preferences rather than important properties |
Troubleshooting guide (Part 2)
| Symptom | Likely cause |
|---|---|
| Rewrite feels safer than migration | Hidden behavior and compatibility cost are underestimated |
| Shared package changes break everyone | Public API and versioning are weak |
| Architecture depends on one expert | Ownership, documentation, and runbooks are insufficient |
| “Simple” system fails at scale | The relevant quality attribute was not measured |
Completion checklist
- the problem and context are explicit;
- quality attributes are prioritized and measurable;
- constraints are distinguished from preferences;
- alternatives and trade-offs are documented;
- unknowns are tested with focused spikes;
- boundaries reduce change cost and blast radius;
- important architecture properties have guardrails;
- failure, security, performance, and accessibility are included;
- migration and rollback are credible;
- the decision has an owner and revisit trigger.
Misconceptions to leave behind (Part 1)
| Misconception | Better mental model |
|---|---|
| Architecture is the technology stack | It is contextual decisions and boundaries |
| A good architecture optimizes everything | It makes explicit trade-offs |
| More abstraction is better | Abstraction has coupling and cognitive cost |
| Duplication is always bad | Local duplication can preserve autonomy |
| Global state is required for scale | Ownership and lifetime determine scope |
| SSR is more architectural than CSR | Rendering is one contextual decision |
| Micro-frontends are the natural future | Distribution is justified by real independence needs |
Misconceptions to leave behind (Part 2)
| Misconception | Better mental model |
|---|---|
| Every shared component belongs centrally | Stable shared concepts deserve promotion |
| ADRs are bureaucracy | They preserve reasoning and revisit conditions |
| A decision cannot change | Architecture should learn from evidence |
| A rewrite is cleaner | Incremental migration often reduces risk |
| Technical debt must always be removed | Prioritize by impact and change cost |
| Simplicity means underengineering | Simplicity can be deliberate, bounded architecture |
The chapter in one sentence
Make architecture decisions from requirements, quality attributes, constraints, evidence, and reversible boundaries - and keep revisiting them as the system and organization learn.
Course completion: Capstone architecture
Congratulations on completing all 18 chapters of modern web application engineering!
Next steps to master front-end architecture:
- Capstone Architecture Project: Build and defend an end-to-end civic application platform;
- Appendix A: Review the Architectural Rosetta Stone (React vs. Vue mechanics);
- Appendix B: Consult the Modern Browser APIs Reference;
- Appendix C: Run the Production Deployment Checklist before every release.
Questions
Which architecture decision in your system is treated as permanent even though its assumptions, evidence, or constraints have already changed?