Fragments at Build Time

Learning Website with Next.js

Chapter 4 ยท Reading Fragments at Build Time

Chapter 3 described every file. This chapter shows one: it turns a content file into the HTML a visitor receives, makes every page of the languages site when the site is built, and makes sure the scripts inside those pages still run. Two of the findings came from opening a real browser, not from reading the code.

Run for real, on 561 pages of the languages site
34 tests pass; the prepared HTML of all 4,414 pages was compared byte for byte with what the Django project stored; the site was built, served and fetched; and a headless Chrome was driven to test scripts. Not tried: next dev, other browsers, and pages with scripts that load other scripts.

Step 1: From a File to a Fragment

A content file is not yet what a site shows. Three steps, the same ones the Django importer takes:

  1. Unwrap. A complete HTML document gives its body, preceded by the styles from its head; a fragment loses the banner comment at its top.
  2. Fix solution links. Every .txt link is pointed at the course's own solutions folder, whatever the old link said.
  3. Site changes. A site may add one of its own. The languages site marks each line in the language being taught with a lang attribute (lang="hu", lang="ja"), so screen readers pronounce it correctly and the browser draws Japanese with the right glyph shapes.
export function prepareFragment(raw: string, path: string, site: SiteName): string { const courseUrlPath = path.includes("/") ? "/" + posix.dirname(path) : ""; let fragment = rewriteSolutionLinks(extractFragment(raw), courseUrlPath); for (const transform of SITE_TRANSFORMS[site] ?? []) fragment = transform(fragment); return fragment; }

A body that never closes is refused with a MalformedDocument error rather than guessed at. Language marking adds lang to whole class names only (hu, not hu-lesson), leaves an element that already has one alone, and is safe to run twice. Percent-decoding of file names never throws: a cut-off sequence becomes a replacement character. The loader also gained an option to read one site only, decided from the path alone, so the languages site reads 561 files, not 4,414.

Is It the Same as Django's?

The Django project stored the prepared fragment of every page. A script there writes a SHA-256 of each; another here prepares every page and compares:

CheckResult
Pages prepared4,415 in 10.6 s (one is this chapter's own new file, not yet in the Django database)
Identical to Django's stored fragment4,414 of 4,414; different: 0; failed to prepare: 0
Language marking switched off on purpose4,253 identical, 161 different: the check works
File names in solution links not decoded, or decoded with the wrong text encoding0 different
What the last row means
No real page has a percent-encoded .txt file name, so that decoding (and its UTF-8 handling) is checked only by its unit tests, not by the real content. The comparison proves what the real pages exercise, and only that.

Step 2: Make Every Page When the Site Is Built

Next.js can make a page at build time from a list of addresses. Here the list is the content folder: one address for each page of the site, so a visitor receives a finished HTML file and no code runs when they ask for it.

// every page of the site is made when the site is built; any other address is a 404 export const dynamicParams = false; export function generateStaticParams() { return [...sitePages().values()].map((page) => ({ path: urlSegments(page) })); }
  • LW_CONTENT_ROOT has no default: a build that quietly reads the wrong folder is worse than one that stops.
  • A content file that cannot be read stops the build, naming the file. (None did.)
  • The page title and description come from the banner. A page with its own <h1> shows it; every other page gets one made from its title.
  • dangerouslySetInnerHTML is React's name for “put this text in as HTML”. It is safe here only because the text is one of the site's own files, read at build time. It must never be given text a visitor typed.
Real build and fetchResult
Build the languages site567 pages made (the 561 pages plus Next's own and the test bench), 4.1 s of page generation, 14.6 s in all, 59 MB
A Hungarian lesson200, 97 KB, heading “Buying Clothes: Sizes, Fit & Returns”, 38 lang="hu" marks
A German lesson, a hiragana lesson200, 65 KB and 51 KB
A page with a Japanese address (hiragana ใ‚)200: made at build time, found by its percent-encoded address
An address that does not exist404
The test bench address, in a normal build404

Step 3: Scripts Inside a Fragment

1,844 of the 4,415 pages have a <script> in their fragment (copy buttons and interactive tools; on the languages site only 2). They have to work, and there is a trap. A browser runs the scripts in the page it first downloads. But when a visitor clicks a link, Next fetches the new page's data and inserts it into the page, and scripts inserted that way never run. The cure is to replace each script with a fresh copy of itself, which the browser does run, and a tiny client component does that.

To see it, I made a test bench (it exists only in a build made with LW_LAB=1): two pages with the same counting script, one with the component and one without, and an index that links to both. A real headless Chrome was driven through the DevTools protocol, using only Node's own WebSocket:

How the visitor arrivesFirst versionFixed version
Open the page without the component, directlyscript ran 1 timescript ran 1 time
Open the page with the component, directlyscript ran 2 times (a bug)script ran 1 time
Click a link to the page without the componentwaiting: the script never ran (the trap)waiting (the same: no component, no run)
Click a link to the page with the componentscript ran 1 timescript ran 1 time
Click, go back, click againnot testedscript ran 2 times in the window: once per visit
What went wrong, and how to avoid it
The mistake: my first component re-created every script it found, including on the page that was downloaded directly, where the browser had already run them. The script ran twice; a script that adds a click handler would have added it twice. Reading the code would not have shown it; counting in a real browser did. The takeaway: re-run scripts only for pages that arrived by navigation, and test both ways of arriving with a script that counts how often it ran. The fix: a mark in the layout, set after the first page is hydrated. A child's effect runs before its parent's, so on the first page the component sees the mark missing (the page was downloaded: do nothing), and on every later page it sees it (the page came by navigation: re-run). Two small markers stop React's development mode, which runs every effect twice, from running a script twice.
What was not verified
next dev (the markers are written for its double effects but it was not run), scripts that load other scripts with src= (the copy keeps the attributes; none was tried), and browsers other than Chrome. There are no styles, menus or navigation yet (Chapters 5 and 6), and no PDFs (Chapter 8).

Hands-On Exercises

Exercise 1

Write the steps that turn a content file into a fragment (unwrap, fix solution links, a site change), test the awkward inputs, and prove the result is identical to Django's on all 4,414 real pages. Show the comparison can fail, and say which parts the real data does not exercise.

๐Ÿ“„ View solution
Exercise 2

Make every page of the languages site at build time from the content folder, with no default content path, a build that stops on an unreadable file, and 404 for everything else. Build it, serve it and fetch pages including one with a non-ASCII address.

๐Ÿ“„ View solution
Exercise 3

Make scripts inside fragments run after a client-side navigation, but not twice on a direct load. Build a test bench and drive a real headless Chrome through it; record the failure of your first version and the fix.

๐Ÿ“„ View solution

Chapter 4 Quick Reference

  • prepareFragment: unwrap a document or drop the banner, point .txt links at the course's solutions folder, apply the site's own change
  • Languages site change: add lang to whole class names hu, de, fr, jp (as ja); safe to run twice
  • loadPages(root, { site }) decides the site from the path alone, so other sites' files are never read
  • Identical to Django on 4,414 of 4,414 pages; the check fails (161 different) when language marking is off; file-name decoding is only unit-tested
  • generateStaticParams + dynamicParams = false: one page per content file, everything else 404
  • LW_CONTENT_ROOT has no default; an unreadable content file stops the build
  • dangerouslySetInnerHTML only for the site's own files, never for visitor text
  • Real build: 567 pages, 14.6 s, 59 MB; non-ASCII addresses work
  • Scripts: they run on the page first downloaded, not after a link click; RunScripts re-creates them, but only for navigated pages
  • Found in Chrome: the first version ran a script twice on a direct load; tested both ways of arriving; not tested in next dev