Fragments at Build Time
Learning Website with Next.js
Chapter 4 ยท Reading Fragments at Build Time
Chapter 3 described every file. This chapter shows one: it turns a content file into the HTML a visitor receives, makes every page of the languages site when the site is built, and makes sure the scripts inside those pages still run. Two of the findings came from opening a real browser, not from reading the code.
next dev, other
browsers, and pages with scripts that load other scripts.
Step 1: From a File to a Fragment
A content file is not yet what a site shows. Three steps, the same ones the Django importer takes:
- Unwrap. A complete HTML document gives its body, preceded by the styles from its head; a fragment loses the banner comment at its top.
- Fix solution links. Every
.txtlink is pointed at the course's ownsolutionsfolder, whatever the old link said. - Site changes. A site may add one of its own. The languages site marks each line in the language being
taught with a
langattribute (lang="hu",lang="ja"), so screen readers pronounce it correctly and the browser draws Japanese with the right glyph shapes.
A body that never closes is refused with a MalformedDocument error rather than guessed at. Language marking adds
lang to whole class names only (hu, not hu-lesson), leaves an element that already has
one alone, and is safe to run twice. Percent-decoding of file names never throws: a cut-off sequence becomes a replacement
character. The loader also gained an option to read one site only, decided from the path alone, so the languages
site reads 561 files, not 4,414.
Is It the Same as Django's?
The Django project stored the prepared fragment of every page. A script there writes a SHA-256 of each; another here prepares every page and compares:
| Check | Result |
|---|---|
| Pages prepared | 4,415 in 10.6 s (one is this chapter's own new file, not yet in the Django database) |
| Identical to Django's stored fragment | 4,414 of 4,414; different: 0; failed to prepare: 0 |
| Language marking switched off on purpose | 4,253 identical, 161 different: the check works |
| File names in solution links not decoded, or decoded with the wrong text encoding | 0 different |
.txt file name, so that decoding (and its UTF-8 handling) is checked only
by its unit tests, not by the real content. The comparison proves what the real pages exercise, and only that.
Step 2: Make Every Page When the Site Is Built
Next.js can make a page at build time from a list of addresses. Here the list is the content folder: one address for each page of the site, so a visitor receives a finished HTML file and no code runs when they ask for it.
LW_CONTENT_ROOThas no default: a build that quietly reads the wrong folder is worse than one that stops.- A content file that cannot be read stops the build, naming the file. (None did.)
- The page title and description come from the banner. A page with its own
<h1>shows it; every other page gets one made from its title. dangerouslySetInnerHTMLis React's name for “put this text in as HTML”. It is safe here only because the text is one of the site's own files, read at build time. It must never be given text a visitor typed.
| Real build and fetch | Result |
|---|---|
| Build the languages site | 567 pages made (the 561 pages plus Next's own and the test bench), 4.1 s of page generation, 14.6 s in all, 59 MB |
| A Hungarian lesson | 200, 97 KB, heading “Buying Clothes: Sizes, Fit & Returns”, 38 lang="hu" marks |
| A German lesson, a hiragana lesson | 200, 65 KB and 51 KB |
| A page with a Japanese address (hiragana ใ) | 200: made at build time, found by its percent-encoded address |
| An address that does not exist | 404 |
| The test bench address, in a normal build | 404 |
Step 3: Scripts Inside a Fragment
1,844 of the 4,415 pages have a <script> in their fragment (copy buttons and interactive tools; on the
languages site only 2). They have to work, and there is a trap. A browser runs the scripts in the page it first
downloads. But when a visitor clicks a link, Next fetches the new page's data and inserts it into the
page, and scripts inserted that way never run. The cure is to replace each script with a fresh copy of itself, which the
browser does run, and a tiny client component does that.
To see it, I made a test bench (it exists only in a build made with LW_LAB=1): two pages with the same counting
script, one with the component and one without, and an index that links to both. A real headless Chrome was driven through the
DevTools protocol, using only Node's own WebSocket:
| How the visitor arrives | First version | Fixed version |
|---|---|---|
| Open the page without the component, directly | script ran 1 time | script ran 1 time |
| Open the page with the component, directly | script ran 2 times (a bug) | script ran 1 time |
| Click a link to the page without the component | waiting: the script never ran (the trap) | waiting (the same: no component, no run) |
| Click a link to the page with the component | script ran 1 time | script ran 1 time |
| Click, go back, click again | not tested | script ran 2 times in the window: once per visit |
next dev (the markers are written for its double effects but it was not run), scripts that load other scripts with
src= (the copy keeps the attributes; none was tried), and browsers other than Chrome. There are no styles,
menus or navigation yet (Chapters 5 and 6), and no PDFs (Chapter 8).
Hands-On Exercises
Write the steps that turn a content file into a fragment (unwrap, fix solution links, a site change), test the awkward inputs, and prove the result is identical to Django's on all 4,414 real pages. Show the comparison can fail, and say which parts the real data does not exercise.
๐ View solutionMake every page of the languages site at build time from the content folder, with no default content path, a build that stops on an unreadable file, and 404 for everything else. Build it, serve it and fetch pages including one with a non-ASCII address.
๐ View solutionMake scripts inside fragments run after a client-side navigation, but not twice on a direct load. Build a test bench and drive a real headless Chrome through it; record the failure of your first version and the fix.
๐ View solutionChapter 4 Quick Reference
prepareFragment: unwrap a document or drop the banner, point.txtlinks at the course'ssolutionsfolder, apply the site's own change- Languages site change: add
langto whole class nameshu,de,fr,jp(asja); safe to run twice loadPages(root, { site })decides the site from the path alone, so other sites' files are never read- Identical to Django on 4,414 of 4,414 pages; the check fails (161 different) when language marking is off; file-name decoding is only unit-tested
generateStaticParams+dynamicParams = false: one page per content file, everything else 404LW_CONTENT_ROOThas no default; an unreadable content file stops the builddangerouslySetInnerHTMLonly for the site's own files, never for visitor text- Real build: 567 pages, 14.6 s, 59 MB; non-ASCII addresses work
- Scripts: they run on the page first downloaded, not after a link click;
RunScriptsre-creates them, but only for navigated pages - Found in Chrome: the first version ran a script twice on a direct load; tested both ways of arriving; not tested in
next dev