learning-website-nextjs1-4 Exercise 1: Prepare a Fragment, and Prove It Matches Django
===============================================================================================
A content file is not yet what a site shows. Three steps (the same ones the Django importer takes) turn it into a
fragment:
1. a complete HTML document gives its body, preceded by the styles from its head; a fragment loses the banner comment
at its top;
2. every .txt link is pointed at the course's own solutions folder, whatever the old link said;
3. a site may add one change of its own: the languages site marks each line in the language being taught with a
lang attribute (lang="hu", lang="ja", ...), so screen readers and browsers treat it correctly.
Save as packages/content/src/fragment.ts:
import { posix } from "node:path";
import type { SiteName } from "@lw/sites";
import { addLanguageAttributes } from "./languages.ts";
import { ContentParseError } from "./parse.ts";
/**
* Turn a raw content file into the fragment a site shows. The same steps as the Django importer:
* a complete document gives its body (plus the styles from its head), a fragment loses its banner comment,
* every .txt link is pointed at the course's own solutions folder, and a site may add its own change.
*/
const BODY_OPEN = /
]*>/i;
const BODY_CLOSE = /<\/body>/i;
const HEAD = /]*>([\s\S]*?)<\/head>/i;
const STYLE_ONE = /" +
"hi
";
assert.equal(extractFragment(raw), "\n\nhi
");
assert.equal(extractFragment("x
"), "x
");
});
test("a body that never closes is refused, not guessed at", () => {
assert.throws(() => extractFragment("x
"), MalformedDocument);
assert.throws(() => extractFragment("x"), MalformedDocument);
});
test("percent-decoding never throws and understands UTF-8", () => {
assert.equal(unquote("a%20b"), "a b");
assert.equal(unquote("%E3%81%82.txt"), "あ.txt");
assert.equal(unquote("100%"), "100%");
assert.equal(unquote("%E3%81"), "�"); // a cut-off sequence is replaced, not an error
});
test("every .txt link is pointed at the course's solutions folder", () => {
const html = 'a b c';
assert.equal(
rewriteSolutionLinks(html, "/web/js"),
'a b c');
assert.equal(rewriteSolutionLinks('', "web/js"), '');
});
test("language marking adds lang to whole class names only", () => {
assert.equal(addLanguageAttributes('x'), 'x');
assert.equal(addLanguageAttributes('x
'), 'x
');
assert.equal(addLanguageAttributes('x
'), 'x
');
assert.equal(addLanguageAttributes('x'), 'x');
assert.equal(addLanguageAttributes('x'), 'x'); // not a property lookup
assert.equal(addLanguageAttributes("
"), "
");
});
test("language marking is safe to run twice", () => {
const once = addLanguageAttributes('x');
assert.equal(addLanguageAttributes(once), once);
});
test("preparing a page applies the shared steps, then the site's own", () => {
const raw = '\nouie';
assert.equal(
prepareFragment(raw, "france/french-basic-1/f_1_1.html", "languages"),
'ouie');
assert.equal(
prepareFragment(raw, "web-development/x/p.html", "webdevelopment"),
'ouie'); // no marking on another site
assert.equal(prepareFragment("x
", "top.html", "languages"), "x
"); // a page with no folder
});
Run:
npm test
ℹ tests 34
ℹ pass 34
ℹ fail 0
(the 17 tests of Chapter 3, 9 of the site package, and 8 new ones)
Now the proof. The Django project stored the prepared fragment of every page; its dumphashes.py writes a SHA-256 of each.
compare-fragments.mjs prepares every page here and compares the hashes:
Save as compare-fragments.mjs:
// Prepare every page's fragment with @lw/content and compare its hash with the one Django stored.
// node compare-fragments.mjs
import { createHash } from "node:crypto";
import { readFileSync } from "node:fs";
import { loadPages, prepareFragment, readContentFile } from "./packages/content/src/index.ts";
const [root, hashesPath] = process.argv.slice(2);
const django = JSON.parse(readFileSync(hashesPath, "utf8"));
const started = performance.now();
const { pages } = loadPages(root);
let same = 0;
const different = [];
const missing = [];
const failed = [];
for (const page of pages) {
const expected = django[page.path];
if (expected === undefined) { missing.push(page.path); continue; }
let fragment;
try {
fragment = prepareFragment(readContentFile(root, page.path), page.path, page.site);
} catch (error) {
failed.push(`${page.path}: ${error.message}`);
continue;
}
const hash = createHash("sha256").update(fragment, "utf8").digest("hex");
if (hash === expected) same++;
else different.push(page.path);
}
const ms = Math.round(performance.now() - started);
console.log(`${pages.length} pages prepared in ${ms} ms`);
console.log(`identical to Django's stored fragment: ${same}`);
console.log(`different: ${different.length}`, different.slice(0, 5));
console.log(`not in Django: ${missing.length} failed to prepare: ${failed.length}`, failed.slice(0, 3));
console.log(`Django pages not parsed here: ${Object.keys(django).filter((p) => !pages.some((x) => x.path === p)).length}`);
python dumphashes.py django-hashes.json (in the Django project)
4414 fragment hashes
node compare-fragments.mjs django-hashes.json
4415 pages prepared in 9078 ms
identical to Django's stored fragment: 4414
different: 0 []
not in Django: 1 failed to prepare: 0 []
Django pages not parsed here: 0
(4,414 identical; the 4,415th page is this chapter's own new file, which the Django database has not imported yet.)
Can this comparison fail? Three deliberate breakages of the TypeScript code, each undone afterwards:
language marking switched off identical 4253, different 161
file names in solution links not percent-decoded different 0
decoded with the wrong text encoding different 0
The first shows the comparison works: 161 pages carry language marks. The other two show something honest: NO real page has a
percent-encoded .txt file name, so that decoding (and its UTF-8 handling) is checked only by its unit tests, not by the real
content.
WHY THIS WORKS AS AN ANSWER
---------------------------
"The same as Django" is shown byte for byte on 4,414 real pages, with a note of which parts of the code the real data does
not exercise.