learning-website-nextjs1-3 Exercise 3: Is It the Same as the Django Importer? Run Both on the Real Content =========================================================================================================== Two implementations of the same rules should give the same answers, and the only honest way to say so is to run both on the real content and compare every page. Step 1: dump what the Django project stored (page and course fields, no page bodies) to JSON: python dumpdb.py django-dump.json (in the Django project folder, after import_content) 4414 pages, 396 courses Step 2: parse the same folder with @lw/content and compare field by field: Save as compare-django.mjs: // Parse the real content folder with @lw/content and compare every page with what the Django project stored. // node compare-django.mjs import { readFileSync } from "node:fs"; import { groupCourses, loadPages } from "./packages/content/src/index.ts"; const [root, dumpPath] = process.argv.slice(2); const started = performance.now(); const { pages, errors } = loadPages(root); const ms = Math.round(performance.now() - started); console.log(`TypeScript: ${pages.length} pages, ${errors.length} errors, ${ms} ms`); for (const e of errors.slice(0, 8)) console.log(" error:", e.path, "|", e.message.slice(0, 90)); const django = JSON.parse(readFileSync(dumpPath, "utf8")); const byPath = new Map(pages.map((p) => [p.path, p])); console.log(`Django: ${django.pages.length} pages`); const onlyTs = pages.filter((p) => !django.pages.some((d) => d.path === p.path)).length; const djangoSet = new Set(django.pages.map((d) => d.path)); const missingInTs = django.pages.filter((d) => !byPath.has(d.path)).map((d) => d.path); console.log(`in Django but not parsed here: ${missingInTs.length}`, missingInTs.slice(0, 5)); console.log(`parsed here but not in Django: ${pages.filter((p) => !djangoSet.has(p.path)).length}`); const fields = ["site", "kind", "title", "chapter_no", "created", "updated", "course_no", "course_name", "summary"]; const diffs = Object.fromEntries(fields.map((f) => [f, []])); for (const d of django.pages) { const p = byPath.get(d.path); if (!p) continue; const mine = { site: p.site, kind: p.kind, title: p.title, chapter_no: p.chapterNo, created: p.created, updated: p.updated, course_no: p.courseNo, course_name: p.courseName, summary: p.summary, }; for (const f of fields) { if (f === "summary" && p.summary === "") continue; // Django fills an empty summary from the page text instead if (f === "course_no" || f === "course_name") { if (d.course_folder === null) continue; // only chapters have a course } const theirs = d[f] ?? null; if ((mine[f] ?? null) !== theirs) diffs[f].push({ path: d.path, mine: mine[f] ?? null, theirs }); } } let total = 0; for (const f of fields) { total += diffs[f].length; console.log(`${f.padEnd(12)} differences: ${diffs[f].length}`); for (const x of diffs[f].slice(0, 3)) console.log(" ", JSON.stringify(x).slice(0, 200)); } const courses = groupCourses(pages); console.log(`courses (by folder): TypeScript ${courses.length}, Django ${django.courses.length}`); const dj = new Map(django.courses.map((c) => [c.folder, c])); let cd = 0; for (const c of courses) { const d = dj.get(c.folder); if (!d || d.name !== c.name || d.site !== c.site || (d.course_no ?? null) !== (c.courseNo ?? null)) { cd++; if (cd <= 3) console.log(" course differs:", c.folder, c.name, "|", d && d.name, c.courseNo, d && d.course_no); } } console.log(`course differences: ${cd}`); console.log(`TOTAL page field differences: ${total}`); node compare-django.mjs django-dump.json TypeScript: 4414 pages, 0 errors, 1849 ms Django: 4414 pages in Django but not parsed here: 0 [] parsed here but not in Django: 0 site differences: 0 kind differences: 0 title differences: 0 chapter_no differences: 0 created differences: 0 updated differences: 0 course_no differences: 0 course_name differences: 0 summary differences: 0 courses (by folder): TypeScript 396, Django 396 course differences: 0 TOTAL page field differences: 0 So on the real content (3,907 course chapters, 410 complete pages, 47 sidebar pages, 43 lessons, 7 others) the TypeScript reader finds the same 4,414 pages, with no errors, and agrees with Django on site, kind, title, chapter number, both dates, course number, course name and summary for every page, and on all 396 courses. It took 1.5 to 1.9 seconds for the whole folder (two runs) (this reads and parses only; Django's import also prepares and indexes every page, so the two times are NOT a comparison). Can this check fail? Zero differences on the first run is only worth something if the comparison is able to find one. Two deliberate breakages of the TypeScript code, then back again: humanize made to upper-case the rest of each word title differences: 0 (!) named entities no longer decoded title differences: 11 both restored TOTAL page field differences: 0 The second shows the comparison does catch a fault, and that the small entity table covers every real case. The FIRST is the useful warning: no real page gets its title from its file name, so that function is checked only by its unit tests, not by the real content. (A check you have not seen fail has not been shown to check.) What the comparison does NOT cover: the page bodies (preparing and showing the HTML is Chapter 4), the links, and files that Django rejects while preparing the body (there were none, since both sides have the same 4,414). WHY THIS WORKS AS AN ANSWER --------------------------- It compares everything on the real data, reports its own limits, and was shown to be able to fail.