learning-website-nextjs1-9 Exercise 2: Prove It Against the Django Search, and Two Findings ========================================================================================= The Django project searches with SQLite's full-text search. This one searches with its own index. They should find the same pages, and they can be compared on the real content, in two steps. STEP 1: is the TEXT the same? The Django index stores each page's plain text; dumptext.py (in the Django project) writes a fingerprint of each, and compare-text.mjs makes the same text here and compares every page byte for byte. Save as compare-text.mjs: // Is the plain text we index the same text the Django project indexed? Compare every page's text, byte for byte. // node compare-text.mjs import { createHash } from "node:crypto"; import { readFileSync } from "node:fs"; import { loadPages, prepareFragment, readContentFile } from "./packages/content/src/index.ts"; import { htmlToText } from "./packages/search/src/index.ts"; const [root, dumpPath] = process.argv.slice(2); const django = JSON.parse(readFileSync(dumpPath, "utf8")); const { pages } = loadPages(root); let same = 0; let missing = 0; const different = []; for (const page of pages) { const expected = django[page.path]; if (!expected) { missing++; continue; } const text = htmlToText(prepareFragment(readContentFile(root, page.path), page.path, page.site)); if (createHash("sha256").update(text, "utf8").digest("hex") === expected[0]) same++; else different.push({ path: page.path, mine: text.length, django: expected[1] }); } console.log(`${pages.length} pages: ${same} identical text, ${different.length} different, ${missing} not in the Django index`); for (const d of different.slice(0, 8)) console.log(" ", JSON.stringify(d)); FIRST RUN: 4,413 of 4,420 identical, 7 different. The seven pages came from three causes (a fourth difference showed up later, in a unit test): 1. "" (a React fragment written in a code sample) was kept as text here; a parser ignores it. Dropped now. 2. "⅖" (the fraction two fifths) was not a name my decoder knew. A real HTML name, added. 3. In TWO pages (a Romaji course chapter and the Build Tooling combined course) the text contains a raw "