The Content Model

Learning Website with Next.js

Chapter 3 ยท The Content Model in TypeScript

Every site needs the same answer to “what is this file?”: which site it belongs to, what kind of page it is, its title, its course and chapter, and its dates. If each app worked that out for itself, they would drift apart. So the answer becomes one set of TypeScript types and one parser, in a package that every app imports. This chapter builds it, and then does the thing a second implementation owes you: it runs it on the real content and compares every page with what the Django project stored.

Run for real, on 4,414 pages
The package has 17 tests of its own (26 with the site package, all passing) and was run over the real content folder. It does not read or show page bodies; that is Chapter 4.

The Types

export type Kind = "course_chapter" | "sidebar" | "lesson" | "full_page" | "other"; export interface Page { readonly path: string; // relative to the content folder, forward slashes readonly site: SiteName; readonly kind: Kind; readonly title: string; readonly summary: string; // the banner's Topic line, at most 300 characters, or "" readonly courseName: string | null; readonly courseNo: number | null; readonly chapterNo: number | null; readonly created: IsoDate | null; readonly updated: IsoDate | null; }
  • The path is BASE-relative: hungary/hungarian-basic-3/x.html, with forward slashes. Nothing in the model knows where the content folder is on disk, so the same pages work on any machine.
  • Dates are ISO text ("2026-10-07"), not Date objects, so no time zone can move a date by a day.
  • Everything is readonly, and absence is null, never "" or 0 (the summary is the one string, empty when there is none).
  • The site is a SiteName from the shared package, so a page can only belong to a site that exists.

Reading the Banner

Every page begins with a comment such as Course: German Basic Conversation 3 / Chapter: At the Station / Topic: ... / Date Created: ... Date Updated: .... The parser reads the comment at the very top (a comment further down is not a banner), strips the ==== decoration, keeps the first value of a repeated key, and finds the dates with their own pattern. A line with a Course field makes a course chapter; a Category field, a sidebar page; any other comment, a lesson; a complete HTML page is recognised by its doctype and described by its <title>; no comment at all is “other”.

From a File to a Page

QuestionHow it is answered
TitleBanner Chapter, then <title>, then the first h1 or h2, then the file name made readable
Course and chapter numbersFrom the file name, which has come in five generations
SitesiteForPath() from Chapter 1; a path no site owns is a NoSiteError, never a guess
DatesMust be real: 2026-13-01 and 2025-02-29 are errors, 2024-02-29 is fine
Character references in a titleDecoded once: &amp;lt; becomes &lt;, not <
Long titlesCut at 300 characters, counting an emoji as one

The five generations of chapter file name, each with a test:

hungarian_basic_conversation_3_1.html     course 3, chapter 1
html_lesson_01_3.html                     chapter 1 of course 3 (the numbers are the other way round)
js1-4.html, psp-croll-1-2.html            course 1, chapter 4 / 2 (the older prompt-prefix names)
setting_up_a_web_server_on_debian_03.html chapter 3 only
linux_appendix_a.html                     no number at all

Reading a Whole Folder

loadPages(root) walks the content folder in a stable order and returns { pages, errors }. A bad file is reported in errors (not valid UTF-8, over 2 MB, a path that climbs out of the folder, an impossible date, a path no site owns) and never stops the others. But it re-throws anything that is not a content problem: a missing folder is ENOENT, not “a bad file”, so a real bug cannot hide in the error list. It skips files ending _print.html, files starting with _, the pdfs and solutions folders and the archive copy of the kanji pages. groupCourses() then groups chapters by folder in number order, so chapter 2 comes before chapter 10 and an unnumbered chapter goes last.

Is It the Same as the Django Importer?

Two implementations of the same rules should give the same answers, and the only honest way to say so is to run both on the real content. Django's stored pages were dumped to JSON, the TypeScript package parsed the same folder, and every field was compared:

ComparisonResult
Pages found4,414 in TypeScript, 4,414 in Django; none missing either way
Errors0
Site, kind, title, chapter number, created, updated, course number, course name, summary0 differences on every page
Courses396 in each; 0 differences
Time to read and parse the whole folder1.5 to 1.9 seconds (two runs)

The real pages are 3,907 course chapters, 410 complete pages, 47 sidebar pages, 43 lessons and 7 others. The time is not a comparison with Django, whose import also prepares and indexes every page.

Can this check fail? I made it
Zero differences on the first run is only worth something if the comparison can find one, so I broke the parser twice on purpose and ran it again:
  • Named character references no longer decoded: 11 title differences. The check works, and the small table of named references covers every real case.
  • The file-name-to-title function changed to give wrong answers: 0 differences. No real page takes its title from its file name, so that function is checked only by its unit tests, not by the real content.
The second result is the useful one: a check you have not seen fail has not been shown to check.
What was not verified
Page bodies (Chapter 4), links, and files Django would reject while preparing the body (there were none: both sides found the same 4,414). Only the common named character references are decoded; Python decodes all of them, so a future page with an unusual one in its title would differ and the comparison would show it.

Hands-On Exercises

Exercise 1

Define the TypeScript types for a page and a course, and write the banner reader. Explain each decision in the types (relative paths, ISO dates, readonly, null for absence).

๐Ÿ“„ View solution
Exercise 2

Turn a file into a page (title order, file-name numbers, real dates, decoded entities) and a whole folder into pages and errors, refusing unsafe paths and re-throwing real bugs. Test the awkward cases.

๐Ÿ“„ View solution
Exercise 3

Compare the TypeScript reader with the Django importer on the real content, field by field, and show that your comparison can fail by breaking the code on purpose.

๐Ÿ“„ View solution

Chapter 3 Quick Reference

  • @lw/content holds the types (Page, Course, Kind) and the reader; every app imports it
  • Paths are relative to the content folder with forward slashes; dates are ISO text; fields are readonly; absence is null
  • Banner: the comment at the very top; Course makes a chapter, Category a sidebar page, any other comment a lesson, a doctype a full page
  • Title: banner Chapter, <title>, first h1/h2, then the file name; cut at 300 characters
  • Five file-name generations give course and chapter numbers; html_lesson_01_3 is chapter 1 of course 3
  • Impossible dates are errors; entities decode once; a path no site owns is a NoSiteError
  • loadPages returns { pages, errors }; one bad file never stops the rest, but real bugs are re-thrown
  • groupCourses: by folder, chapters in number order (2 before 10), unnumbered last
  • Real run: 4,414 pages, 0 errors, 396 courses, 0 field differences from Django, 1.5 to 1.9 s
  • The comparison was shown able to fail (11 differences with entities off), and it does not cover the file-name title fallback