The Content Model

Learning Website: Framework & Architecture

Chapter 3 · The Content Model

Chapter 2 decided which folder belongs to which site. Before any site can be built, you need to know what a page is: what kinds of page exist, where their information lives, and how a script can read it. On this site there is no database. The files are the database, and the content model is the description of how to read them. This chapter inspects the real files, finds the shapes they actually take, and defines a model that copes with all of them.

Every example here was run
The scripts in this chapter's exercises were run on your real content/ folder, and the outputs quoted in the solutions are what they printed. Where the real files turned out messier than the rules describe, the chapter says so.

The Kinds of Page

Several kinds of page share the content/ folder. They differ in where they live, how they are named and what information they carry:

KindWhere it livesNamingNotes
Course chapter<subject>/<course>/<course_name>_<course#>_<chapter#>.htmlNumbered; has a banner with Course and Chapter; may have pdfs/ and solutions/
Sidebar pagesidebar/<subject>/links_, sidebar_, cheat_sheet_ prefixes; tools use noneLinks, tools, standalone lessons, cheat sheets; banner has Category and Subcategory
Standalone lessonA lessons folder, such as python-lessons/<language>_lesson_<topic>.htmlNot numbered; a free-form banner
Reference pageAlongside lessons or under reference foldersDescriptive, such as hungarian_alphabet.htmlA free-form banner with no title line
Kanji pageTwo copies (archive and live)kanji_<character>.htmlA complete HTML page with its own <title>, not a fragment

Cheat sheets also have a printable twin: cheat_sheet_git_command_line.html sits beside cheat_sheet_git_command_line_print.html. The model has to know the second file is a variant of the first and not a page of its own.

Anatomy of a Course Folder

linux/system-administration/debian-development-machine-setup/ ├── _outline.md # planning notes, not published ├── debian_development_machine_setup_1_1.html # chapter files sit directly in the folder ├── debian_development_machine_setup_1_2.html ├── ... ├── pdfs/ # per-chapter PDFs and the course PDF └── solutions/ ├── devsetup1-1_exercise1.txt # keyed by the short prompt prefix └── devsetup1-1_exercise2.txt

Three details matter for the build. Files that start with an underscore (_outline.md) are working notes and are never published. The chapter files follow <course_name>_<course#>_<chapter#>.html, so the course number and chapter number can be read from the filename. And solution files use the short prompt prefix (devsetup1), not the long course name, so they cannot be found by guessing from the chapter filename. The chapter page links to them, and the build must follow the links.

Fragments and Full Pages

Nearly everything is an HTML fragment: a banner comment, a <style> block and markup, with no <html>, <head> or <body>. The site wraps each fragment in its own layout. The kanji pages are the exception. They are complete HTML pages, with their own head and title, and so they need different handling: the build must not wrap a full page inside a second page.

The Banner Shapes

Because the site has been built over a long time, the banner comment comes in several shapes. A parser that handles only the newest will silently mislabel the rest. These are the shapes that really exist:

ShapeFieldsExample
Course chapterCourse, Chapter, File, Topic, (Level), dates when presentHungarian Basic Conversation 3, chapter 1
Course chapter, olderThe same, with no datesApache In Depth, chapter 1
SidebarCategory, Subcategory, Chapter, File, datesThe Git command line cheat sheet
Free-form with a headlineA line such as PYTHON LESSON: …, then Level and datesThe Python boolean logic lesson
Free-form without a titleLevel and Focus onlyThe Hungarian alphabet reference
None (full page)Nothing: the title is in <title>The kanji pages
Tolerate the past, fix it gradually
You could rewrite every old banner to a single format, but that touches a very large number of files and risks breaking pages that work. A better approach is to make the parser accept every shape and return only the fields that are really there. New content then follows the current rules, and old content is updated only when it is rewritten for another reason.

Where Each Field Comes From

Every piece of information about a page lives in exactly one natural place, and the model should read it from there:

FieldRead it fromNotes
Subject and siteThe first folder in the pathThrough the site map from Learning Website: Framework & Architecture 2
Course name, chapter titleThe bannerThe reader-facing text, so it needs no guessing
Course number, chapter numberThe filenameThe banner does not carry them
Created and updated datesThe bannerOptional: missing on older pages
Chapter PDFThe pdfs/ folder next to the chapterOptional: not every course has per-chapter PDFs
Solution filesLinks inside the chapterFollow the links; do not guess filenames
KindBanner shape plus folderCourse chapter, sidebar page, lesson, reference or full page

The dates are worth keeping carefully. The site's own rule is that the created date never changes and the updated date moves when the content really changes. That makes the updated date a ready-made value for the lastmod field in a sitemap and for a “recently updated” list (Learning Website: Framework & Architecture 7 uses both).

A Model That Survives Real Data

A Page record gathers those fields in one place. The design choices that keep it from breaking on old content:

  • Optional fields are really optional. Dates, PDFs and course numbers can be None. A page is not invalid for lacking them.
  • Paths are relative to content/ and use forward slashes, so the same record works on Windows, on the Debian server and in a URL.
  • The record holds metadata, not the HTML. Read the fragment from disk when you build the page, so the index stays small.
  • Sort numbers as numbers. Filenames sort as text, so chapter 10 appears before chapter 2. A chapter list built from a directory listing is wrong until it sorts on the chapter number.
  • Skip what is not a page: files beginning with an underscore, print variants, pdfs/ and solutions/.
@dataclass class Page: path: str # relative to content/, forward slashes site: str # from the site map kind: str # course-chapter | sidebar | free-banner | full-page title: Optional[str] course: Optional[str] = None course_no: Optional[int] = None chapter_no: Optional[int] = None created: Optional[str] = None updated: Optional[str] = None pdf: Optional[str] = None
Do not store the same fact twice
The model reads the course name from the banner and the course number from the filename, and it does not copy either into a separate list. If a second file ever holds the same information (a hand-edited course index, for example), the two will drift apart. Prefer deriving the data from the files every time the site is built.

Validate Before You Build

Once pages are records, content problems become checkable: a banner whose File: line does not match the real filename, a chapter missing from the sequence, a file that does not follow the naming pattern. A validator that runs before the build finds these cheaply. Exercise 3 writes one and runs it on two real courses of different ages. Both pass, and the output shows exactly why the model has to treat dates and PDFs as optional.

Hands-On Exercises

Exercise 1

Write parse_banner(text) so it handles every banner shape listed in this chapter, including pages with no banner. Run it on one real file of each shape and check the fields it returns.

📄 View solution
Exercise 2

Define a Page data type and a load_page() function that builds one from a file path: site, kind, title, course and chapter numbers from the filename, the dates, and the chapter PDF if it exists. Run it on a new course chapter and an older one.

📄 View solution
Exercise 3

Write validate_course.py, which indexes one course folder and reports a wrong banner kind, a File: line that does not match the filename, a filename that breaks the pattern, and gaps in the chapter numbers. Run it on a recent course and an older one and explain the differences.

📄 View solution

Chapter 3 Quick Reference

  • The files are the database: the content model describes how to read them
  • Kinds of page: course chapter, sidebar page, standalone lesson, reference page, kanji (full) page
  • Nearly all pages are fragments; kanji pages are full HTML pages and must not be wrapped again
  • Banner shapes: course chapter (with or without dates), sidebar, free-form with a headline, free-form without a title, none
  • Course and chapter numbers come from the filename; names, titles and dates from the banner
  • Solution files use the short prefix, so follow the links in the chapter instead of guessing
  • Optional means optional: older pages have no dates and no per-chapter PDF
  • Sort chapter numbers as integers, never as text
  • Skip underscore files, _print variants and the pdfs/ and solutions/ folders
  • Validate before building so content problems fail early