learning-website-framework1-3 Exercise 2: A Page Model ======================================================= Define a Page data type and a load_page() function that turns one file into one record, with the site, kind, title, course and chapter numbers (from the filename), dates, and the matching chapter PDF if one exists. Save as model.py next to banner.py: import os, re from dataclasses import dataclass from typing import Optional from banner import parse_banner SITE_OF = {"hungary": "languages", "japan": "languages", "france": "languages", "germany": "languages", "web-servers": "webdevelopment", "sidebar": "sidebar-routed"} NEW = re.compile(r"_(\d+)_(\d+)\.html$") @dataclass class Page: path: str # relative to content/, always with forward slashes site: str kind: str # course-chapter | sidebar | free-banner | full-page title: Optional[str] course: Optional[str] = None course_no: Optional[int] = None chapter_no: Optional[int] = None created: Optional[str] = None updated: Optional[str] = None pdf: Optional[str] = None def load_page(content_root, rel): folder, name = os.path.split(rel) with open(os.path.join(content_root, rel), encoding="utf-8") as fh: info = parse_banner(fh.read(4000)) m = NEW.search(name) pdf = os.path.join(content_root, folder, "pdfs", name[:-5] + ".pdf") return Page( path=rel, site=SITE_OF[rel.split("/")[0]], kind=info["kind"], title=info.get("chapter") or info.get("title") or info.get("headline"), course=info.get("course"), course_no=int(m[1]) if m else None, chapter_no=int(m[2]) if m else None, created=info.get("created"), updated=info.get("updated"), pdf=os.path.relpath(pdf, content_root).replace(os.sep, "/") if os.path.exists(pdf) else None, ) if __name__ == "__main__": import sys, pprint root = sys.argv[1] for rel in sys.argv[2:]: pprint.pprint(load_page(root, rel), sort_dicts=False, width=100) Run it: python model.py "/claude-projects/website-content/content" hungary/hungarian-basic-3/hungarian_basic_conversation_3_1.html web-servers/apache-in-depth/apache_in_depth_1_1.html Output (checked by running it): Page(path='hungary/hungarian-basic-3/hungarian_basic_conversation_3_1.html', site='languages', kind='course-chapter', title='Buying Clothes: Sizes, Fit & Returns', course='Hungarian Basic Conversation 3', course_no=3, chapter_no=1, created='2026-10-05', updated='2026-10-05', pdf='hungary/hungarian-basic-3/pdfs/hungarian_basic_conversation_3_1.pdf') Page(path='web-servers/apache-in-depth/apache_in_depth_1_1.html', site='webdevelopment', kind='course-chapter', title="Scope: What This Course Builds On, and Why It's Apache-Only", course='Apache In Depth', course_no=1, chapter_no=1, created=None, updated=None, pdf=None) Design points ------------- - Fields a page may not have (dates, a PDF, a course number) are Optional and default to None. The Apache chapter shows why: it has no dates and no per-chapter PDF. - The path is stored relative to content/ with forward slashes, so the same record works on Windows, Linux and in a URL. - The site comes from the map in Learning Website: Framework & Architecture 2 (shortened here to the few folders the example needs). In the real build you would import the full map instead of repeating it. - Chapter and course numbers are read from the filename, not from the banner, because the banner holds the human-readable course name and not the number. - The model does not hold the HTML. Keep it small: metadata only, and read the fragment from disk when the page is built. WHY THIS WORKS AS AN ANSWER --------------------------- Everything later in the course (navigation, search, sitemaps, redirects) asks questions about pages. Putting the answers in one typed record, built in one place, means those questions are answered consistently and a change to the model is a change in one file.