learning-website-django1-3 Exercise 1: Parse a Content File =========================================================== Teach the project to read what a content file says about itself, using the banner shapes that really exist on the site (Learning Website: Framework & Architecture 3), plus the site lookup for a path. PART A. The script below adds, to the Chapter 2 project: - site_for_path() in config/sites_config.py: which site owns a content path (sidebar pages follow their subject folder; an unknown path raises an error) - apps/content/banner.py: parse_banner() - apps/content/parsing.py: parse_page(), which returns every value a Page needs - apps/content/models.py: the Course and Page models (Exercise 2) Save as make_content_app.py and run it with the project folder: python make_content_app.py . """make_content_app.py Adds Chapter 3 (the content app) to the project: the site lookup for a path, the banner parser, the page parser, and the Course and Page models. Safe to run twice.""" import os, sys root = sys.argv[1] def write(rel, text): p = os.path.join(root, *rel.split("/")) os.makedirs(os.path.dirname(p), exist_ok=True) with open(p, "w", encoding="utf-8", newline="\n") as fh: fh.write(text) def patch(rel, old, new): p = os.path.join(root, *rel.split("/")) text = open(p, encoding="utf-8").read() if new in text: return assert old in text, f"{rel}: expected text not found" open(p, "w", encoding="utf-8", newline="\n").write(text.replace(old, new, 1)) # 1. which site does a content path belong to? -------------------------------------------------- patch("config/sites_config.py", "def site_hosts(env, only=None):", '''# pages under sidebar// follow their subject; a few URLs are not content folders SIDEBAR_ROUTES = { "ai": "ai", "drinks": "lifeskills", "football": "humanities", "history": "humanities", "linux": "systems", "programming": "programming", "web-development": "webdevelopment", } SPECIAL_PREFIXES = {"resources/japanese": "languages", "resources/hungarian": "languages"} class NoSiteError(LookupError): """A path that belongs to no site: fail loudly instead of guessing.""" def site_for_path(path): """The site that owns a root-relative content path such as "hungary/hungarian-basic-3/x.html".""" parts = path.strip("/").split("/") joined = "/".join(parts) for prefix, site in SPECIAL_PREFIXES.items(): if joined == prefix or joined.startswith(prefix + "/"): return site if parts[0] == "sidebar": if len(parts) > 1 and parts[1] in SIDEBAR_ROUTES: return SIDEBAR_ROUTES[parts[1]] raise NoSiteError(path) for site, info in SITES.items(): if parts[0] in info["folders"]: return site raise NoSiteError(path) def site_hosts(env, only=None):''') # 2. the banner parser: the shapes of banner that really exist on the site ----------------------------- write("apps/content/banner.py", '''"""Read the banner comment at the top of a content file (or the of a full page).""" import re KEY = re.compile(r"^([A-Za-z][A-Za-z ]*?):\\s*(.*)$") DATES = re.compile(r"Date Created:\\s*(\\d{4}-\\d{2}-\\d{2})\\s+Date Updated:\\s*(\\d{4}-\\d{2}-\\d{2})") def parse_banner(text): """Return a dict. Its "kind" is one of: course-chapter, sidebar, free-banner, full-page, no-banner. Only fields that are really present are returned (older pages have no dates).""" if text.lstrip().lower().startswith("<!doctype"): m = re.search(r"<title>(.*?)", text, re.S) return {"kind": "full-page", "title": m.group(1).strip() if m else None} m = re.match(r"\\s*", text, re.S) if not m: return {"kind": "no-banner"} block = m.group(1) info = {} d = DATES.search(block) if d: info["created"], info["updated"] = d.group(1), d.group(2) lines = [ln.strip().strip("=").strip() for ln in block.splitlines()] lines = [ln for ln in lines if ln] for ln in lines: km = KEY.match(ln) if km and km.group(1).strip() not in ("Date Created", "Date Updated"): info.setdefault(km.group(1).strip().lower(), km.group(2).strip()) if "course" in info: info["kind"] = "course-chapter" elif "category" in info: info["kind"] = "sidebar" else: info["kind"] = "free-banner" info["headline"] = lines[0] if lines else "" return info ''') # 3. the page parser -------------------------------------------------------------------------------- write("apps/content/parsing.py", '''"""Turn one content file into the values stored on a Page.""" import html import re from dataclasses import dataclass from datetime import date from typing import Optional from config.sites_config import site_for_path from .banner import parse_banner # chapter file names come in several generations: # hungarian_basic_conversation_3_1.html (course 3, chapter 1) # html_lesson_01_3.html (chapter 1 of course 3: the numbers are the other way round) # js1-4.html, psp-croll-1-2.html (course 1, chapter 4 / 2: the older prompt-prefix names) # setting_up_a_web_server_on_debian_03.html (chapter 3 only) # linux_appendix_a.html (no number at all) LESSON_NUMBERS = re.compile(r"_lesson_(\\d+)_(\\d+)\\.html$") NEW_NUMBERS = re.compile(r"_(\\d+)_(\\d+)\\.html$") OLD_NUMBERS = re.compile(r"[A-Za-z_-](\\d+)-(\\d+)\\.html$") ONE_NUMBER = re.compile(r"_(\\d+)\\.html$") def numbers_from(path): """(course_no, chapter_no) read from the file name; either may be None.""" m = LESSON_NUMBERS.search(path) if m: return int(m[2]), int(m[1]) for pattern in (NEW_NUMBERS, OLD_NUMBERS): m = pattern.search(path) if m: return int(m[1]), int(m[2]) m = ONE_NUMBER.search(path) return (None, int(m[1])) if m else (None, None) HEADING = re.compile(r"]*>(.*?)", re.I | re.S) TAGS = re.compile(r"<[^>]+>") KINDS = {"course-chapter": "course_chapter", "sidebar": "sidebar", "free-banner": "lesson", "full-page": "full_page", "no-banner": "other"} @dataclass(frozen=True) class ParsedPage: path: str site: str kind: str title: str course_name: Optional[str] = None course_no: Optional[int] = None chapter_no: Optional[int] = None created: Optional[date] = None updated: Optional[date] = None @property def course_folder(self): return self.path.rsplit("/", 1)[0] if self.kind == "course_chapter" else None def humanize(filename): stem = filename.rsplit(".", 1)[0] return " ".join(w.capitalize() for w in re.split(r"[-_]+", stem) if w) def best_title(raw, banner, filename): """The same order the live site uses: banner Chapter, , first h1/h2, then the file name.""" if banner.get("chapter"): return html.unescape(banner["chapter"]) if banner.get("title"): return html.unescape(banner["title"]) m = HEADING.search(raw) if m: text = html.unescape(TAGS.sub("", m.group(1))).strip() if text: return text return humanize(filename) def parse_page(raw, path): banner = parse_banner(raw[:6000]) course_no, chapter_no = numbers_from(path) if banner["kind"] == "course-chapter" else (None, None) return ParsedPage( path=path, site=site_for_path(path), kind=KINDS[banner["kind"]], title=best_title(raw, banner, path.rsplit("/", 1)[-1]), course_name=banner.get("course"), course_no=course_no, chapter_no=chapter_no, created=date.fromisoformat(banner["created"]) if "created" in banner else None, updated=date.fromisoformat(banner["updated"]) if "updated" in banner else None, ) ''') # 4. the models -------------------------------------------------------------------------------------- write("apps/content/models.py", '''from django.core.exceptions import ValidationError from django.db import models from config.sites_config import NoSiteError, SITES, site_for_path SITE_CHOICES = [(name, info["title"]) for name, info in SITES.items()] def validate_content_path(value): """A content path is relative, uses forward slashes, ends in .html and never climbs out of the tree.""" if "\\\\" in value or value.startswith("/") or ".." in value.split("/"): raise ValidationError("A content path must be relative, use forward slashes and contain no '..'.") if not value.endswith(".html"): raise ValidationError("A content path must end in .html.") class Course(models.Model): """A folder of numbered chapters, such as hungary/hungarian-basic-3.""" site = models.CharField(max_length=40, choices=SITE_CHOICES) folder = models.CharField(max_length=400, unique=True) name = models.CharField(max_length=200) course_no = models.PositiveSmallIntegerField(null=True, blank=True) class Meta: ordering = ["site", "folder"] def __str__(self): return self.name class Page(models.Model): class Kind(models.TextChoices): COURSE_CHAPTER = "course_chapter", "Course chapter" SIDEBAR = "sidebar", "Sidebar page" LESSON = "lesson", "Standalone lesson or reference" FULL_PAGE = "full_page", "Complete HTML page" OTHER = "other", "Other" path = models.CharField(max_length=500, unique=True, validators=[validate_content_path]) site = models.CharField(max_length=40, choices=SITE_CHOICES, db_index=True) kind = models.CharField(max_length=20, choices=Kind.choices) title = models.CharField(max_length=300) course = models.ForeignKey(Course, null=True, blank=True, on_delete=models.SET_NULL, related_name="pages") chapter_no = models.PositiveSmallIntegerField(null=True, blank=True) created = models.DateField(null=True, blank=True) updated = models.DateField(null=True, blank=True) class Meta: # chapter_no is a NUMBER, so chapter 2 sorts before chapter 10 (a text sort would not) ordering = ["course_id", "chapter_no", "path"] indexes = [models.Index(fields=["site", "kind"])] def __str__(self): return self.title @property def url_path(self): """/hungary/hungarian-basic-3/hungarian_basic_conversation_3_1/ (the live site's address form)""" return "/" + self.path[: -len(".html")] + "/" def clean(self): super().clean() try: expected = site_for_path(self.path) except NoSiteError: raise ValidationError({"path": "This path belongs to no site."}) if self.site != expected: raise ValidationError({"site": f"This path belongs to the {expected} site, not {self.site}."}) ''') print("content app written to", root) PART B. Copy six real files into tests/fixtures (one of each banner shape): a recent course chapter, an older chapter with no dates, a sidebar page, a standalone lesson with a headline banner, a reference page with no title line, and a complete HTML page (a kanji page; copy it with a plain file name such as kanji_jou.html). Save as tests/test_content_parsing.py: from datetime import date from pathlib import Path from django.test import SimpleTestCase from apps.content.banner import parse_banner from apps.content.parsing import parse_page from config.sites_config import NoSiteError, site_for_path FIXTURES = Path(__file__).parent / "fixtures" def raw(name): return (FIXTURES / name).read_text(encoding="utf-8") class SiteForPathTests(SimpleTestCase): def test_ordinary_folders(self): self.assertEqual(site_for_path("hungary/hungarian-basic-3/x.html"), "languages") self.assertEqual(site_for_path("web-servers/apache-in-depth/x.html"), "webdevelopment") def test_sidebar_pages_follow_their_subject(self): self.assertEqual(site_for_path("sidebar/linux/x.html"), "systems") self.assertEqual(site_for_path("sidebar/programming/x.html"), "programming") def test_special_resource_prefixes(self): self.assertEqual(site_for_path("resources/japanese/kanji/kanji_x.html"), "languages") def test_a_path_in_no_site_fails_loudly(self): for bad in ("nonsense/x.html", "sidebar/x.html", "sidebar/unknown/x.html"): with self.assertRaises(NoSiteError): site_for_path(bad) class BannerShapeTests(SimpleTestCase): def test_course_chapter_with_dates(self): info = parse_banner(raw("hungarian_basic_conversation_3_1.html")[:6000]) self.assertEqual(info["kind"], "course-chapter") self.assertEqual(info["course"], "Hungarian Basic Conversation 3") self.assertEqual((info["created"], info["updated"]), ("2026-10-05", "2026-10-05")) def test_older_course_chapter_has_no_dates(self): info = parse_banner(raw("apache_in_depth_1_1.html")[:6000]) self.assertEqual(info["kind"], "course-chapter") self.assertNotIn("created", info) def test_sidebar_banner(self): info = parse_banner(raw("cheat_sheet_git_command_line.html")[:6000]) self.assertEqual((info["kind"], info["category"], info["subcategory"]), ("sidebar", "Sidebar", "Programming")) def test_free_form_banner_with_a_headline(self): info = parse_banner(raw("python_lesson_boolean_logic_truthy_falsy.html")[:6000]) self.assertEqual(info["kind"], "free-banner") self.assertTrue(info["headline"].startswith("PYTHON LESSON")) def test_free_form_banner_without_a_title_line(self): info = parse_banner(raw("hungarian_alphabet.html")[:6000]) self.assertEqual(info["kind"], "free-banner") def test_complete_page_has_a_title_and_no_banner(self): info = parse_banner(raw("kanji_jou.html")[:6000]) self.assertEqual(info["kind"], "full-page") self.assertIn("Kanji", info["title"]) def test_text_with_no_banner(self): self.assertEqual(parse_banner("<p>hello</p>")["kind"], "no-banner") class ParsePageTests(SimpleTestCase): def test_a_new_course_chapter(self): p = parse_page(raw("hungarian_basic_conversation_3_1.html"), "hungary/hungarian-basic-3/hungarian_basic_conversation_3_1.html") self.assertEqual((p.site, p.kind, p.course_no, p.chapter_no), ("languages", "course_chapter", 3, 1)) self.assertEqual(p.title, "Buying Clothes: Sizes, Fit & Returns") self.assertEqual(p.created, date(2026, 10, 5)) self.assertEqual(p.course_folder, "hungary/hungarian-basic-3") def test_an_older_chapter_still_parses(self): p = parse_page(raw("apache_in_depth_1_1.html"), "web-servers/apache-in-depth/apache_in_depth_1_1.html") self.assertEqual((p.site, p.chapter_no, p.created), ("webdevelopment", 1, None)) self.assertTrue(p.title.startswith("Scope:")) def test_the_title_falls_back_to_the_first_heading(self): p = parse_page(raw("python_lesson_boolean_logic_truthy_falsy.html"), "programming/general-purpose-languages/python/python-lessons/python_lesson_boolean_logic_truthy_falsy.html") self.assertEqual(p.kind, "lesson") self.assertIn("Boolean Logic", p.title) self.assertIsNone(p.chapter_no) def test_a_complete_page_keeps_its_title(self): p = parse_page(raw("kanji_jou.html"), "japan/japanese-language/reference-materials/kanji/kanji_jou.html") self.assertEqual((p.kind, p.site), ("full_page", "languages")) def test_the_title_falls_back_to_the_file_name_when_nothing_else_exists(self): p = parse_page("<p>no banner, no heading</p>", "france/french-lessons/french_lesson_the_market.html") self.assertEqual(p.title, "French Lesson The Market") class FileNameNumberTests(SimpleTestCase): BANNER = """<!-- Course: Some Course Chapter: A Title File: x.html --> <p>hi</p>""" def numbers(self, path): p = parse_page(self.BANNER, path) return p.course_no, p.chapter_no def test_the_newest_names_carry_course_and_chapter(self): self.assertEqual(self.numbers("hungary/h/hungarian_basic_conversation_3_10.html"), (3, 10)) def test_lesson_names_have_the_chapter_first(self): self.assertEqual(self.numbers("web-development/front-end-development/html/advanced/html_lesson_01_3.html"), (3, 1)) def test_the_older_prompt_prefix_names(self): self.assertEqual(self.numbers("web-development/scripting-and-backend/javascript/fundamentals/js1-4.html"), (1, 4)) self.assertEqual(self.numbers("projects/crunchyroll-downloader/psp-croll-1-2.html"), (1, 2)) def test_a_single_trailing_number_is_a_chapter(self): self.assertEqual(self.numbers("projects/web-server-on-debian/setting_up_a_web_server_on_debian_03.html"), (None, 3)) def test_a_name_with_no_number_has_none(self): self.assertEqual(self.numbers("linux/system-administration/linux_appendix_a.html"), (None, None)) Run all the tests, Django 6.1.2: python manage.py test tests Found 53 test(s). System check identified no issues (0 silenced). Creating test database for alias 'default'... ..................................................... ---------------------------------------------------------------------- Ran 53 tests in 0.061s OK Destroying test database for alias 'default'... PART C. Print what the parser finds in each sample (show_parsed.py): """show_parsed.py: parse the six sample files in tests/fixtures and print what the parser found.""" import os os.environ.setdefault("DJANGO_SETTINGS_MODULE", "config.settings.dev") import django django.setup() from pathlib import Path from apps.content.parsing import parse_page SAMPLES = [ ("hungarian_basic_conversation_3_1.html", "hungary/hungarian-basic-3/hungarian_basic_conversation_3_1.html"), ("apache_in_depth_1_1.html", "web-servers/apache-in-depth/apache_in_depth_1_1.html"), ("cheat_sheet_git_command_line.html", "sidebar/programming/cheat_sheet_git_command_line.html"), ("python_lesson_boolean_logic_truthy_falsy.html", "programming/general-purpose-languages/python/python-lessons/python_lesson_boolean_logic_truthy_falsy.html"), ("hungarian_alphabet.html", "hungary/hungarian-lessons/hungarian_alphabet.html"), ("kanji_jou.html", "japan/japanese-language/reference-materials/kanji/kanji_jou.html"), ] for fixture, path in SAMPLES: raw = (Path("tests/fixtures") / fixture).read_text(encoding="utf-8") p = parse_page(raw, path) print(path.rsplit("/", 1)[-1]) print(f" site={p.site} kind={p.kind} course={p.course_name!r} course_no={p.course_no} chapter_no={p.chapter_no}") print(f" title={p.title[:70]!r}") print(f" created={p.created} updated={p.updated}") python show_parsed.py hungarian_basic_conversation_3_1.html site=languages kind=course_chapter course='Hungarian Basic Conversation 3' course_no=3 chapter_no=1 title='Buying Clothes: Sizes, Fit & Returns' created=2026-10-05 updated=2026-10-05 apache_in_depth_1_1.html site=webdevelopment kind=course_chapter course='Apache In Depth' course_no=1 chapter_no=1 title="Scope: What This Course Builds On, and Why It's Apache-Only" created=None updated=None cheat_sheet_git_command_line.html site=programming kind=sidebar course=None course_no=None chapter_no=None title='Git Command Line Cheat Sheet' created=2026-10-06 updated=2026-10-06 python_lesson_boolean_logic_truthy_falsy.html site=programming kind=lesson course=None course_no=None chapter_no=None title='πŸ”€ Boolean Logic β€” What Python Actually Considers True or False' created=2026-08-16 updated=2026-08-16 hungarian_alphabet.html site=languages kind=lesson course=None course_no=None chapter_no=None title='Hungarian Alphabet' created=None updated=None kanji_jou.html site=languages kind=full_page course=None course_no=None chapter_no=None title='上 | Kanji | osztromok.com' created=None updated=None Design points ------------- - The title uses the same order as the live site: the banner's Chapter: line, then the page's <title>, then the first h1 or h2, then the file name made readable. Course chapters have no real heading of their own, so the banner is the only place their title lives. - Only fields that are present are filled in. The Apache chapter has no dates, so created and updated are None, not a made-up date. - Chapter file names come in generations, each with a test: _3_10 (course then chapter), _lesson_01_3 (chapter THEN course), js1-4 and psp-croll-1-2 (the older prompt-prefix names), a single trailing number, and no number at all. - A path in no site raises NoSiteError. The importer (Chapter 4) must stop or report, never guess. WHY THIS WORKS AS AN ANSWER --------------------------- The parser is plain Python with no database, so it is fast to test, and every banner shape has a real example in the tests. A shape that is not tested is a shape that will surprise you when you import the whole tree.