learning-website-framework1-3 Exercise 3: Validate a Course Folder =================================================================== Write a script that indexes one course folder using your parser and reports problems before the build does: wrong banner kind, a File: line that does not match the real filename, a filename that does not follow the pattern, and gaps in the chapter numbers. Save as validate_course.py next to banner.py: import os, re, sys from banner import parse_banner NEW = re.compile(r"^(?P[a-z0-9_]+)_(?P\d+)_(?P\d+)\.html$") def index_course(folder): pages, problems = [], [] for name in sorted(os.listdir(folder)): if not name.endswith(".html") or name.endswith("_print.html"): continue with open(os.path.join(folder, name), encoding="utf-8") as fh: info = parse_banner(fh.read(4000)) m = NEW.match(name) page = { "file": name, "kind": info.get("kind"), "title": info.get("chapter"), "course": info.get("course"), "chapter_no": int(m["chapter"]) if m else None, "has_dates": "created" in info, "pdf": os.path.exists(os.path.join(folder, "pdfs", name[:-5] + ".pdf")), } pages.append(page) if info.get("kind") != "course-chapter": problems.append(f"{name}: banner kind is {info.get('kind')}, expected course-chapter") if info.get("file") and info["file"] != name: problems.append(f"{name}: banner says File: {info['file']}") if not m: problems.append(f"{name}: filename does not match __.html") numbers = sorted(p["chapter_no"] for p in pages if p["chapter_no"] is not None) if numbers and numbers != list(range(1, len(numbers) + 1)): problems.append(f"chapter numbers are not 1..{len(numbers)}: {numbers}") return pages, problems if __name__ == "__main__": folder = sys.argv[1] pages, problems = index_course(folder) print(f"{len(pages)} chapter files in {os.path.basename(folder)}") print(f"with dates in banner: {sum(p['has_dates'] for p in pages)}") print(f"with a chapter PDF: {sum(p['pdf'] for p in pages)}") for p in pages[:3]: print(f" {p['chapter_no']:>2} {p['title']}") print("Problems:" if problems else "No problems found.") for pr in problems: print(" " + pr) Run it on two real courses: python validate_course.py "/claude-projects/website-content/content/hungary/hungarian-basic-3" python validate_course.py "/claude-projects/website-content/content/web-servers/apache-in-depth" Output (checked by running it): 12 chapter files in hungarian-basic-3 with dates in banner: 12 with a chapter PDF: 12 1 Buying Clothes: Sizes, Fit & Returns 10 Wishes, Hopes & What You Would Do 11 Customs, Festivals & Traditions No problems found. 10 chapter files in apache-in-depth with dates in banner: 0 with a chapter PDF: 0 1 Scope: What This Course Builds On, and Why It's Apache-Only 10 Capstone: Designing and Hardening a Production Apache Deployment 2 MPMs In Depth: Prefork, Worker & Event No problems found. Two things to notice -------------------- 1. The Apache course is older. It has no dates in its banners and no per-chapter PDFs, yet it validates, because the model treats those as optional. A stricter rule ("every page must have dates") would fail hundreds of legitimate old pages. 2. The first three chapters printed are 1, 10, 11 for one course and 1, 10, 2 for the other. Files sort as text, so "10" comes before "2". The script sorts the chapter numbers as integers when it checks for gaps, and the build must do the same when it builds the chapter list. This is a classic bug on content sites. WHY THIS WORKS AS AN ANSWER --------------------------- A validator is a cheap way to find content problems before they become broken pages. It also documents the rules the content is expected to follow. Run it over every course folder as part of the build and fail on errors, so a typo in a banner or a missing chapter never reaches a visitor.