The Content App
Learning Website with Django
Chapter 3 ยท The Content App
The project can now tell which site a request is for. It cannot yet tell what pages exist. This chapter builds
the content app: the models that describe a page and a course, and the code that reads a content
file and works out those facts from the file itself. It does not import anything yet; Chapter 4 does that. The
order is deliberate: first teach the project what a page is, then run it over thousands of real ones.
What the Database Is For
The pages themselves are files, and they stay files: you write them, you commit them, and the PDF builders read them. The database does not replace that. It holds what the site must query quickly and repeatedly:
| Question the site asks | Answered from |
|---|---|
| Which pages are on this site? | The database (site, kind) |
| What are the chapters of this course, in order? | The database (course, chapter number) |
| What was updated recently? | The database (updated date) |
| What does this page say? | The file (Chapter 4 decides how it is read) |
A page's identity is its path, relative to the content folder, with forward slashes:
hungary/hungarian-basic-3/hungarian_basic_conversation_3_1.html. It is unique, it is stable, and the
page's address is just that path without .html. Nothing needs a separate id.
The Two Models
| Model | Field | Notes |
|---|---|---|
| Course | site, folder | The folder is unique, such as hungary/hungarian-basic-3 |
name | From the banner's Course: line | |
course_no | Optional | |
| A course is just a folder of numbered chapters | ||
| Page | path | Unique, relative, forward slashes, ends in .html, never contains .. |
site, kind | The site comes from the path; the kind is one of five (below) | |
title | Found in the file (below) | |
course, chapter_no | Optional; deleting a course keeps its pages | |
created, updated | Optional: many older pages have no dates | |
| Ordered by course, then chapter number, then path | ||
Ordering by a number matters. A test inserts chapters 10, 2, 1 and 11 and reads them back as 1, 2, 10, 11. A sort on file names would give 1, 10, 11, 2, a trap from Learning Website: Framework & Architecture 3.
Keeping the Table Honest
- Unsafe paths are refused. A path must be relative, use forward slashes, end in
.htmland contain no... Four bad examples are tested, so a path can never point outside the content folder. - The site must agree with the path.
clean()looks the path up in the site map and rejects a row that sayssystemsfor ahungary/...path. The table cannot contradict the map. - A path in no site is an error. The lookup,
site_for_path(), raisesNoSiteErrorinstead of guessing, and the importer will report it. - Sidebar pages follow their subject.
sidebar/linux/...belongs to the systems site, as decided in Learning Website: Framework & Architecture 2.
Django validates these rules in full_clean(), not on every save(). The importer will call
full_clean() explicitly, so a bad row is reported, not stored.
Reading a File
The parser takes a file's text and its path and returns every value a Page needs. It starts from the
banner shapes catalogued in Learning Website: Framework & Architecture 3, and the tests use a real file of each
shape:
| Banner shape | Kind stored | Example parsed |
|---|---|---|
| Course and Chapter, with dates | course_chapter | Hungarian Basic Conversation 3, chapter 1, dated 2026-10-05 |
| Course and Chapter, no dates (older) | course_chapter | Apache In Depth, chapter 1: dates stay empty |
| Category and Subcategory | sidebar | The Git command line cheat sheet, on the programming site |
| Free-form headline | lesson | The Python boolean logic lesson: title from its first heading |
| Free-form, no title line | lesson | The Hungarian alphabet reference: title from the file name |
| A complete HTML page | full_page | A kanji page: title from <title> |
| No banner | other | Seven pages in the real content (below) |
The title follows the same order as the live site: the banner's Chapter: line, then
the page's own <title>, then the first <h1> or <h2>, then
the file name made readable. Course chapters have no real heading of their own, so the banner is the only place
their title lives. Entities such as & are decoded, so a title reads “Fit &
Returns”, not “Fit & Returns”.
Chapter numbers come from the file name, and the names are not uniform
Course and chapter numbers are not in the banner, so the parser reads them from the file name. The real content has several generations of names, and the first version of the parser handled only the newest. Counting the course-chapter pages by pattern shows how much would have been wrong:
| Pattern | Example | Pages | Course, chapter |
|---|---|---|---|
| Newest | hungarian_basic_conversation_3_10.html | 2,950 | 3, 10 |
| Older prompt-prefix | js1-4.html, psp-croll-1-2.html | 495 | 1, 4 and 1, 2 |
| One trailing number | setting_up_a_web_server_on_debian_03.html | 367 | none, 3 |
| Lesson, chapter first | html_lesson_01_3.html | 62 | 3, 1 (the numbers are the other way round) |
| No number | linux_appendix_a.html | 21 at first, 3 once the other forms were handled | none, none |
html_lesson_01_3.html, put the chapter first. A parser that
always reads “course then chapter” gives every chapter of the advanced HTML course the number 3, and
nothing complains. It was found only because the dry run checked that each course's chapters run 1 to N, and
twelve of them came out as 3, 3, 3, 3… Check the numbers against a sequence, not just against “is it a number”.
The Dry Run
Before any import, Exercise 3 parses every real page and reports what it finds, writing nothing to the database:
| Result | Value |
|---|---|
| Pages parsed | 4,402 of 4,402 |
| By kind | 3,895 course chapters, 410 complete pages, 47 sidebar, 43 lessons, 7 with no banner |
| By site | programming 1,370; humanities 707; systems 653; languages 561; webdevelopment 514; lifeskills 253; ai 187; creative 157 |
| Course folders | 395, with no folder using two different course names |
| Paths in no site | 0: the site map is complete |
| Chapter numbers not running 1 to N | 1 course (PHP Fundamentals, because its first chapter has no banner) |
Banner File: line differs from the real file | 7 pages, all in Setting Up a Web Server on Debian |
php1-1.html have no banner comment at all. They are real content problems, not parser problems, and
php1-1.html is why PHP Fundamentals seems to start at chapter 2. Adding a banner to each is a small
fix, listed in the review notes.
Hands-On Exercises
Add the site lookup for a path, the banner parser and the page parser to the Chapter 2 project. Copy a real file of each banner shape into tests/fixtures, write tests for each shape and for every file-name pattern, and print what the parser finds in each sample.
Create the Course and Page models and their migration, and test them against a real database: one row per path, chapters in numeric order, unsafe paths and mismatched sites rejected, and a course deleted without losing its pages. Read the SQL Django generates.
Write a dry-run script that parses every page of the real content folder and reports the counts by kind and site, paths in no site, titles that fell back to the file name, banner file-name mismatches, and courses whose chapter numbers do not run 1 to N. Fix the parser for whatever it finds, then list the content problems that remain.
๐ View solutionChapter 3 Quick Reference
- The files stay the source; the database holds what the site must query: pages by site, chapters in order, recent updates
- A page's identity is its path (relative, forward slashes,
.html); its address is the path without.html - Models:
Course(a folder of chapters) andPage(path, site, kind, title, course, chapter number, dates) - Order by the chapter number; a text sort puts 10 before 2
- Kinds: course chapter, sidebar, lesson, complete page, other; only fields that exist are filled in
- Title order: banner
Chapter:,<title>, first h1/h2, then the file name - File names have generations:
_3_10,js1-4, a single number,_lesson_01_3(chapter first), none site_for_path()raisesNoSiteErrorfor a path in no site;clean()checks the stored site matches- Validation runs in
full_clean(), so the importer must call it - Dry-run before importing: 4,402 pages parse, 0 unmapped, 7 pages have no banner