The Content App

Learning Website with Django

Chapter 3 ยท The Content App

The project can now tell which site a request is for. It cannot yet tell what pages exist. This chapter builds the content app: the models that describe a page and a course, and the code that reads a content file and works out those facts from the file itself. It does not import anything yet; Chapter 4 does that. The order is deliberate: first teach the project what a page is, then run it over thousands of real ones.

Run for real, on Django 6.1
The app has 53 tests, all passing, and was run over every page of the real content folder in a dry run (Exercise 3). The numbers quoted below are from that run. The page counts differ from the framework course because that course also counted directory index pages; this chapter counts content files.

What the Database Is For

The pages themselves are files, and they stay files: you write them, you commit them, and the PDF builders read them. The database does not replace that. It holds what the site must query quickly and repeatedly:

Question the site asksAnswered from
Which pages are on this site?The database (site, kind)
What are the chapters of this course, in order?The database (course, chapter number)
What was updated recently?The database (updated date)
What does this page say?The file (Chapter 4 decides how it is read)

A page's identity is its path, relative to the content folder, with forward slashes: hungary/hungarian-basic-3/hungarian_basic_conversation_3_1.html. It is unique, it is stable, and the page's address is just that path without .html. Nothing needs a separate id.

The Two Models

ModelFieldNotes
Coursesite, folderThe folder is unique, such as hungary/hungarian-basic-3
nameFrom the banner's Course: line
course_noOptional
A course is just a folder of numbered chapters
PagepathUnique, relative, forward slashes, ends in .html, never contains ..
site, kindThe site comes from the path; the kind is one of five (below)
titleFound in the file (below)
course, chapter_noOptional; deleting a course keeps its pages
created, updatedOptional: many older pages have no dates
Ordered by course, then chapter number, then path
class Page(models.Model): class Kind(models.TextChoices): COURSE_CHAPTER = "course_chapter", "Course chapter" SIDEBAR = "sidebar", "Sidebar page" LESSON = "lesson", "Standalone lesson or reference" FULL_PAGE = "full_page", "Complete HTML page" OTHER = "other", "Other" path = models.CharField(max_length=500, unique=True, validators=[validate_content_path]) site = models.CharField(max_length=40, choices=SITE_CHOICES, db_index=True) ... class Meta: ordering = ["course_id", "chapter_no", "path"] # chapter_no is a NUMBER

Ordering by a number matters. A test inserts chapters 10, 2, 1 and 11 and reads them back as 1, 2, 10, 11. A sort on file names would give 1, 10, 11, 2, a trap from Learning Website: Framework & Architecture 3.

Keeping the Table Honest

  • Unsafe paths are refused. A path must be relative, use forward slashes, end in .html and contain no ... Four bad examples are tested, so a path can never point outside the content folder.
  • The site must agree with the path. clean() looks the path up in the site map and rejects a row that says systems for a hungary/... path. The table cannot contradict the map.
  • A path in no site is an error. The lookup, site_for_path(), raises NoSiteError instead of guessing, and the importer will report it.
  • Sidebar pages follow their subject. sidebar/linux/... belongs to the systems site, as decided in Learning Website: Framework & Architecture 2.

Django validates these rules in full_clean(), not on every save(). The importer will call full_clean() explicitly, so a bad row is reported, not stored.

Reading a File

The parser takes a file's text and its path and returns every value a Page needs. It starts from the banner shapes catalogued in Learning Website: Framework & Architecture 3, and the tests use a real file of each shape:

Banner shapeKind storedExample parsed
Course and Chapter, with datescourse_chapterHungarian Basic Conversation 3, chapter 1, dated 2026-10-05
Course and Chapter, no dates (older)course_chapterApache In Depth, chapter 1: dates stay empty
Category and SubcategorysidebarThe Git command line cheat sheet, on the programming site
Free-form headlinelessonThe Python boolean logic lesson: title from its first heading
Free-form, no title linelessonThe Hungarian alphabet reference: title from the file name
A complete HTML pagefull_pageA kanji page: title from <title>
No bannerotherSeven pages in the real content (below)

The title follows the same order as the live site: the banner's Chapter: line, then the page's own <title>, then the first <h1> or <h2>, then the file name made readable. Course chapters have no real heading of their own, so the banner is the only place their title lives. Entities such as &amp; are decoded, so a title reads “Fit & Returns”, not “Fit &amp; Returns”.

Chapter numbers come from the file name, and the names are not uniform

Course and chapter numbers are not in the banner, so the parser reads them from the file name. The real content has several generations of names, and the first version of the parser handled only the newest. Counting the course-chapter pages by pattern shows how much would have been wrong:

PatternExamplePagesCourse, chapter
Newesthungarian_basic_conversation_3_10.html2,9503, 10
Older prompt-prefixjs1-4.html, psp-croll-1-2.html4951, 4 and 1, 2
One trailing numbersetting_up_a_web_server_on_debian_03.html367none, 3
Lesson, chapter firsthtml_lesson_01_3.html623, 1 (the numbers are the other way round)
No numberlinux_appendix_a.html21 at first, 3 once the other forms were handlednone, none
The reversed names would have passed unnoticed
The HTML course's files, such as html_lesson_01_3.html, put the chapter first. A parser that always reads “course then chapter” gives every chapter of the advanced HTML course the number 3, and nothing complains. It was found only because the dry run checked that each course's chapters run 1 to N, and twelve of them came out as 3, 3, 3, 3… Check the numbers against a sequence, not just against “is it a number”.

The Dry Run

Before any import, Exercise 3 parses every real page and reports what it finds, writing nothing to the database:

ResultValue
Pages parsed4,402 of 4,402
By kind3,895 course chapters, 410 complete pages, 47 sidebar, 43 lessons, 7 with no banner
By siteprogramming 1,370; humanities 707; systems 653; languages 561; webdevelopment 514; lifeskills 253; ai 187; creative 157
Course folders395, with no folder using two different course names
Paths in no site0: the site map is complete
Chapter numbers not running 1 to N1 course (PHP Fundamentals, because its first chapter has no banner)
Banner File: line differs from the real file7 pages, all in Setting Up a Web Server on Debian
Seven pages with no banner
Four older lessons (two Hungarian and two Japanese), the hiragana and katakana tile fragments, and php1-1.html have no banner comment at all. They are real content problems, not parser problems, and php1-1.html is why PHP Fundamentals seems to start at chapter 2. Adding a banner to each is a small fix, listed in the review notes.

Hands-On Exercises

Exercise 1

Add the site lookup for a path, the banner parser and the page parser to the Chapter 2 project. Copy a real file of each banner shape into tests/fixtures, write tests for each shape and for every file-name pattern, and print what the parser finds in each sample.

๐Ÿ“„ View solution
Exercise 2

Create the Course and Page models and their migration, and test them against a real database: one row per path, chapters in numeric order, unsafe paths and mismatched sites rejected, and a course deleted without losing its pages. Read the SQL Django generates.

๐Ÿ“„ View solution
Exercise 3

Write a dry-run script that parses every page of the real content folder and reports the counts by kind and site, paths in no site, titles that fell back to the file name, banner file-name mismatches, and courses whose chapter numbers do not run 1 to N. Fix the parser for whatever it finds, then list the content problems that remain.

๐Ÿ“„ View solution

Chapter 3 Quick Reference

  • The files stay the source; the database holds what the site must query: pages by site, chapters in order, recent updates
  • A page's identity is its path (relative, forward slashes, .html); its address is the path without .html
  • Models: Course (a folder of chapters) and Page (path, site, kind, title, course, chapter number, dates)
  • Order by the chapter number; a text sort puts 10 before 2
  • Kinds: course chapter, sidebar, lesson, complete page, other; only fields that exist are filled in
  • Title order: banner Chapter:, <title>, first h1/h2, then the file name
  • File names have generations: _3_10, js1-4, a single number, _lesson_01_3 (chapter first), none
  • site_for_path() raises NoSiteError for a path in no site; clean() checks the stored site matches
  • Validation runs in full_clean(), so the importer must call it
  • Dry-run before importing: 4,402 pages parse, 0 unmapped, 7 pages have no banner