Loading Content

Learning Website with Django

Chapter 4 ยท Loading Content from Files

Chapter 3 taught the project what a page is. This chapter fills the database with the real thing: a management command, import_content, that reads every file in the content folder, works out what each one is, and creates, updates or skips it. It has to be safe to run over and over, fast enough to run on every deploy, and honest about anything it cannot handle. Every number below comes from running it on the real content folder.

Run for real, on Django 6.1
The importer has 79 tests in the project (all passing, with a temporary content folder for each importer test), and was run on all 4,403 real pages. The timings depend on a folder synced by OneDrive on an ordinary PC, so treat them as relative, not as promises.

Where Does the Page Body Live?

The Chapter 3 models store facts about a page. The page's HTML has to be somewhere, and there are two honest choices:

In the databaseRead from the file on each request
What is deployedCode plus one database; the content folder need not be on the serverCode, the database and the whole content folder, kept in step
ConsistencyA snapshot: metadata and body always come from the same importA file edited after the import can disagree with its database row
CostDisk space, and care not to load bodies for listsFile reads, ideally cached
Speed of a pageOne queryOne query plus a file read

The course stores the body in the database. The real sizes make it easy: the 4,403 pages total about 92 MB of HTML (an average of 22 KB, the largest 0.77 MB, none over 1 MB), and the finished SQLite file is about 99 MB. A build is then a snapshot: import, test, deploy the database with the code. The cost to manage is loading bodies by accident, so Page.objects.light() defers the body field, and every list, menu and sitemap uses it.

Storing a Fragment Safely

A stored fragment is HTML that will be sent to browsers without escaping, which is exactly the thing Django normally prevents. It is safe here for one reason: it came from your own files, through the importer, and nothing else can write to it. Six habits protect that:

  • One place marks it safe. Page.html is the only mark_safe in the project. A template that wants the body uses that property; everything else is escaped as usual.
  • Never accept visitor text into fragment. When Chapter 10 adds accounts, comments go in their own model and are escaped.
  • Refuse paths that leave the folder. safe_join resolves the real path, including symbolic links, and rejects anything outside the content root.
  • Set limits. A file over 2 MB is an error; the real maximum is 0.77 MB, so the limit only catches accidents.
  • Be strict about encoding. A file that is not valid UTF-8 is reported, not guessed at.
  • Write in one transaction. A failure part-way leaves the database exactly as it was.
The pages contain scripts on purpose
Interactive tools carry their own <script>, and every page has inline styles. That means a strict Content Security Policy that forbids inline code cannot be switched on without rewriting the content. The protection is where the HTML comes from, not a header, which is why that point is worth repeating: only your files, only through the importer.

Turning a File Into a Fragment

Three small, tested functions do the work, matching what the live site does (Learning Website: Framework & Architecture 6):

StepWhat it does
extract_fragmentA complete document gives its <body> plus the styles from its head. A fragment loses only its leading banner comment. A body with no end is an error.
rewrite_solution_linksEvery .txt link is rebuilt as <course>/solutions/<file name>, whatever stale form it had (five are tested).
safe_joinResolves a path under the content root and refuses to leave it.

The Importer

# python manage.py import_content [--dry-run] [--prune] [--force] [--site languages] for rel in iter_content_paths(root): data = open(full, "rb").read() digest = hashlib.sha256(data).hexdigest() if not force and existing.get(rel) == digest: unchanged += 1; continue # nothing to do parsed = parse_page(raw, rel) # Chapter 3 fragment = prepare_fragment(raw, rel) # this chapter ...create or update... # bulk_create / bulk_update, one transaction
SituationWhat the importer does
A new fileCreates the page, and its course if it is a new folder
An unchanged fileSkips it (same SHA-256 hash): no parsing, no write
An edited fileUpdates that page only; the other pages keep their import time
A file that has goneReports it as missing and keeps the row; --prune deletes it
Not UTF-8, over 2 MB, in no site, or a body with no endAn error for that file only; the rest import; the command exits with an error code
_print pages, underscore files, pdfs, solutions, the kanji archiveSkipped, as on the live site
--dry-runReports everything; writes nothing

Deleting by default would be dangerous: a mistake in the content path would look like “every page was deleted” and empty the site. So a missing file is a report, and removal is a deliberate flag.

The Real Run

CommandResult
--dry-runWould create 4,403 pages, 0 errors (13 to 43 seconds)
First importCreated 4,403 pages in 395 courses, 0 errors, about 20 seconds
Second import, nothing changed4,403 unchanged, about 10 to 12 seconds (almost all of it reading files to hash them)
--forceUpdated all 4,403, about 23 to 27 seconds
--site languages561 pages, 1 to 5 seconds

The stored result: programming 1,371 pages (33 MB of HTML), humanities 707 (9.5 MB), systems 653 (19 MB), languages 561 (7.3 MB), web development 514 (14 MB), life skills 253 (3.2 MB), AI 187 (3.9 MB) and creative 157 (2.0 MB). No stored page still begins with its banner comment, and none lacks a hash.

Measure, then move the cheap test first
The first version of --site languages took 16 seconds, because it parsed every file and only then checked the site. The site is known from the path alone, so the fixed version checks that first and never reads the other sites' files: 1 to 5 seconds. The same command also showed that a timing is not a fact: the dry run took 13 seconds on one run and 43 on another, depending on how many files were already cached. Repeat a measurement before trusting it.

Slugs: The Path Is the Slug

A slug is the short, URL-safe name of a page. Here there is no separate slug field: the page's path is its slug, and its address is the path without .html. That keeps the address identical to the old site's, which makes redirects trivial (Learning Website: Framework & Architecture 11). Three checks on the real paths made sure that is safe:

  • Letter case: no two paths differ only by case, so the database's unique path cannot collide when the site moves from Windows to Linux.
  • Unicode form: all paths are in the NFC form, so the same character cannot be stored two ways.
  • Non-ASCII names: 282 paths contain non-ASCII characters (the hiragana and katakana pages). A test sends both the percent-encoded and the raw address and gets the page. Links should use the percent-encoded form, which is what browsers send.

Hands-On Exercises

Exercise 1

Write the functions that turn a raw file into a stored fragment: unwrap a complete document, drop the banner, rewrite every stale solution link, and refuse unsafe paths. Test each, including five stale link forms and a path that climbs out of the folder. Then test addresses that contain non-ASCII characters.

๐Ÿ“„ View solution
Exercise 2

Add the hash, fragment and import-time fields to the page model, then write the importer and the import_content command with --dry-run, --prune, --force and --site. Test it against a temporary content folder: creating, skipping unchanged files, updating an edited file, a deleted file, bad encoding, an oversized file, and a path in no site.

๐Ÿ“„ View solution
Exercise 3

Run the importer on the real content folder: a dry run, the first import, a repeat, a forced import and a single site. Record the counts and times, look at what was stored, and find and fix the slowest avoidable step.

๐Ÿ“„ View solution

Chapter 4 Quick Reference

  • The body lives in the database (92 MB for 4,403 pages; the database file is about 99 MB); use Page.objects.light() for every list
  • A stored fragment is trusted HTML: one mark_safe (Page.html), only your own files, never visitor text
  • Protect the import: safe_join, a 2 MB limit, strict UTF-8, one transaction
  • Change detection is a SHA-256 of the file: unchanged files are skipped without parsing
  • A missing file is reported and kept; --prune deletes; --dry-run writes nothing; --force re-imports all; --site limits to one site
  • One bad file is an error for that file only; the command exits with an error code so a build fails
  • The path is the slug and the address is the path without .html; no case collisions, all NFC, 282 non-ASCII paths work
  • Real run: 4,403 pages, 395 courses, 0 errors, about 20 seconds; repeat runs about 10 seconds
  • Check the cheap test (the site, from the path) before the expensive one (reading and parsing the file)