Loading Content
Learning Website with Django
Chapter 4 ยท Loading Content from Files
Chapter 3 taught the project what a page is. This chapter fills the database with the real thing: a management
command, import_content, that reads every file in the content folder, works out what each one is,
and creates, updates or skips it. It has to be safe to run over and over, fast enough to run on every deploy, and
honest about anything it cannot handle. Every number below comes from running it on the real content folder.
Where Does the Page Body Live?
The Chapter 3 models store facts about a page. The page's HTML has to be somewhere, and there are two honest choices:
| In the database | Read from the file on each request | |
|---|---|---|
| What is deployed | Code plus one database; the content folder need not be on the server | Code, the database and the whole content folder, kept in step |
| Consistency | A snapshot: metadata and body always come from the same import | A file edited after the import can disagree with its database row |
| Cost | Disk space, and care not to load bodies for lists | File reads, ideally cached |
| Speed of a page | One query | One query plus a file read |
The course stores the body in the database. The real sizes make it easy: the 4,403 pages total
about 92 MB of HTML (an average of 22 KB, the largest 0.77 MB, none over 1 MB), and the finished SQLite file
is about 99 MB. A build is then a snapshot: import, test, deploy the database with the code. The cost to manage
is loading bodies by accident, so Page.objects.light() defers the body field, and every list, menu
and sitemap uses it.
Storing a Fragment Safely
A stored fragment is HTML that will be sent to browsers without escaping, which is exactly the thing Django normally prevents. It is safe here for one reason: it came from your own files, through the importer, and nothing else can write to it. Six habits protect that:
- One place marks it safe.
Page.htmlis the onlymark_safein the project. A template that wants the body uses that property; everything else is escaped as usual. - Never accept visitor text into
fragment. When Chapter 10 adds accounts, comments go in their own model and are escaped. - Refuse paths that leave the folder.
safe_joinresolves the real path, including symbolic links, and rejects anything outside the content root. - Set limits. A file over 2 MB is an error; the real maximum is 0.77 MB, so the limit only catches accidents.
- Be strict about encoding. A file that is not valid UTF-8 is reported, not guessed at.
- Write in one transaction. A failure part-way leaves the database exactly as it was.
<script>, and every page has inline styles. That means a
strict Content Security Policy that forbids inline code cannot be switched on without rewriting the content. The
protection is where the HTML comes from, not a header, which is why that point is worth repeating: only your
files, only through the importer.
Turning a File Into a Fragment
Three small, tested functions do the work, matching what the live site does (Learning Website: Framework & Architecture 6):
| Step | What it does |
|---|---|
extract_fragment | A complete document gives its <body> plus the styles from its head. A fragment loses only its leading banner comment. A body with no end is an error. |
rewrite_solution_links | Every .txt link is rebuilt as <course>/solutions/<file name>, whatever stale form it had (five are tested). |
safe_join | Resolves a path under the content root and refuses to leave it. |
The Importer
| Situation | What the importer does |
|---|---|
| A new file | Creates the page, and its course if it is a new folder |
| An unchanged file | Skips it (same SHA-256 hash): no parsing, no write |
| An edited file | Updates that page only; the other pages keep their import time |
| A file that has gone | Reports it as missing and keeps the row; --prune deletes it |
| Not UTF-8, over 2 MB, in no site, or a body with no end | An error for that file only; the rest import; the command exits with an error code |
_print pages, underscore files, pdfs, solutions, the kanji archive | Skipped, as on the live site |
--dry-run | Reports everything; writes nothing |
Deleting by default would be dangerous: a mistake in the content path would look like “every page was deleted” and empty the site. So a missing file is a report, and removal is a deliberate flag.
The Real Run
| Command | Result |
|---|---|
--dry-run | Would create 4,403 pages, 0 errors (13 to 43 seconds) |
| First import | Created 4,403 pages in 395 courses, 0 errors, about 20 seconds |
| Second import, nothing changed | 4,403 unchanged, about 10 to 12 seconds (almost all of it reading files to hash them) |
--force | Updated all 4,403, about 23 to 27 seconds |
--site languages | 561 pages, 1 to 5 seconds |
The stored result: programming 1,371 pages (33 MB of HTML), humanities 707 (9.5 MB), systems 653 (19 MB), languages 561 (7.3 MB), web development 514 (14 MB), life skills 253 (3.2 MB), AI 187 (3.9 MB) and creative 157 (2.0 MB). No stored page still begins with its banner comment, and none lacks a hash.
--site languages took 16 seconds, because it parsed every file and only then
checked the site. The site is known from the path alone, so the fixed version checks that first and never
reads the other sites' files: 1 to 5 seconds. The same command also showed that a timing is not a fact: the
dry run took 13 seconds on one run and 43 on another, depending on how many files were already cached. Repeat a
measurement before trusting it.
Slugs: The Path Is the Slug
A slug is the short, URL-safe name of a page. Here there is no separate slug field: the page's path
is its slug, and its address is the path without .html. That keeps the address identical to
the old site's, which makes redirects trivial (Learning Website: Framework & Architecture 11). Three checks on the
real paths made sure that is safe:
- Letter case: no two paths differ only by case, so the database's unique path cannot collide when the site moves from Windows to Linux.
- Unicode form: all paths are in the NFC form, so the same character cannot be stored two ways.
- Non-ASCII names: 282 paths contain non-ASCII characters (the hiragana and katakana pages). A test sends both the percent-encoded and the raw address and gets the page. Links should use the percent-encoded form, which is what browsers send.
Hands-On Exercises
Write the functions that turn a raw file into a stored fragment: unwrap a complete document, drop the banner, rewrite every stale solution link, and refuse unsafe paths. Test each, including five stale link forms and a path that climbs out of the folder. Then test addresses that contain non-ASCII characters.
๐ View solutionAdd the hash, fragment and import-time fields to the page model, then write the importer and the import_content command with --dry-run, --prune, --force and --site. Test it against a temporary content folder: creating, skipping unchanged files, updating an edited file, a deleted file, bad encoding, an oversized file, and a path in no site.
Run the importer on the real content folder: a dry run, the first import, a repeat, a forced import and a single site. Record the counts and times, look at what was stored, and find and fix the slowest avoidable step.
๐ View solutionChapter 4 Quick Reference
- The body lives in the database (92 MB for 4,403 pages; the database file is about 99 MB); use
Page.objects.light()for every list - A stored fragment is trusted HTML: one
mark_safe(Page.html), only your own files, never visitor text - Protect the import:
safe_join, a 2 MB limit, strict UTF-8, one transaction - Change detection is a SHA-256 of the file: unchanged files are skipped without parsing
- A missing file is reported and kept;
--prunedeletes;--dry-runwrites nothing;--forcere-imports all;--sitelimits to one site - One bad file is an error for that file only; the command exits with an error code so a build fails
- The path is the slug and the address is the path without
.html; no case collisions, all NFC, 282 non-ASCII paths work - Real run: 4,403 pages, 395 courses, 0 errors, about 20 seconds; repeat runs about 10 seconds
- Check the cheap test (the site, from the path) before the expensive one (reading and parsing the file)