PDFs and Static Files

Learning Website with Django

Chapter 8 · PDFs, Solutions & Static Files

A page is not the only thing a visitor downloads. Every course has a PDF, most chapters link to solution files, and every page loads a stylesheet and a script. This chapter sorts those files into three kinds, decides who serves each one, and keeps the addresses that old pages already link to working after the split. It also fixes a small problem with the copy button on code blocks that only shows up when the clipboard is unavailable.

Run for real, on Django 6.1 and WhiteNoise 6.12
The project has 180 tests (all passing). Real PDFs and solution files were mirrored from the content folder and fetched with curl from the development server, and the copy script was loaded in headless Chrome. The Apache configuration could not be run here and is marked as such.

Three Kinds of File

KindExamplesSizeServed by
Page bodiesThe 4,403 pages92 MB in the databaseDjango, from the database (Chapter 4)
Static filesThe theme's three stylesheets and copy-code.jsA few kilobytesWhiteNoise, inside the Django process
AssetsCourse PDFs, solution .txt files8,394 files, 666 MBApache in production (Django only in development)

The split matters because the three kinds want different things. Static files are tiny, shared by every site and can be cached for years. Assets are large, belong to one site, and benefit from a real web server's handling of big files. Pages are the only thing that needs Django's templates.

Assets: One Folder per Site

Pages link to assets at the same relative path as in the content folder: the importer rewrote every stale solution link to /<course>/solutions/<file>. To keep those addresses working, the collect_assets command mirrors every pdfs/ and solutions/ file into ASSET_ROOT/<site>/<the same path>:

assets/ ├── languages/ │ └── hungary/hungarian-basic-3/ │ ├── pdfs/Hungarian_Basic_Conversation_3_Course.pdf │ └── solutions/... └── systems/ └── linux/system-administration/debian-development-machine-setup/ ├── pdfs/... └── solutions/devsetup1-1_exercise1.txt

On the real content, a dry run for every site finds 8,394 files (666 MB), all of which belong to a site: nothing was skipped. Languages has 236 files (195 MB) and systems 1,419 (90 MB). A repeat run copies nothing and takes 0.3 seconds, because a file is copied only when it is missing or its size or modified time differs. A file in a folder that belongs to no site is reported, not guessed. Removal needs --prune.

The site boundary is a folder
Because each site has its own folder, a site cannot serve another site's files: there is nothing there to serve. The same PDF requested on the systems site gets a 404. This holds whether Django or Apache serves the file, because both read the per-site folder.

Serving Assets: Django in Development, Apache in Production

Django's django.views.static.serve can send the files, and its documentation is blunt that the view is “not hardened for production use”. So it is used only while SERVE_ASSETS_WITH_DJANGO is on (development), with a route that only matches .pdf and .txt files inside a pdfs or solutions folder of that site. In production the route does not exist, and Apache serves the same folder directly.

Real requests to the development server, with curl:

RequestResult
The Hungarian course PDF, on the languages site200, application/pdf, 3.9 MB, with a Last-Modified date
The same PDF, on the systems site404
A Debian chapter's solution link200, text/plain, X-Content-Type-Options: nosniff
The same file again, with If-Modified-Since304: the browser keeps its copy
The first 100 bytes of the PDF200 with the whole file: Django does not honour range requests
  • Solutions are plain text with nosniff, so a browser may not decide a .txt file is something it should run.
  • Range requests are why Apache is better. A PDF viewer asks for pieces of a large file so the first page appears quickly; a 206 reply is what Apache sends and what Django's view cannot.
  • Tests cover the dangerous cases: a path that climbs out with .. or %2e%2e is refused, a missing file and a directory are 404, a .sh file in a solutions folder is not routed, and the link the importer wrote reaches the file.
  • The course page lists its downloads: the PDF as a download link, and the solution files in a collapsible section when there are any (the language courses have none).
The Apache configuration was not run
The production configuration (an AliasMatch to the site's own asset folder, a ProxyPassMatch exclusion so the proxy does not claim those addresses first, and a cache header for PDFs) is in the Exercise 2 solution. Apache is not installed on the machine used to write this course, so test it with apache2ctl configtest and curl -I: a range request should return 206 from Apache.

Static Files with WhiteNoise

The theme's files are a handful of small static files. WhiteNoise serves them from inside the Django process, which needs no web-server configuration at all. Two settings do the work: the middleware, placed directly after SecurityMiddleware, and a storage class that gives every file a hashed name:

# config/settings/prod.py STORAGES = { "default": {"BACKEND": "django.core.files.storage.FileSystemStorage"}, "staticfiles": {"BACKEND": "whitenoise.storage.CompressedManifestStaticFilesStorage"}, } # then, before deploying: python manage.py collectstatic

The tests read the real behaviour instead of assuming it:

  • After collectstatic each file has a name with a 12-character hash (tokens.5f6ee8c98574.css) and a compressed .gz copy, and the pages link to the hashed names.
  • A hashed file is sent with Cache-Control: max-age=315360000, public, immutable: ten years, never revalidated. A first guess in the test, one year, was wrong, and the test failed and said so. The unhashed name is cached for 60 seconds only.
  • Static files are sent with Access-Control-Allow-Origin: *, which suits a stylesheet shared between sites (Learning Website: Framework & Architecture 1), and means anything under /static/ can be read by any web page.
  • A file missing from the manifest is an error in production, so a typo in a template shows up at once.

One small thing from the first full run: WhiteNoise printed a warning on every start because the folder collectstatic fills did not exist in development. The development settings now create it.

The Copy Button, Revisited

The first version used the modern clipboard and did nothing if it was unavailable. It is only available on HTTPS and on localhost, so the development server opened from another machine over plain HTTP would fail without a word. The new version falls back to a hidden text box and the older copy command, and shows “Copy failed” if both fail, so the button never lies.

It was tested in headless Chrome with a page that inserts a body as text, as the site does, with a script in it, and then clicks the button:

CheckResult
A script inserted as text: has it run before the page loads?No (as expected)
And after the page loads and copy-code.js has re-created it?Yes: the interactive tools will work
Click the button (normal context)“Copy failed”
Click the button (forced insecure context, so the fallback runs)“Copy failed”
What this does and does not prove
Script re-creation is verified. The copy itself could not be verified: in headless Chrome the clipboard is not permitted and the older command also fails, so both paths report failure. That does prove the fallback is reached and that the button says so instead of staying silent. Whether copying succeeds has to be tried by hand in a real browser, on an HTTPS address and on a plain HTTP one. It is in the review notes.

Hands-On Exercises

Exercise 1

Write the command that mirrors every pdfs/ and solutions/ file into one folder per site at the same relative path, copying only what is missing or changed, with a dry run, a one-site option and an opt-in prune. Run it on the real content and report the counts.

📄 View solution
Exercise 2

Serve the mirrored files from the development server through a route that only matches PDFs and text files in pdfs and solutions folders of that site. Test the boundaries and the dangerous paths, fetch real files with curl, and write the Apache configuration for production, saying clearly that it is untested.

📄 View solution
Exercise 3

Serve the theme's static files with WhiteNoise and hashed names, and test the real cache headers. Then make the copy button fall back when the clipboard is unavailable, and test the script in headless Chrome. State exactly which parts the browser test can and cannot verify.

📄 View solution

Chapter 8 Quick Reference

  • Three kinds of file: page bodies (database), static files (WhiteNoise), assets (PDFs and solutions: Apache in production)
  • Assets keep the same relative path; collect_assets mirrors them to ASSET_ROOT/<site>/...: 8,394 files, 666 MB, 0 skipped
  • The site boundary is a folder: another site's PDF is a 404 because it is not there
  • Django's static.serve is development only (its docs say so); it does not honour range requests
  • Solutions are served as text/plain with nosniff; a repeat request with If-Modified-Since gets 304
  • Only .pdf and .txt in pdfs or solutions folders are routed; .. and %2e%2e are refused
  • WhiteNoise: middleware right after SecurityMiddleware; CompressedManifestStaticFilesStorage; run collectstatic before deploying
  • Hashed static files: max-age=315360000, public, immutable (ten years); unhashed: 60 seconds; both with Access-Control-Allow-Origin: *
  • Copy button: modern clipboard, then a fallback, then “Copy failed”; script re-creation verified in Chrome; the copy itself needs a manual test
  • The Apache configuration is untested: apache2ctl configtest and curl -I (expect 206 for a range request)