learning-website-django1-8 Exercise 2: Serve Assets Safely ============================================================= Django can serve a file for you. Its own documentation says that view is not hardened for production, so use it only in development, and let Apache serve the files in production from the same per-site folder. Either way a site only sees its own files. Add the view to apps/content/views.py: from pathlib import Path from django.conf import settings from django.views.static import serve def asset(request, site, asset_path): return serve(request, asset_path, document_root=Path(settings.ASSET_ROOT) / site) and a route to each site's URL configuration, ahead of the page route, only while SERVE_ASSETS_WITH_DJANGO is on (config/urlconfs/__init__.py): assets = [] if settings.SERVE_ASSETS_WITH_DJANGO: assets = [re_path(rf"^(?P(?:{alternatives})/(?:.+/)?(?:pdfs|solutions)/[^/]+\.(?:pdf|txt))$", views.asset, {"site": site}, name="asset")] module.urlpatterns = assets + [ ...the home and page routes... ] Only .pdf and .txt files in a pdfs or solutions folder are routed, and only for that site's own folders. The listing of a course folder shows its downloads. Add this to apps/theme/templates/theme/listing.html (the PDFs, then the solutions in a details element) and "assets": list_assets(settings.ASSET_ROOT, site, bare) to the listing's context: Save as apps/theme/templates/theme/listing.html: {% extends "theme/base.html" %} {% block title %}{{ title }}{% endblock %} {% block content %}

{{ title }}

{% if content.folders %}

Folders

{% endif %} {% if content.pages %}

{% if content.course %}Chapters{% else %}Pages{% endif %}

    {% for p in content.pages %}
  1. {{ p.title }}
  2. {% endfor %}
{% endif %} {% if assets.pdfs or assets.solutions %}

Downloads

{% if assets.pdfs %} {% endif %} {% if assets.solutions %}
Exercise solutions ({{ assets.solutions|length }} files)
{% endif %} {% endif %} {% endblock %} The second half of tests/test_assets.py (ServeAssetTests) is in Exercise 1's file. Run everything (Django 6.1.2): python manage.py test tests Found 180 test(s). System check identified no issues (0 silenced). Creating test database for alias 'default'... .................................................................................................................................................................................... ---------------------------------------------------------------------- Ran 180 tests in 2.607s OK Destroying test database for alias 'default'... Then try it for real: start the server and fetch real files with curl. Save as try_assets.py in the project folder: """try_assets.py: start the development server and fetch real PDFs, solution files and static files with curl. Run from the project folder with the project's Python.""" import os, re, socket, subprocess, sys, time PORT = 8765 proc = subprocess.Popen([sys.executable, "manage.py", "runserver", f"127.0.0.1:{PORT}", "--noreload"], stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL) for _ in range(60): with socket.socket() as s: if s.connect_ex(("127.0.0.1", PORT)) == 0: break time.sleep(0.25) def curl(host, path, *extra, body=False): args = ["curl.exe", "-s", "--resolve", f"{host}:{PORT}:127.0.0.1", *extra, f"http://{host}:{PORT}{path}"] if not body: args[1:1] = ["-D", "-", "-o", "NUL"] return subprocess.run(args, capture_output=True, text=True, encoding="utf-8", errors="replace").stdout def headers(host, path, *extra): text = curl(host, path, *extra) lines = [l.strip() for l in text.splitlines() if l.strip()] status = lines[0].split(" ", 2)[1] if lines else "?" pick = {l.split(":", 1)[0].lower(): l.split(":", 1)[1].strip() for l in lines[1:] if ":" in l} return status, pick try: pdf = "/hungary/hungarian-basic-3/pdfs/Hungarian_Basic_Conversation_3_Course.pdf" status, h = headers("languages.localhost", pdf) print(f"PDF on its own site {status} {h.get('content-type')} {int(h.get('content-length', 0)) / 1e6:.1f} MB Last-Modified: {'yes' if 'last-modified' in h else 'no'}") status, h = headers("languages.localhost", pdf, "-r", "0-99") print(f"PDF, first 100 bytes only {status} (a 206 would mean the range was honoured)") status, h = headers("systems.localhost", pdf) print(f"the same PDF on systems {status}") chapter = "/linux/system-administration/debian-development-machine-setup/debian_development_machine_setup_1_1/" page = curl("systems.localhost", chapter, body=True) links = re.findall(r'href="(/linux/[^"]+/solutions/[^"]+\.txt)"', page) print(f"solution links on the Debian chapter page: {len(links)}") status, h = headers("systems.localhost", links[0]) print(f"first solution link {status} {h.get('content-type')} nosniff: {h.get('x-content-type-options')}") status, h = headers("systems.localhost", links[0], "-H", f"If-Modified-Since: {h.get('last-modified')}") print(f"same file, repeat request {status} (304 means the browser's copy is still good)") status, h = headers("systems.localhost", "/static/theme/tokens.css") print(f"stylesheet {status} {h.get('content-type')} Cache-Control: {h.get('cache-control')}") listing = curl("languages.localhost", "/hungary/hungarian-basic-3/", body=True) print(f"course page: PDF link shown: {'download' in listing and '.pdf' in listing}; solutions section shown: {'Exercise solutions' in listing}") finally: proc.terminate() proc.wait(timeout=10) python try_assets.py PDF on its own site 200 application/pdf 3.9 MB Last-Modified: yes PDF, first 100 bytes only 200 (a 206 would mean the range was honoured) the same PDF on systems 404 solution links on the Debian chapter page: 3 first solution link 200 text/plain nosniff: nosniff same file, repeat request 304 (304 means the browser's copy is still good) stylesheet 200 text/css Cache-Control: None course page: PDF link shown: True; solutions section shown: False Reading the results ------------------- - A PDF is served as application/pdf with a Last-Modified date; asking for it on a DIFFERENT site gives 404, because that site's folder does not hold it. - The Debian chapter's three solution links (written by the importer) resolve to real files, served as text/plain with X-Content-Type-Options: nosniff, so a browser may not treat a .txt file as something executable. - A repeat request with If-Modified-Since gets 304, so the browser re-uses its copy. - The Hungarian course page shows its PDF link; its solutions section does not appear, because the language courses have none. - A request for the first 100 bytes of the PDF got the whole file (200, not 206). Django's view does not honour Range requests. A browser's PDF viewer uses ranges to show the first page of a large PDF quickly, which is one more reason Apache should serve them in production. - Tests also cover: a missing file and a directory are 404, a path that climbs out with .. or %2e%2e is refused, a .sh file in a solutions folder is not routed, and in production mode (the setting off) Django does not serve assets at all. PRODUCTION: Apache serves the files. THIS CONFIGURATION WAS NOT RUN (Apache is not installed on the machine used to write the course). Test it with apache2ctl configtest and curl -I on the server. In each site's *:443 virtual host, before the ProxyPass lines: # PDFs and solutions come from this site's own mirrored folder AliasMatch "^/(.+/(?:pdfs|solutions)/[^/]+\.(?:pdf|txt))$" "/var/www/languages/assets/$1" Require all granted Options -Indexes Header set Cache-Control "public, max-age=86400" ForceType "text/plain; charset=utf-8" # make sure the proxy does not claim those addresses first ProxyPassMatch "^/(.+/(?:pdfs|solutions)/[^/]+\.(?:pdf|txt))$" "!" Check that an address under assets is answered by Apache (the response has Apache's headers and a 206 for a range request) and not by Django. WHY THIS WORKS AS AN ANSWER --------------------------- The same folder layout serves development (Django) and production (Apache), so the boundary and the addresses are identical in both. The tests cover the dangerous cases (traversal, other sites, unexpected file types), and the real requests show the limits of the development view.