learning-website-django1-7 Exercise 3: Run It on the Real Content ================================================================ A change to how pages are PROCESSED does not change the files, so the importer's hash of each file (Chapter 4) would say "unchanged" and skip every page: the new markup would only reach pages that were edited. The fix is a pipeline version that is part of the hash. In apps/content/importer.py: PIPELINE_VERSION = "2" ... digest = hashlib.sha256(PIPELINE_VERSION.encode() + b"\0" + data).hexdigest() Raise it whenever a step in prepare_fragment, or a site transform, changes. A test (PipelineVersionTests in tests/test_languages.py) imports a page, changes the version, and checks that the page is updated once and is then unchanged again. Import the real content (the version had just been raised): python manage.py import_content created 3, updated 4403, unchanged 0, errors 0 (41.9s) python manage.py import_content created 0, updated 0, unchanged 4406, errors 0 (18.5s) (The "created 3" are three new course chapters added since the last import. Both runs were measured on a folder synced by OneDrive and vary a lot between runs.) Then look at what was stored. Save as language_stats.py in the project folder: """language_stats.py: what the languages pipeline did to the real, imported pages. Run from the project folder.""" import os import re os.environ.setdefault("DJANGO_SETTINGS_MODULE", "config.settings.dev") import django django.setup() from collections import Counter from django.core.cache import cache from django.test import Client from apps.content.models import Page lang_attr = re.compile(r'<[a-zA-Z0-9]+ lang="([a-z-]+)" class="([^"]*)"') by_code, pages_with, other_sites = Counter(), Counter(), 0 for site, frag in Page.objects.values_list("site", "fragment"): found = lang_attr.findall(frag) if site == "languages": for code, cls in found: by_code[code] += 1 for code in {c for c, _ in found}: pages_with[code] += 1 else: other_sites += len(found) print("lang attributes added on the languages site, by language code:") for code, n in by_code.most_common(): print(f" lang=\"{code}\" {n:>5} lines on {pages_with[code]} pages") print("lang attributes of this kind on every other site:", other_sites) cache.clear() html = Client(HTTP_HOST="languages.localhost").get("/").content.decode() print("\nthe languages front page:") for section in re.findall(r'
', html, re.S): name = re.search(r'lang-section__name">([^<]+)', section)[1] groups = re.findall(r'lang-group">([^<]+)', section) courses = re.findall(r'
  • ([^<]+) (\d+) chapters', section) lessons = re.search(r'Lessons and reference (\d+) pages', section) print(f" {name}: {len(courses)} courses; groups: {', '.join(groups) or 'none'}" + (f"; lessons and reference: {lessons[1]} pages" if lessons else "")) python language_stats.py lang attributes added on the languages site, by language code: lang="de" 1285 lines on 48 pages lang="hu" 1274 lines on 63 pages lang="ja" 687 lines on 35 pages lang="fr" 363 lines on 15 pages lang attributes of this kind on every other site: 0 the languages front page: French: 1 courses; groups: Survival French; lessons and reference: 3 pages German: 4 courses; groups: Survival German, Everyday German Hungarian: 4 courses; groups: Reading and writing, Survival Hungarian; lessons and reference: 12 pages Japanese: 11 courses; groups: Reading and writing, Survival Japanese, Culture; lessons and reference: 300 pages What it shows ------------- - The lang attributes added to the real pages match the counts of the classes exactly: 1,285 German lines, 1,274 Hungarian, 687 Japanese (as ja) and 363 French. Nothing was missed and nothing was added to any other site. - The front page groups the real courses as designed. German has Survival German (courses 1 to 3) and Everyday German (course 4); Hungarian has Reading and writing (the alphabet course) and Survival Hungarian; Japanese has the writing courses (hiragana and katakana), Survival Japanese, and the five culture courses. - 196 pages keep their answers inside a native details element, which needs no script and works here unchanged. Looked at: a headless Chrome screenshot of the front page showed each language's card with its own colour, the group titles in capitals, the courses with chapter counts, and the lessons link where lessons exist. (Desktop width only.) Not checked: the Japanese culture courses use their own lesson wrapper classes, so they appear in the site's accent colour, not a language colour; the chapters' PDFs and solution files are Chapter 8. WHY THIS WORKS AS AN ANSWER --------------------------- The results are checked against counts taken from the pages before the code was written, so a mismatch would have shown at once. The pipeline version solved a problem that only appears AFTER the first import: processing that changes with no file changing.