learning-website-django1-11 Exercise 2: Redirects You Review Before They Exist ================================================================================ Old addresses that moved need a permanent (301) redirect, or every old bookmark and search result becomes a 404. Redirects are DECISIONS, not imported text, so they live in the database (editable in the admin) and come from a CSV that a person reviews. Save as apps/migration/models.py: from django.db import models from apps.content.models import SITE_CHOICES class Redirect(models.Model): """An old address that moved. Unlike pages, redirects are NOT imported from files: they are decisions, so they live in the database and are edited in the admin or loaded from a reviewed CSV file.""" site = models.CharField(max_length=40, choices=SITE_CHOICES) # the site that gets the old request old_path = models.CharField(max_length=500) # "hungary/hungarian-language/lesson_01" new_site = models.CharField(max_length=40, choices=SITE_CHOICES) new_path = models.CharField(max_length=500) # "hungary/hungarian-lessons/lesson_01" note = models.CharField(max_length=200, blank=True) created = models.DateTimeField(auto_now_add=True) class Meta: constraints = [models.UniqueConstraint(fields=["site", "old_path"], name="one_redirect_per_old_address")] def __str__(self): return f"{self.old_path} -> {self.new_path}" Save as apps/migration/redirects.py: """Finding where moved pages went, and answering for them.""" import csv import re from django.conf import settings from django.http import HttpResponsePermanentRedirect from apps.content.models import Page from .models import Redirect def normalise(slug): """Compare file names loosely: case and - versus _ are not differences that matter here.""" return slug.lower().replace("-", "_") def bare_path(path): """Any form of an address -> "hungary/x/y": no slashes at the ends, no index.html, no .html.""" path = path.strip("/") path = re.sub(r"(^|/)index\.html$", "", path) path = re.sub(r"\.html$", "", path) return path.strip("/") def redirect_target(request_site, bare): """The address to send an old request to, or None. A page that moved to ANOTHER site gets a full address.""" row = Redirect.objects.filter(site=request_site, old_path=bare).first() if row is None: return None path = "/" + row.new_path.strip("/") + "/" if row.new_site == request_site: return path return settings.SITE_URL_TEMPLATE.format(site=row.new_site) + path def redirect_response(request_site, bare): target = redirect_target(request_site, bare) return HttpResponsePermanentRedirect(target) if target else None def propose(missing): """For old page addresses that no longer exist, suggest where they went. missing: iterable of (site, old_path). Returns (proposals, ambiguous, nothing): proposals [(site, old_path, new_site, new_path, why)] exactly ONE page has that file name ambiguous [(site, old_path, [candidate paths])] several do: a person must choose nothing [(site, old_path)] no page has that file name Suggestions are only suggestions: a file name is not proof that it is the same lesson.""" by_slug = {} for page in Page.objects.light(): by_slug.setdefault(normalise(page.path[:-5].rsplit("/", 1)[-1]), []).append(page) proposals, ambiguous, nothing = [], [], [] for site, old in missing: slug = normalise(old.rsplit("/", 1)[-1]) found = by_slug.get(slug, []) why = "same file name" if not found and slug.endswith("_print"): # the old site had a printable twin of many pages; the page itself is the closest thing now found = by_slug.get(slug[: -len("_print")], []) why = "was the print version of this page" if len(found) == 1: page = found[0] proposals.append((site, old, page.site, page.path[:-5], why)) elif found: ambiguous.append((site, old, [p.path[:-5] for p in found])) else: nothing.append((site, old)) return proposals, ambiguous, nothing CSV_COLUMNS = ["site", "old_path", "new_site", "new_path", "why", "decision"] def write_csv(path, proposals): """decision is left EMPTY: the owner writes yes against each row they accept.""" with open(path, "w", newline="", encoding="utf-8") as handle: writer = csv.writer(handle) writer.writerow(CSV_COLUMNS) for row in proposals: writer.writerow([*row, ""]) def load_csv(path, *, accept_all=False): """Create a Redirect for every row marked yes. Returns (created, updated, skipped).""" created = updated = skipped = 0 with open(path, newline="", encoding="utf-8") as handle: for row in csv.DictReader(handle): if not accept_all and row.get("decision", "").strip().lower() != "yes": skipped += 1 continue _, was_created = Redirect.objects.update_or_create( site=row["site"], old_path=row["old_path"], defaults={"new_site": row["new_site"], "new_path": row["new_path"], "note": row.get("why", "")[:200]}) created += was_created updated += not was_created return created, updated, skipped The page view tries the redirect table only when there is no page and no folder listing at that address, so a real page can never be hidden by a redirect (a test checks this). In apps/content/views.py the end of the page view became: content = listing(site, bare) if content is None: moved = redirect_response(site, bare) if moved is not None: return moved raise Http404("No such page or folder") A redirect to a page on ANOTHER site is sent as a full address built from SITE_URL_TEMPLATE (for example http://systems.localhost:8000/linux/... in development). Save as apps/migration/management/commands/propose_redirects.py: from django.core.management.base import BaseCommand from apps.migration.oldsite import scan_old_site from apps.migration.redirects import propose, write_csv from apps.migration.verify import verify class Command(BaseCommand): help = "Suggest redirects for old page addresses that no longer work. Writes a CSV for you to review." def add_arguments(self, parser): parser.add_argument("old_root") parser.add_argument("--out", required=True, help="the CSV file to write") def handle(self, old_root, out, **options): pages = [u for u in scan_old_site(old_root) if u.kind == "page"] _, missing = verify(pages) proposals, ambiguous, nothing = propose([(site, path) for _, site, path in missing]) write_csv(out, proposals) self.stdout.write(f"{len(missing)} old pages do not work: {len(proposals)} have exactly one page with the same file name, " f"{len(ambiguous)} have several, {len(nothing)} have none.") for site, old, candidates in ambiguous[:10]: self.stdout.write(f" ambiguous: {old} -> {candidates[:3]}") for site, old in nothing[:10]: self.stdout.write(f" nothing: {old}") self.stdout.write(f"Suggestions written to {out}. Put 'yes' in the decision column of each row you accept, then run load_redirects.") Save as apps/migration/management/commands/load_redirects.py: from django.core.management.base import BaseCommand from apps.migration.redirects import load_csv class Command(BaseCommand): help = "Create redirects from a reviewed CSV: only rows whose decision column says yes." def add_arguments(self, parser): parser.add_argument("csv_file") parser.add_argument("--accept-all", action="store_true", help="load every row, reviewed or not (for tests)") def handle(self, csv_file, accept_all=False, **options): created, updated, skipped = load_csv(csv_file, accept_all=accept_all) self.stdout.write(f"{created} created, {updated} updated, {skipped} skipped (not marked yes)") The matching rule is deliberately simple and deliberately modest: if exactly ONE new page has the same file name (ignoring case and - versus _), propose it; if several do, list them for a person; if none do, say so; a name ending in _print falls back to the page without it. A file name is not proof that it is the same lesson, so every row starts with an EMPTY decision column and only rows where you write yes are loaded. Run on the real data: python manage.py propose_redirects "" --out proposed.csv 80 old pages do not work: 32 have exactly one page with the same file name, 0 have several, 48 have none. (Adding the _print rule afterwards turned 23 of those 48 into proposals: 55 proposals, 0 ambiguous, 25 with nothing.) Rows look like this: site,old_path,new_site,new_path,why,decision ai,ai/claude-tools/claude-lessons/japanese_lesson_dates,languages,japan/japanese-language/japanese-lessons/japanese_lesson_dates,same file name, languages,hungary/hungarian-language/hungarian_lesson_01,languages,hungary/hungarian-lessons/hungarian_lesson_01,same file name, systems,linux/system-administration/tentative-package-management/pm_lesson_01,systems,linux/system-administration/linux-package-managers/pm_lesson_01,same file name, languages,hungary/hungarian-language/hungarian_lesson_visiting_a_vineyard_print,languages,hungary/hungarian-lessons/hungarian_lesson_visiting_a_vineyard,was the print version of this page, Loading without any "yes" does nothing: "0 created, 0 updated, 32 skipped (not marked yes)". To measure what the mechanism achieves, I loaded ALL the proposals with --accept-all into my scratch database (a test shortcut: I looked at the rows above, but I have NOT reviewed all 55; you should before they go anywhere real), then ran the check again: pages owned by a site: 3182 80 missing, 3102 same (before) pages owned by a site: 3182 48 missing, 32 redirected, 3102 same pages owned by a site: 3182 25 missing, 55 redirected, 3102 same (with the _print rule) The 25 that still need a person (rows I could not match by name): ai/claude-tools/claude-lessons (folder index) ai/claude-tools/claude-lessons/japanese_where_is_the_cat (renamed to japanese_lesson_where_is_the_cat) ai/claude-tools/tentative-claude-projects + proj_lesson_01 ... 08 (course removed?) ai/tentative-freelance-ai-content-writing, linux/.../tentative-package-management (folder indexes) projects/tentative-web-stack-overview + ws_01 ... ws_07 (course removed?) security/tentative-web-security-course france/french-language, hungary/hungarian-language (folder indexes: the folders were renamed) japan/kanji-tiles (a dedicated page on the old site; see Exercise 3) For each, decide: redirect to the right new page or folder (add a row by hand or in the admin), or leave it a 404 (or 410 Gone) because the content was deliberately dropped. Tests for all of this are in tests/test_migration.py (shown in Exercise 3). What was not verified: redirects for PDF and solution files are not implemented (the 166 missing files in Exercise 1 would need them, or a decision that they are gone); a 301 is cached hard by browsers, so a wrong redirect is hard to undo: that is why they are reviewed first. WHY THIS WORKS AS AN ANSWER --------------------------- The tool does the boring 70% (it found 55 of 80 by file name) and refuses to guess the rest, and nothing becomes a permanent redirect without a person writing yes.