learning-website-framework1-11 Exercise 3: The Complete Redirect Map ========================================================================== Generate the Apache redirect rules for ALL sites from the site map, in the right order, and test every old URL against them: each must match exactly one destination site, the right one. Rule order matters: specific rules (resource folders, sidebar subjects) come BEFORE the general folder rules. Needs sitemap.py (Chapter 5) and sitemaps.py (Chapter 7). Save as redirect_map.py: import re, sys from collections import Counter from sitemap import SITES, SIDEBAR_ROUTES, SPECIAL_PREFIXES, site_for from sitemaps import collect DOMAIN = "osztromok.com" def build_rules(): """Ordered list of (pattern, site). Specific rules come before general ones.""" rules = [] for prefix, site in SPECIAL_PREFIXES.items(): # resources/japanese -> languages rules.append((rf"^/{re.escape(prefix)}(/.*)?$", site, "special resource prefix")) rules.append((r"^/resources/hungarian(/.*)?$", "languages", "special resource prefix")) for subject, site in sorted(SIDEBAR_ROUTES.items()): # sidebar/ -> that subject's site rules.append((rf"^/sidebar/{re.escape(subject)}(/.*)?$", site, "sidebar subject")) for site, folders in SITES.items(): # ordinary folders alt = "|".join(re.escape(f) for f in folders) rules.append((rf"^/({alt})(/.*)?$", site, "folders")) return rules def apache(rules): lines = [] for pattern, site, why in rules: lines.append(f"# {why} -> {site}") lines.append(f"RedirectMatch 301 {pattern} https://{site}.{DOMAIN}$0") return "\n".join(lines) def first_match(rules, path): for pattern, site, why in rules: if re.match(pattern, path): return site return None def all_matches(rules, path): return [site for pattern, site, why in rules if re.match(pattern, path)] if __name__ == "__main__": rules = build_rules() text = apache(rules) print(f"{len(rules)} rules, {len(text.splitlines())} lines. The first ten lines:") print("\n".join(text.splitlines()[:10])) print("...") urls = sorted(collect(sys.argv[1])) urls = [u for u in urls if u != "/sidebar/"] # the sidebar index is handled separately ok = unmatched = disagree = multi = 0 extra = ["/resources/japanese/kanji/kanji_%E4%B8%8A.html", "/resources/hungarian/some-file.pdf"] for u in urls + extra: matches = all_matches(rules, u) first = matches[0] if matches else None if first is None: unmatched += 1 elif u in extra or first == site_for(u): ok += 1 else: disagree += 1 if len(set(matches)) > 1: multi += 1 print(f"\ntested {len(urls) + len(extra)} old URLs (every page, plus two resource files)") print(f" redirected to the right site: {ok}") print(f" matched by no rule: {unmatched}") print(f" sent to the wrong site: {disagree}") print(f" matched by rules for DIFFERENT sites (order matters): {multi}") print(" /sidebar/ itself:", first_match(rules, "/sidebar/"), "(not redirected: decide where the old index goes)") Run it: python redirect_map.py "/claude-projects/website-content/content" Output (checked by running it): 17 rules, 34 lines. The first ten lines: # special resource prefix -> languages RedirectMatch 301 ^/resources/japanese(/.*)?$ https://languages.osztromok.com$0 # special resource prefix -> languages RedirectMatch 301 ^/resources/hungarian(/.*)?$ https://languages.osztromok.com$0 # sidebar subject -> ai RedirectMatch 301 ^/sidebar/ai(/.*)?$ https://ai.osztromok.com$0 # sidebar subject -> lifeskills RedirectMatch 301 ^/sidebar/drinks(/.*)?$ https://lifeskills.osztromok.com$0 # sidebar subject -> humanities RedirectMatch 301 ^/sidebar/football(/.*)?$ https://humanities.osztromok.com$0 ... tested 4926 old URLs (every page, plus two resource files) redirected to the right site: 4926 matched by no rule: 0 sent to the wrong site: 0 matched by rules for DIFFERENT sites (order matters): 0 /sidebar/ itself: None (not redirected: decide where the old index goes) What this proves, and what it does not -------------------------------------- - Proves: the generated rules cover every one of the 4,925 pages (plus the resource files), send each to the site the map says, and no URL can match rules for two different sites. - Does NOT prove: that Apache behaves the same as Python's regular expressions. This was simulated, not run on an Apache server. Before launch, test each rule on the real server with curl -I, and in particular: * the substitution $0 (the whole matched path) in RedirectMatch; * URLs with NON-ASCII characters, which are common here (the hiragana and katakana pages have names such as hiragana_あ.html). Apache matches the decoded path and the target must be re-encoded correctly. Test several. * the trailing slash, with and without it. - The old /sidebar/ index page matches no rule on purpose: the sidebar splits across six sites, so decide where the old index goes (the root landing page is a sensible choice). - This set replaces the single languages rule from Chapter 2, Exercise 3. - Use 302 while testing and 301 for the switch (Chapter 10). Keep the rules for at least a year. WHY THIS WORKS AS AN ANSWER --------------------------- The rules are generated from the same site map as everything else, so they cannot drift from the menus and sitemaps. Testing every real URL finds a gap or a conflict before a visitor does, and the notes say plainly where a script's test ends and a server test must begin.