learning-website-framework1-10 Exercise 1: Inventory the Links Around One Area ================================================================================ Before moving an area, measure how tangled it is with the rest. For the languages area, count: links from its pages to its own pages, to other areas, to special routes and to resource files, and links INTO it from other areas. Also check that every target exists. Needs sitemap.py (Chapter 5). Save as link_inventory.py: import os, re, sys from collections import Counter, defaultdict from urllib.parse import urlsplit, unquote from sitemap import site_for, SPECIAL_PREFIXES HREF = re.compile(r'href="(/[^"#?]*)', re.I) SKIP = {"pdfs", "solutions"} # pages that are real routes of the site but are not files in content/ (see Chapter 5) KNOWN_ROUTES = {"/", "/courses/", "/japan/kanji-tiles/"} def pages(content): for dirpath, dirnames, filenames in os.walk(content): dirnames[:] = [d for d in dirnames if d not in SKIP] for n in filenames: if n.endswith(".html") and not n.endswith("_print.html") and not n.startswith("_"): yield os.path.join(dirpath, n) def exists(content, url_path): rel = unquote(url_path).strip("/") p = os.path.join(content, rel) return os.path.isdir(p) or os.path.isfile(p + ".html") or os.path.isfile(p) def classify(content, href): if href.startswith("//"): return "protocol-relative" path = href if path in KNOWN_ROUTES: return "special route" try: target = site_for(path) except (KeyError, IndexError): return "not a content path" if path.strip("/").split("/")[0] in ("resources",): return f"resource:{target}" return target if exists(content, path) else f"broken:{target}" if __name__ == "__main__": content, area = sys.argv[1], sys.argv[2] outbound = Counter(); inbound = Counter(); broken = [] n_area = 0 for p in pages(content): rel = os.path.relpath(p, content).replace(os.sep, "/") here = site_for("/" + rel) text = open(p, encoding="utf-8", errors="ignore").read() links = set(HREF.findall(text)) if here == area: n_area += 1 for href in links: c = classify(content, href) outbound["same site" if c == area else c] += 1 if c.startswith("broken") and len(broken) < 6: broken.append((rel, href)) else: for href in links: try: if site_for(href) == area: inbound[here] += 1 except (KeyError, IndexError): pass print(f"area: {area} pages scanned: {n_area}") print("links FROM this area (distinct per page):") for k, v in outbound.most_common(): print(f" {k:<28}{v:>6}") print("links INTO this area from other sites:") for k, v in inbound.most_common(): print(f" from {k:<22}{v:>6}") print(f" total {sum(inbound.values())}") if broken: print("examples of links to a missing target:") for rel, href in broken: print(f" {rel} -> {href}") Run it: python link_inventory.py "/claude-projects/website-content/content" languages Output (checked by running it): area: languages pages scanned: 593 links FROM this area (distinct per page): broken:languages 282 same site 225 special route 30 resource:languages 23 links INTO this area from other sites: total 0 examples of links to a missing target: japan/japanese-language/reference-materials/hiragana/hiragana_あ.html -> /japan/hiragana/hiragana-tiles japan/japanese-language/reference-materials/hiragana/hiragana_い.html -> /japan/hiragana/hiragana-tiles japan/japanese-language/reference-materials/hiragana/hiragana_う.html -> /japan/hiragana/hiragana-tiles japan/japanese-language/reference-materials/hiragana/hiragana_うぁ.html -> /japan/hiragana/hiragana-tiles japan/japanese-language/reference-materials/hiragana/hiragana_うぃ.html -> /japan/hiragana/hiragana-tiles japan/japanese-language/reference-materials/hiragana/hiragana_うぇ.html -> /japan/hiragana/hiragana-tiles What the numbers say -------------------- - 593 pages scanned (the exercise counts distinct links per page). - ZERO links from the languages area to any other area, and ZERO links into it from other areas. It is a clean cut: nothing else needs rewriting for it to move. That is the best reason to migrate it first. - 225 links stay inside the area (they stay relative after the move). - 23 links go to resource files under /resources/japanese/ (these assets must move with the area, or the links must be rewritten). - 30 links go to the kanji tiles page, a special route built by its own page file, not a file in content/. The scanner has a short list of such routes. - 282 links point at a page that does not exist: 141 hiragana pages link to /japan/hiragana/hiragana-tiles and 141 katakana pages link to /japan/katakana/katakana-tiles. The folders hold hiragana_tiles_cms_content.html and katakana_tiles_cms_content.html, but only the kanji tiles file has a page that routes it. These back-links are dead on the current site (checked against the built output). Fix them before migrating, so you do not move a bug. WHY THIS WORKS AS AN ANSWER --------------------------- A migration plan should start from measurements. This one shows that the first area is a clean cut, lists the exceptions that need decisions, and found a real bug that would otherwise be copied to the new site.