learning-website-framework1-11 Exercise 1: Audit Every Link in the Content =============================================================================== PART A: Census. Count every href and src attribute by kind. Save as link_census.py: import os, re, sys from collections import Counter from urllib.parse import urlsplit ATTR = re.compile(r'\b(href|src)="([^"]*)"', re.I) OWN = {"osztromok.com", "www.osztromok.com"} def kind(url): u = url.strip() if not u: return "empty" if u.startswith("#"): return "anchor only" if u.lower().startswith(("mailto:", "tel:", "javascript:", "data:")): return "mailto/tel/script/data" if u.startswith("//"): return "protocol-relative" parts = urlsplit(u) if parts.scheme in ("http", "https"): return "absolute, own domain" if (parts.hostname or "").lower() in OWN else "absolute, external" if u.startswith("/"): return "root-relative (/path)" return "relative (path, ../path)" def census(content): kinds, external, own_examples = Counter(), Counter(), [] files = 0 for dirpath, dirnames, filenames in os.walk(content): dirnames[:] = [d for d in dirnames if d not in ("pdfs", "solutions")] for n in filenames: if not n.endswith(".html"): continue files += 1 text = open(os.path.join(dirpath, n), encoding="utf-8", errors="ignore").read() for attr, url in ATTR.findall(text): k = kind(url) kinds[k] += 1 if k == "absolute, external": external[(urlsplit(url).hostname or "").lower().removeprefix("www.")] += 1 if k == "absolute, own domain" and len(own_examples) < 5: own_examples.append(url[:90]) return files, kinds, external, own_examples if __name__ == "__main__": files, kinds, external, own = census(sys.argv[1]) total = sum(kinds.values()) print(f"{files} HTML files, {total} href/src attributes") for k, v in kinds.most_common(): print(f" {k:<28}{v:>8} {100*v/total:5.1f}%") print("top external hosts:", ", ".join(f"{h} ({n})" for h, n in external.most_common(6))) if own: print("own-domain absolute examples:") for u in own: print(" " + u) Run it: python link_census.py "/claude-projects/website-content/content" Output (checked by running it): 4483 HTML files, 12358 href/src attributes root-relative (/path) 10516 85.1% relative (path, ../path) 1264 10.2% absolute, external 305 2.5% anchor only 271 2.2% empty 1 0.0% mailto/tel/script/data 1 0.0% top external hosts: code.claude.com (64), github.com (18), linkedin.com (17), youtube.com (9), reddit.com (6), mxtoolbox.com (5) PART B: Which links leave their own site? Resolve each link (root-relative, or relative to the page's own folder) and map it to a site. Set aside the solution (.txt) links, which the build rewrites from the file name (Chapter 6). Needs sitemap.py (Chapter 5). Save as cross_site.py: import os, posixpath, re, sys from collections import Counter, defaultdict from urllib.parse import urlsplit, unquote from sitemap import site_for ATTR = re.compile(r'\b(href|src)="([^"]*)"', re.I) KNOWN_ROUTES = {"/", "/courses/", "/japan/kanji-tiles/", "/admin/"} def resolve(page_rel, url): """Return the root-relative path a link points at, or None for external/anchor links.""" u = url.strip() if not u or u.startswith(("#", "//")) or u.lower().startswith(("mailto:", "tel:", "javascript:", "data:")): return None parts = urlsplit(u) if parts.scheme or parts.netloc: return None path = unquote(parts.path) if not path: return None if path.startswith("/"): return path folder = posixpath.dirname("/" + page_rel) return posixpath.normpath(posixpath.join(folder, path)) def target_site(path): if path in KNOWN_ROUTES: return "(special route)" try: return site_for(path) except (KeyError, IndexError): top = path.strip("/").split("/")[0] return f"(unmapped: /{top}/)" if top == "resources" else "(unknown)" def scan(content): pairs = Counter(); examples = defaultdict(list); rel_cross = 0; rel_total = 0; txt_skipped = 0 for dirpath, dirnames, filenames in os.walk(content): dirnames[:] = [d for d in dirnames if d not in ("pdfs", "solutions")] for n in filenames: if not n.endswith(".html"): continue p = os.path.join(dirpath, n) rel = os.path.relpath(p, content).replace(os.sep, "/") try: here = site_for("/" + rel) except (KeyError, IndexError): continue text = open(p, encoding="utf-8", errors="ignore").read() for attr, url in ATTR.findall(text): path = resolve(rel, url) if path is None: continue if path.lower().endswith(".txt"): txt_skipped += 1 # solution links: the build rewrites them by file name (Chapter 6) continue is_rel = not url.strip().startswith("/") rel_total += is_rel tgt = target_site(path) if tgt != here: pairs[(here, tgt)] += 1 if is_rel: rel_cross += 1 if len(examples[(here, tgt)]) < 1: examples[(here, tgt)].append((rel, url[:70])) return pairs, examples, rel_cross, rel_total, txt_skipped if __name__ == "__main__": pairs, examples, rel_cross, rel_total, txt_skipped = scan(sys.argv[1]) print(f"solution (.txt) links set aside, rewritten by the build: {txt_skipped}") print(f"links that leave their own site: {sum(pairs.values())}") print(f" of the relative links ({rel_total}), those that leave their site: {rel_cross}") for (a, b), n in pairs.most_common(12): rel, url = examples[(a, b)][0] print(f" {a:<15} -> {b:<26}{n:>5} e.g. {url}") Run it: python cross_site.py "/claude-projects/website-content/content" Output (checked by running it): solution (.txt) links set aside, rewritten by the build: 11152 links that leave their own site: 55 of the relative links (43), those that leave their site: 0 languages -> (special route) 30 e.g. /japan/kanji-tiles/ webdevelopment -> (unknown) 17 e.g. /about programming -> (unknown) 8 e.g. /posts/5/edit What the audit found -------------------- - 12,358 link attributes in 4,483 HTML files. - 11,152 of them (90%) are solution (.txt) links. They are written in many stale forms, and the build already rewrites every one from its file name, so they need no work. - No link anywhere uses the site's own domain as an absolute URL (the "absolute, own domain" kind does not appear), so there is nothing to rewrite from https://www.osztromok.com/.... - Of the real links, 85% of all attributes are root-relative, and only 43 are relative paths. None of the relative links leaves its own site. - Only 55 links leave their site, and none is a real cross-site link: 30 go to the kanji tiles page (a languages route, handled as a special route), and 25 are example URLs written inside lessons (such as /about or /posts/5/edit in a web framework tutorial), not links to your own pages. - Why so few cross-links? Cross-references between chapters are written as PLAIN TEXT (a course name and chapter number, rule P11), not as links. - 305 links (2.5%) go to external sites; the most common are code.claude.com, github.com, linkedin.com and youtube.com. External links rot over time and need a periodic check (Learning Website: Framework & Architecture 12). WHY THIS WORKS AS AN ANSWER --------------------------- Rewriting links sounds like the biggest job of a split, and the audit shows it is nearly nothing: the build already handles the solution links, and there are no cross-site content links to rewrite. The real work is redirects for visitors and bookmarks from outside (Exercise 3). Measuring first saved a large piece of work that was never needed.