learning-website-framework1-11 Exercise 2: A Safe Link Rewriter =========================================================================== Even if the real content needs no rewriting today, you need a tool that does it correctly for the day a cross-site link is added. Write one that: - rewrites a root-relative href only when its target belongs to another site - leaves external links, anchors, relative paths and solution (.txt) links alone - REPORTS a path that belongs to no site instead of guessing - has a dry-run mode, so it can be run over everything without changing a file Needs links.py and sitemap.py (Chapter 5). Save as rewrite_links.py: import os, re, sys from urllib.parse import urlsplit from links import link_for from sitemap import site_for HREF = re.compile(r'(\bhref=")([^"]*)(")', re.I) def rewrite_html(html, here, env="production"): """Rewrite root-relative links that point at another site. Returns (new_html, changes, unknown).""" changes, unknown = [], [] def fix(m): url = m.group(2) parts = urlsplit(url) if not url.startswith("/") or url.startswith("//") or parts.path.lower().endswith(".txt"): return m.group(0) # external, relative, anchor, or a solution link try: new = link_for(here, url, env) except (KeyError, IndexError): unknown.append(url) # report, never guess return m.group(0) if new != url: changes.append((url, new)) return m.group(1) + new + m.group(3) return HREF.sub(fix, html), changes, unknown if __name__ == "__main__": demo = ('same site ' 'other site ' 'sidebar ' 'solution ' 'external ' 'unknown') new, changes, unknown = rewrite_html(demo, "languages") print("demo page, built for the languages site:") for old, n in changes: print(f" rewrote {old} -> {n}") print(f" left alone: {len(HREF.findall(demo)) - len(changes) - len(unknown)}; reported as unknown: {unknown}") print() root = sys.argv[1] pages = total_changes = total_unknown = 0 ex = [] for dirpath, dirnames, filenames in os.walk(root): dirnames[:] = [d for d in dirnames if d not in ("pdfs", "solutions")] for n in filenames: if not n.endswith(".html"): continue p = os.path.join(dirpath, n) rel = os.path.relpath(p, root).replace(os.sep, "/") try: here = site_for("/" + rel) except (KeyError, IndexError): continue pages += 1 _, ch, unk = rewrite_html(open(p, encoding="utf-8", errors="ignore").read(), here) total_changes += len(ch) total_unknown += len(unk) if unk and len(ex) < 3: ex.append((rel.split("/")[-1], unk[0])) print(f"dry run over {pages} real pages (nothing written): {total_changes} links would be rewritten, {total_unknown} reported as unknown") for f, u in ex: print(f" unknown example: {u} (in {f})") Run it: python rewrite_links.py "/claude-projects/website-content/content" Output (checked by running it): demo page, built for the languages site: rewrote /linux/shell-and-scripting/vim/#modes -> https://systems.osztromok.com/linux/shell-and-scripting/vim/#modes rewrote /sidebar/football/ -> https://humanities.osztromok.com/sidebar/football/ left alone: 3; reported as unknown: ['/no-such-area/page/'] dry run over 4483 real pages (nothing written): 0 links would be rewritten, 23 reported as unknown unknown example: /posts/5/edit (in rails1-4.html) unknown example: /mark-used/<id>/ (in django_food_tracker_1_course_combined.html) unknown example: /mark-used/<id>/ (in food_tracker_django_1_9.html) Reading the result ------------------ - The demo shows the two cases that matter: a link to another site becomes an absolute URL (with the #modes anchor kept), and a sidebar link follows its subject to the right site. - The dry run over all 4,483 real pages changes nothing (0 links), as the audit predicted. It reports 23 links whose targets belong to no site; they are the example URLs inside lessons, and the tool correctly leaves them alone. - This version only handles href attributes; src attributes (images, scripts) would need the same treatment if any ever pointed across sites. WHY THIS WORKS AS AN ANSWER --------------------------- A rewriter that never guesses, reports what it cannot place and starts with a dry run is safe to run over thousands of files. Running it on the real content and getting zero changes is itself a useful result: it confirms the audit.