learning-website-framework1-11 Exercise 2: A Safe Link Rewriter
===========================================================================
Even if the real content needs no rewriting today, you need a tool that does it
correctly for the day a cross-site link is added. Write one that:
- rewrites a root-relative href only when its target belongs to another site
- leaves external links, anchors, relative paths and solution (.txt) links alone
- REPORTS a path that belongs to no site instead of guessing
- has a dry-run mode, so it can be run over everything without changing a file
Needs links.py and sitemap.py (Chapter 5). Save as rewrite_links.py:
import os, re, sys
from urllib.parse import urlsplit
from links import link_for
from sitemap import site_for
HREF = re.compile(r'(\bhref=")([^"]*)(")', re.I)
def rewrite_html(html, here, env="production"):
"""Rewrite root-relative links that point at another site. Returns (new_html, changes, unknown)."""
changes, unknown = [], []
def fix(m):
url = m.group(2)
parts = urlsplit(url)
if not url.startswith("/") or url.startswith("//") or parts.path.lower().endswith(".txt"):
return m.group(0) # external, relative, anchor, or a solution link
try:
new = link_for(here, url, env)
except (KeyError, IndexError):
unknown.append(url) # report, never guess
return m.group(0)
if new != url:
changes.append((url, new))
return m.group(1) + new + m.group(3)
return HREF.sub(fix, html), changes, unknown
if __name__ == "__main__":
demo = ('same site '
'other site '
'sidebar '
'solution '
'external '
'unknown')
new, changes, unknown = rewrite_html(demo, "languages")
print("demo page, built for the languages site:")
for old, n in changes:
print(f" rewrote {old} -> {n}")
print(f" left alone: {len(HREF.findall(demo)) - len(changes) - len(unknown)}; reported as unknown: {unknown}")
print()
root = sys.argv[1]
pages = total_changes = total_unknown = 0
ex = []
for dirpath, dirnames, filenames in os.walk(root):
dirnames[:] = [d for d in dirnames if d not in ("pdfs", "solutions")]
for n in filenames:
if not n.endswith(".html"):
continue
p = os.path.join(dirpath, n)
rel = os.path.relpath(p, root).replace(os.sep, "/")
try:
here = site_for("/" + rel)
except (KeyError, IndexError):
continue
pages += 1
_, ch, unk = rewrite_html(open(p, encoding="utf-8", errors="ignore").read(), here)
total_changes += len(ch)
total_unknown += len(unk)
if unk and len(ex) < 3:
ex.append((rel.split("/")[-1], unk[0]))
print(f"dry run over {pages} real pages (nothing written): {total_changes} links would be rewritten, {total_unknown} reported as unknown")
for f, u in ex:
print(f" unknown example: {u} (in {f})")
Run it:
python rewrite_links.py "/claude-projects/website-content/content"
Output (checked by running it):
demo page, built for the languages site:
rewrote /linux/shell-and-scripting/vim/#modes -> https://systems.osztromok.com/linux/shell-and-scripting/vim/#modes
rewrote /sidebar/football/ -> https://humanities.osztromok.com/sidebar/football/
left alone: 3; reported as unknown: ['/no-such-area/page/']
dry run over 4483 real pages (nothing written): 0 links would be rewritten, 23 reported as unknown
unknown example: /posts/5/edit (in rails1-4.html)
unknown example: /mark-used/<id>/ (in django_food_tracker_1_course_combined.html)
unknown example: /mark-used/<id>/ (in food_tracker_django_1_9.html)
Reading the result
------------------
- The demo shows the two cases that matter: a link to another site becomes an
absolute URL (with the #modes anchor kept), and a sidebar link follows its
subject to the right site.
- The dry run over all 4,483 real pages changes nothing (0 links), as the
audit predicted. It reports 23 links whose targets belong to no site; they
are the example URLs inside lessons, and the tool correctly leaves them
alone.
- This version only handles href attributes; src attributes (images, scripts)
would need the same treatment if any ever pointed across sites.
WHY THIS WORKS AS AN ANSWER
---------------------------
A rewriter that never guesses, reports what it cannot place and starts with a
dry run is safe to run over thousands of files. Running it on the real content
and getting zero changes is itself a useful result: it confirms the audit.