Links and Redirects

Learning Website: Framework & Architecture

Chapter 11 ยท Rewriting Links & Redirects

Splitting a site breaks two kinds of link: the links inside your pages that assume one site, and the links to your pages that live in search engines, bookmarks and other people's websites. The first kind sounds like the bigger job, because there are thousands of pages. The second kind is the one that really matters. This chapter starts by measuring, because the measurement changes the plan.

Measure before you plan
Every figure in this chapter comes from a script run over the real content folder (Exercises 1 to 3). The result is surprising: the feared job, rewriting links in thousands of pages, turns out to be almost nothing.

The Link Audit

The census read every href and src attribute in 4,483 HTML files:

Kind of linkCountShare
Root-relative (/path/)10,51685.1%
Relative (path, ../path)1,26410.2%
Absolute, external3052.5%
Anchor only (#top)2712.2%
Absolute, on the site's own domain00%
Other (empty, mailto)20.0%

Four facts follow from the census and the cross-site check that came after it:

  1. 90% of all links (11,152) are solution .txt links. They are written in many stale forms, and the build already rewrites every one of them from its file name (Learning Website: Framework & Architecture 6). No work.
  2. No page links to its own site by absolute URL. There is nothing like https://www.osztromok.com/... to rewrite.
  3. Only 55 links leave their own site, and none of them is a real cross-site link. Thirty go to the kanji tiles page, which belongs to the languages area. Twenty-five are example addresses inside lessons, such as /about in a web framework tutorial, not links to your own pages.
  4. Relative links are safe: there are only 43 once solution links are set aside, and none crosses a site boundary.
Why there are so few cross-links
Cross-references between chapters are written as plain text: a course name and chapter number, as the site's own rule requires, never as a link. That habit, adopted for readability, turns out to make the split almost free. If you later want those references clickable, generate the links from the course name deliberately (a “see also” list) instead of scattering paths through the text.

A Rewriter You Hope Not to Need

The content needs no rewriting today, but the first cross-site link someone adds will, so a correct tool belongs in the build. Rules for a safe rewriter:

  • Rewrite only what is needed: a root-relative link whose target belongs to a different site becomes an absolute URL on that site; everything else is left exactly as it is.
  • Never touch external links, anchors, relative paths or solution links.
  • Report what you cannot place (a path in no site) instead of guessing.
  • Keep the path, query, fragment and trailing slash exactly.
  • Have a dry run that reports changes without writing files.

Exercise 2 builds this on top of the link resolver from Chapter 5. Its dry run over all 4,483 real pages changes nothing, as the audit predicted, and it reports 23 links it cannot place, which are the example addresses inside lessons.

# demo page built for the languages site /linux/shell-and-scripting/vim/#modes -> https://systems.osztromok.com/linux/shell-and-scripting/vim/#modes /sidebar/football/ -> https://humanities.osztromok.com/sidebar/football/ /japan/hiragana-1/ (same site: left alone) solutions/x_exercise1.txt (solution link: left alone) /no-such-area/page/ (reported as unknown, never guessed)

Redirects: Where the Real Work Is

Visitors, bookmarks, search results and other people's pages all use the old addresses. Every one of them must reach the new address in a single hop. The rules should come from the site map, not be written by hand, and their order matters: specific rules must come before general ones.

OrderRule typeExample
1Resource folders that move with an area/resources/japanese/... and /resources/hungarian/... go to the languages site
2Sidebar subjects/sidebar/linux/... goes to the systems site, /sidebar/football/... to humanities
3Ordinary folders/hungary/... goes to languages, /linux/... to systems
# generated from the site map: the first lines of 17 rules RedirectMatch 301 ^/resources/japanese(/.*)?$ https://languages.osztromok.com$0 RedirectMatch 301 ^/resources/hungarian(/.*)?$ https://languages.osztromok.com$0 RedirectMatch 301 ^/sidebar/ai(/.*)?$ https://ai.osztromok.com$0 RedirectMatch 301 ^/sidebar/linux(/.*)?$ https://systems.osztromok.com$0 # ... then one rule per site for its ordinary folders

Exercise 3 generates all 17 rules and tests every old URL. The result: all 4,925 pages, plus two sample resource files, are redirected to the right site; none is matched by no rule; none is sent to the wrong site; and no URL matches rules for two different sites. The only path left unmatched on purpose is /sidebar/ itself, because the sidebar splits across six sites, so decide where the old index page should go (the root landing page is a sensible choice).

A script test is not a server test
These rules were checked with Python's regular expressions. They have not been run through Apache. Before launch, test each rule on the real server with curl -I, and look especially at: the $0 substitution (the whole matched path); non-ASCII addresses, which are common here (the hiragana and katakana pages have names such as hiragana_ใ‚.html) and must be matched and re-encoded correctly; and the address with and without a trailing slash.

Keeping Links Healthy Afterwards

  • Check the built site, not the source. After each build, follow every internal link in the output and confirm it reaches a file. The 282 dead tile links found in Chapter 10 are exactly what this finds.
  • Check external links on a schedule. The 305 external links (the most common hosts are code.claude.com, github.com, linkedin.com and youtube.com) rot over time. A monthly check is plenty.
  • Watch the 404s. The server log shows which old addresses visitors still ask for, which is the best list of redirects you missed.
  • Never reuse an old address for different content while its redirect still exists.

Hands-On Exercises

Exercise 1

Write a link census that counts every href and src by kind (root-relative, relative, absolute on your own domain, external, anchor). Then write a second script that resolves each link to a site and counts the links that leave their own site, setting solution links aside. Summarise what it means for the migration.

๐Ÿ“„ View solution
Exercise 2

Write a link rewriter that changes only root-relative links to another site, leaves everything else alone, reports unknown targets instead of guessing, and has a dry-run mode. Test it on a small demo page and then in dry-run mode over all the real pages.

๐Ÿ“„ View solution
Exercise 3

Generate the complete ordered set of Apache redirect rules from the site map, including resource folders and sidebar subjects, and test every old page URL against it: each must match exactly one site, the right one. List what the test cannot prove and how you will check it on the server.

๐Ÿ“„ View solution

Chapter 11 Quick Reference

  • Audit first: 12,358 link attributes in 4,483 files; 90% are solution links the build already fixes
  • 0 absolute links to the site's own domain; only 43 relative links, none crossing a site
  • Only 55 links leave their site and none is a real cross-site link (kanji tiles route and example URLs in lessons)
  • Cross-references are plain text (course name and chapter number), which is why the split is cheap
  • A safe rewriter changes only cross-site root-relative links, reports unknowns, and has a dry run
  • Redirect order: resource folders, then sidebar subjects, then ordinary folders; generate from the site map
  • 17 rules cover all 4,925 pages with no conflicts; /sidebar/ itself needs a decision
  • Python tests are not Apache tests: try curl -I on the server, especially non-ASCII and trailing-slash cases
  • After launch: check the built site's links, check external links monthly, read the 404 log