Migrating from the Old Site

Learning Website with Django

Chapter 11 · Migrating from the Old Site

The old site is live and has bookmarks, search results and links from other places pointing at it. A migration that just switches it off turns every one of them into a 404. This chapter does the opposite: it measures exactly what the old site answers, finds what is missing on the new sites, suggests redirects for a person to approve, and moves one site at a time. The numbers come from the real old site and the real new database.

Run for real, on Django 6.1
The old site is the folder it is served from (debserver/website), checked against the 4,410 pages and 8,403 mirrored files of the new project. The project has 280 tests (all passing). The Apache lines were not run, and the redirect suggestions have not been reviewed by a person.

Step 1: Measure, Don't Remember

The old site is a folder of built files: /hungary/x/y/ is the file hungary/x/y/index.html. That folder is the complete list of old addresses, so a command reads it and asks each new site about each address:

ResultMeaning
sameThe old address answers 200 on the site that now owns its folder
redirectedIt answers a redirect that ends on a page that exists
missingAnything else, including a redirect to a page that does not exist
Old siteCountBefore any redirect
Pages (index.html)3,2053,102 same (97.5% of the 3,182 a site owns), 80 missing
PDFs and solution files5,8315,544 in place, 166 missing (of the 5,710 a site owns)
Owned by no site144 pages and filesMostly resources/php (83), resources/js (30) and resources/csset (19), plus the old front page and /admin/
Dynamic pages (.php)23The Anime Vault: not migrated at all

That is the useful result: 97.5% of old pages already have the same address, because the folder structure was kept, and 80 do not. Reading the list showed why:

  • 32 moved or renamed lessons, for example hungary/hungarian-language/… became hungary/hungarian-lessons/…, and ten Japanese lessons that used to sit under ai/claude-tools/claude-lessons/ are now under japan/.
  • 23 _print pages: the old site had a printable twin of many pages, and the new one does not.
  • 25 others: index pages of renamed folders, old “tentative” course folders that were removed, and japan/kanji-tiles.

The 166 missing files are mostly older copies of Node.js and Express solution files (120, in text-files/ folders the courses no longer use) and 36 in projects/nextjs-rebuild/solutions. The command exits with an error until the list is empty, so it can be run again after every change.

Step 2: Redirects That Someone Reviews

An old address that moved needs a permanent (301) redirect. Redirects are decisions, not text imported from files, so they live in the database, are editable in the admin, and come from a CSV. The matching is deliberately modest:

  1. If exactly one new page has the same file name (ignoring case, and - versus _), propose it.
  2. If several do, list them for a person. If none do, say so.
  3. A name ending in _print falls back to the page without it.

A file name is not proof of the same lesson, so every row of the CSV starts with an empty decision column, and only rows where you write yes are loaded. Loading a file with no “yes” does nothing (“0 created, 0 updated, 32 skipped”). On the real data the first version proposed 32 of the 80, and the _print rule added 23, giving 55 proposals, none ambiguous, 25 with nothing to suggest.

site,old_path,new_site,new_path,why,decision languages,hungary/hungarian-language/hungarian_lesson_01,languages,hungary/hungarian-lessons/hungarian_lesson_01,same file name, ai,ai/claude-tools/claude-lessons/japanese_lesson_dates,languages,japan/japanese-language/japanese-lessons/japanese_lesson_dates,same file name,

The second row is a redirect to a different site: the old ai address becomes a full address on the languages site, built from the same site-address setting as the global bar. The redirect is tried only when there is no page and no folder listing at that address, so a redirect can never hide a real page (a test checks this), and a redirect belongs to one site.

What this does and does not prove
To measure the mechanism I loaded all 55 proposals into my scratch database with --accept-all, a test shortcut. The check then reported 3,102 same, 55 redirected, 25 missing (before: 80 missing). I looked at a sample of the rows but did not review all 55; you should before any of them reaches a real database. A 301 is cached hard by browsers, so a wrong one is hard to undo. Redirects for PDF and solution files are not implemented, so the 166 missing files still need a decision.

Step 3: Links Between Sites

A page written for one site that links to /linux/… with a plain absolute path breaks after the split: that path is on the systems site. Editing thousands of files, and writing a host name into them, would be the wrong fix. Instead the link is rewritten when the page is shown: a link into another site's folders gets that site's address in front, taken from the same setting that differs between development and production. The stored page is unchanged (a test checks this).

On the real content the rewriter found nothing to do: of 11,691 internal links, every one points into its own site. That is a good result for the site map, and it means the rewriter is exercised only by tests today. It leaves alone external links, //cdn… links, relative links, #fragments, /search/ and static files, and keeps query strings.

Broken Links

A second tool reads every stored page and checks each absolute internal link against the pages, folders, redirects and mirrored files. On the real data (3 seconds): 11,715 links checked, 483 broken.

Broken linksCountWhat it is
/japan/hiragana/hiragana-tiles and /japan/katakana/katakana-tiles282Tile pages that exist on the old site and not yet on the new one (like japan/kanji-tiles)
Solutions under /web-development/scripting-and-backend/…174The solution-link problem already noted for several courses in the review file
/resources/japanese23To check
Excel and web-application-security4To check

Step 4: One Site at a Time

While the old site still answers on osztromok.com, a site moves by telling the old host's Apache to send that site's folders to the new subdomain. A command prints the line for one or more sites:

python manage.py apache_cutover languages RedirectMatch 301 ^/(france|germany|hungary|japan|culture|resources/japanese|resources/hungarian)(/.*)?$ https://languages.osztromok.com/$1$2

Without Apache, I ran the same pattern as a Python regular expression over every old address: of 8,892 page and file addresses that a site owns, none was matched by the wrong site or by two sites, and none of the 144 addresses that no site owns was matched. Tests check that no folder is claimed by two sites, that /linuxfoo/ does not match the systems rule, and that sidebar pages follow their subject (for example /sidebar/football/ goes to the humanities site).

What was not verified
The Apache lines were not run (there is no Apache here). Before using them: apachectl configtest, then curl -I on one address of that site, expecting a 301 and the right Location, and keep the old files in place until you have checked. The redirect suggestions are unreviewed, and nothing was tried against a browser's cache.

Hands-On Exercises

Exercise 1

Write the command that lists every address of the old site (a folder of built files) and checks each against the new sites as same, redirected or missing, failing until nothing is missing. Run it on the real old site and explain what the missing ones turned out to be.

📄 View solution
Exercise 2

Build the redirect table, the view hook and the suggest-then-review CSV workflow, matching by file name only when exactly one page fits. Run it on the real data, load the suggestions into a scratch database, measure again, and list what still needs a person.

📄 View solution
Exercise 3

Rewrite cross-site links when a page is shown, write the broken-link checker, and generate the per-site Apache redirect lines. Test the Apache pattern against every real old address, and say clearly what could not be run.

📄 View solution

Chapter 11 Quick Reference

  • Measure the old site from its own folder: 3,205 pages, 5,831 PDFs and solution files, 23 dynamic pages
  • 97.5% of old pages already had the same address (3,102 of 3,182); 80 did not
  • Results: same, redirected (must end on a real page), missing; the command fails until nothing is missing
  • Redirects are decisions: a database table, editable in the admin, loaded from a CSV where only rows marked yes count
  • Matching: exactly one page with the same file name; several or none are left for a person; _print falls back to the page
  • A redirect never hides a real page, belongs to one site, and may point to another site's full address
  • Cross-site links are rewritten when shown, not in the files; real content had none (11,691 internal links, all within their own site)
  • Broken-link check on real data: 483 of 11,715, mostly the hiragana and katakana tile pages and old solution links
  • apache_cutover <site> prints one RedirectMatch 301 line; tested as a regular expression against 8,892 real addresses; not run in Apache
  • Not done: redirects for PDF and solution files, the tile pages, and a person's review of the 55 suggestions