Migrating from the Old Site
Learning Website with Django
Chapter 11 · Migrating from the Old Site
The old site is live and has bookmarks, search results and links from other places pointing at it. A migration that just switches it off turns every one of them into a 404. This chapter does the opposite: it measures exactly what the old site answers, finds what is missing on the new sites, suggests redirects for a person to approve, and moves one site at a time. The numbers come from the real old site and the real new database.
debserver/website), checked against the 4,410 pages and
8,403 mirrored files of the new project. The project has 280 tests (all passing). The Apache lines were not
run, and the redirect suggestions have not been reviewed by a person.
Step 1: Measure, Don't Remember
The old site is a folder of built files: /hungary/x/y/ is the file hungary/x/y/index.html. That
folder is the complete list of old addresses, so a command reads it and asks each new site about each address:
| Result | Meaning |
|---|---|
| same | The old address answers 200 on the site that now owns its folder |
| redirected | It answers a redirect that ends on a page that exists |
| missing | Anything else, including a redirect to a page that does not exist |
| Old site | Count | Before any redirect |
|---|---|---|
Pages (index.html) | 3,205 | 3,102 same (97.5% of the 3,182 a site owns), 80 missing |
| PDFs and solution files | 5,831 | 5,544 in place, 166 missing (of the 5,710 a site owns) |
| Owned by no site | 144 pages and files | Mostly resources/php (83), resources/js (30) and resources/csset (19), plus the old front page and /admin/ |
Dynamic pages (.php) | 23 | The Anime Vault: not migrated at all |
That is the useful result: 97.5% of old pages already have the same address, because the folder structure was kept, and 80 do not. Reading the list showed why:
- 32 moved or renamed lessons, for example
hungary/hungarian-language/…becamehungary/hungarian-lessons/…, and ten Japanese lessons that used to sit underai/claude-tools/claude-lessons/are now underjapan/. - 23
_printpages: the old site had a printable twin of many pages, and the new one does not. - 25 others: index pages of renamed folders, old “tentative” course folders that were
removed, and
japan/kanji-tiles.
The 166 missing files are mostly older copies of Node.js and Express solution files (120, in text-files/
folders the courses no longer use) and 36 in projects/nextjs-rebuild/solutions. The command exits with an
error until the list is empty, so it can be run again after every change.
Step 2: Redirects That Someone Reviews
An old address that moved needs a permanent (301) redirect. Redirects are decisions, not text imported from files, so they live in the database, are editable in the admin, and come from a CSV. The matching is deliberately modest:
- If exactly one new page has the same file name (ignoring case, and
-versus_), propose it. - If several do, list them for a person. If none do, say so.
- A name ending in
_printfalls back to the page without it.
A file name is not proof of the same lesson, so every row of the CSV starts with an empty decision column,
and only rows where you write yes are loaded. Loading a file with no “yes” does nothing
(“0 created, 0 updated, 32 skipped”). On the real data the first version proposed 32 of the 80, and the
_print rule added 23, giving 55 proposals, none ambiguous, 25 with nothing to suggest.
The second row is a redirect to a different site: the old ai address becomes a full address on
the languages site, built from the same site-address setting as the global bar. The redirect is tried only when there
is no page and no folder listing at that address, so a redirect can never hide a real page (a test checks this), and a
redirect belongs to one site.
--accept-all, a test
shortcut. The check then reported 3,102 same, 55 redirected, 25 missing (before: 80 missing). I looked at a
sample of the rows but did not review all 55; you should before any of them reaches a real database. A
301 is cached hard by browsers, so a wrong one is hard to undo. Redirects for PDF and solution files are not
implemented, so the 166 missing files still need a decision.
Step 3: Links Between Sites
A page written for one site that links to /linux/… with a plain absolute path breaks after the split:
that path is on the systems site. Editing thousands of files, and writing a host name into them, would be the wrong
fix. Instead the link is rewritten when the page is shown: a link into another site's folders gets
that site's address in front, taken from the same setting that differs between development and production. The stored
page is unchanged (a test checks this).
On the real content the rewriter found nothing to do: of 11,691 internal links, every one points into
its own site. That is a good result for the site map, and it means the rewriter is exercised only by tests today. It
leaves alone external links, //cdn… links, relative links, #fragments, /search/
and static files, and keeps query strings.
Broken Links
A second tool reads every stored page and checks each absolute internal link against the pages, folders, redirects and mirrored files. On the real data (3 seconds): 11,715 links checked, 483 broken.
| Broken links | Count | What it is |
|---|---|---|
/japan/hiragana/hiragana-tiles and /japan/katakana/katakana-tiles | 282 | Tile pages that exist on the old site and not yet on the new one (like japan/kanji-tiles) |
Solutions under /web-development/scripting-and-backend/… | 174 | The solution-link problem already noted for several courses in the review file |
/resources/japanese | 23 | To check |
| Excel and web-application-security | 4 | To check |
Step 4: One Site at a Time
While the old site still answers on osztromok.com, a site moves by telling the old host's Apache
to send that site's folders to the new subdomain. A command prints the line for one or more sites:
Without Apache, I ran the same pattern as a Python regular expression over every old address: of 8,892
page and file addresses that a site owns, none was matched by the wrong site or by two sites, and none
of the 144 addresses that no site owns was matched. Tests check that no folder is claimed by two sites, that
/linuxfoo/ does not match the systems rule, and that sidebar pages follow their subject (for example
/sidebar/football/ goes to the humanities site).
apachectl configtest, then
curl -I on one address of that site, expecting a 301 and the right Location, and keep the old
files in place until you have checked. The redirect suggestions are unreviewed, and nothing was tried against a
browser's cache.
Hands-On Exercises
Write the command that lists every address of the old site (a folder of built files) and checks each against the new sites as same, redirected or missing, failing until nothing is missing. Run it on the real old site and explain what the missing ones turned out to be.
📄 View solutionBuild the redirect table, the view hook and the suggest-then-review CSV workflow, matching by file name only when exactly one page fits. Run it on the real data, load the suggestions into a scratch database, measure again, and list what still needs a person.
📄 View solutionRewrite cross-site links when a page is shown, write the broken-link checker, and generate the per-site Apache redirect lines. Test the Apache pattern against every real old address, and say clearly what could not be run.
📄 View solutionChapter 11 Quick Reference
- Measure the old site from its own folder: 3,205 pages, 5,831 PDFs and solution files, 23 dynamic pages
- 97.5% of old pages already had the same address (3,102 of 3,182); 80 did not
- Results: same, redirected (must end on a real page), missing; the command fails until nothing is missing
- Redirects are decisions: a database table, editable in the admin, loaded from a CSV where only rows marked
yescount - Matching: exactly one page with the same file name; several or none are left for a person;
_printfalls back to the page - A redirect never hides a real page, belongs to one site, and may point to another site's full address
- Cross-site links are rewritten when shown, not in the files; real content had none (11,691 internal links, all within their own site)
- Broken-link check on real data: 483 of 11,715, mostly the hiragana and katakana tile pages and old solution links
apache_cutover <site>prints oneRedirectMatch 301line; tested as a regular expression against 8,892 real addresses; not run in Apache- Not done: redirects for PDF and solution files, the tile pages, and a person's review of the 55 suggestions