Migrating from the Old Site

Learning Website with Next.js

Chapter 11 ยท Migrating from the Old Site

The old site is live and has bookmarks, search results and links from elsewhere. A migration that just switches it off turns every one of them into a 404. This chapter does what the Django course's chapter on migration did, in TypeScript and in Next.js's own configuration: measure exactly what the old site answers, suggest redirects for a person to approve, keep them in the app's configuration, check the links, and move one site at a time. Because the Django project has done the same measurements, every number here can be compared with an independent implementation.

Run for real, on the old site and 4,422 pages
100 tests pass. The old site's folder (debserver/website) was read, 9,000 of its addresses were checked, and the languages app was asked about 329 of them over HTTP. The Apache lines were not run, and the redirect suggestions have not been reviewed by a person: the lists that are committed are empty.

Step 1: Measure, Don't Remember

The old site is a folder of built files (/hungary/x/y/ is the file hungary/x/y/index.html), so that folder is the complete list of old addresses. A check asks, of each one, whether the new sites have it, without starting a server: from their pages, the folders those pages sit in (each has a listing), the files that will be copied, and the decided redirects. An address is same, redirected (a decided redirect that ends on something that exists) or missing. A redirect to a page that does not exist is not success.

Old siteCountBefore any redirect
Pages3,2053,102 same (of the 3,182 a site owns), 80 missing
PDFs and solution files5,8315,544 in place, 166 missing
Owned by no site144Mostly resources/php, resources/js, resources/csset, the old front page and /admin/
Dynamic pages (.php)23The Anime Vault: not migrated

So most of the site already works at the same address, because the folder structure was kept. Reading the 80 missing pages: 32 are lessons whose folder was renamed or moved, 23 are _print pages (the old site had a printable twin of many pages), and 25 are the index pages of renamed folders, removed “tentative” course folders, and japan/kanji-tiles. These are exactly the numbers the Django project found, from a separate implementation.

Step 2: Redirects You Review First

A permanent redirect is remembered hard by browsers, so a wrong one is hard to undo. So: suggest, then review, then build.

  1. Suggest. If exactly one new page has the same file name (ignoring capitals and - against _), propose it; if several do, list them; if none, say so; a name ending in _print falls back to the page without it. Result: 55 suggestions, 0 ambiguous, 25 with nothing to suggest, and all 55 point where the Django project's do.
  2. Review. The suggestions go in a file whose decision column is empty. Only rows where you write yes become redirects. Loading the file before any review gives 0 redirects (55 rows skipped), so the lists that are committed are empty.
  3. Build. next.config.ts reads the site's list and Next answers the redirects itself. A redirect's source is a pattern in Next, so the characters ( ) : * + ? { } in an old address are made plain. A page that moved to another site gets a full address.
async redirects() { const file = join(process.env["LW_REDIRECTS_DIR"] ?? join(process.cwd(), "..", "..", "redirects"), "languages.json"); return existsSync(file) ? nextRedirects(JSON.parse(readFileSync(file, "utf8")) as Redirect[], "languages", environment()) : []; },

Over Real HTTP

To see the mechanism work, I made a scratch list with every suggestion accepted (--accept-all, which prints “NOT reviewed”: 19 for languages, 8 programming, 11 systems, 11 AI, 6 humanities; not committed), built the languages app with it, and asked the running app about every old address the languages site owns:

Asked over HTTPSameRedirectedMissing
319 old pages297193 (france/french-language, hungary/hungarian-language, japan/kanji-tiles)
10 old files406 (in resources/hungarian and the old pdfs folder)

The HTTP answers and the check from Step 1 agree on all 329 addresses. All redirects are 308. Note that an old link with a trailing slash takes two hops: first Next's own tidy-up to the address without the slash, then the real redirect. It works and each hop is cheap, but it is not one.

Can the suggester fail?
Taking out the print-version rule and running the report again gave “32 suggestions, compared with the Django project's 55: same 32, only there 23”. Put back.

Step 3: Links

Broken links

Every absolute internal link of every page is checked against the pages, folders, files and redirects: 11,751 links, 483 broken (305 to pages, 178 to files), in a few seconds (not timed precisely). That is the Django project's result, category by category. 282 of the page links are to the hiragana and katakana tile pages, which exist on the old site and not yet on the new one; 174 of the file links are solution files under /web-development/scripting-and-backend that were moved or never made. It is not set to fail the build: 483 known problems would stop every build until each is decided.

Links between sites

A page written for the old single site that links to /linux/โ€ฆ breaks after the split, because /linux/ is on the systems site. The fix is made when the page is built, not in the files: a link into another site's folders gets that site's address in front. External links, // links, anchors, /search/ and the site's own links are left alone. On the real content this found 0 links to rewrite: every page links only inside its own site. So the function is tested only on made-up examples, and a function that has never met real data deserves suspicion.

Step 4: One Site at a Time

While the old site still answers on osztromok.com, a site moves by telling the old host's Apache to send that site's folders to its new subdomain. To go back, remove the line.

RedirectMatch 301 ^/(france|germany|hungary|japan|culture|resources/japanese|resources/hungarian)(/.*)?$ https://languages.osztromok.com/$1$2

Without Apache, the same pattern was run as a regular expression over every old address: of 8,892 that a site owns, none was matched by the wrong site or by two, and none of the 144 that no site owns was matched by any rule (the Django project's numbers).

What the real data can and cannot prove
I planted a mistake: the pattern no longer needed a slash or the end after the folder, so /linuxfoo/ would match the systems rule. On the real addresses the result was 0 wrong matches, because no real address looks like that; a unit test with that awkward address failed, as it should. The real addresses show the rules cover what exists; only an awkward test shows they do not cover too much. Both are needed.
What was not verified
The Apache lines were never run: use apachectl configtest, then curl -I on one address of that site expecting a 301 and the right Location. The 55 suggestions are unreviewed. Only the languages site has an app, so only its redirects were tried over HTTP. The tile pages and the 25 unmatched pages were not built or decided. One of my own tests was wrong, not the code: I expected 11 links and the page set had 10 (I had counted by eye).

Hands-On Exercises

Exercise 1

Write the check that reads the old site's folder and says, for every address, whether the new sites have it (same, redirected or missing). Run it on the real old site and explain what the missing ones turned out to be.

๐Ÿ“„ View solution
Exercise 2

Suggest redirects for the missing pages into a file for review, load only the rows marked yes into the app's configuration, and prove over real HTTP what the languages site does with every old address it owns. Keep the committed list empty until a person has reviewed it.

๐Ÿ“„ View solution
Exercise 3

Check every internal link of every page, rewrite links into other sites when a page is built, and generate the per-site Apache cut-over lines. Test the lines against every real old address and against an awkward one, and say what each of the two can prove.

๐Ÿ“„ View solution

Chapter 11 Quick Reference

  • The old site is a folder of built files: read it to get the real list of addresses; 3,205 pages, 5,831 PDFs and solutions, 23 .php, 144 owned by no site
  • Same / redirected / missing, worked out without a server; a redirect to nothing is not success; 3,102 same, 80 missing before; 55 redirected and 25 missing after the suggestions
  • Suggest (one page with that file name, or a print twin), review (decision column), build: only rows marked yes become redirects; the committed lists are empty
  • Redirects live in next.config.ts (redirects()), read from redirects/<site>.json; special characters in an old address are escaped; another site gets a full address
  • Over HTTP: 319 pages and 10 files asked; 19 redirected (308); the HTTP and logic checks agree on all 329; an old link with a slash takes two hops
  • Link check: 11,751 links, 483 broken (305 pages, 178 files), the same as Django; 0 links point into another site
  • Links into another site are rewritten when the page is built, not in the files
  • Apache cut-over: one RedirectMatch 301 line per site; 8,892 old addresses matched by exactly their own site, 0 of the 144 unowned matched; not run in Apache
  • Real data proves rules cover what exists; an awkward test proves they do not cover too much
  • Not done: Apache, reviewing the 55 suggestions, the tile pages, the other sites' apps