Deployment and Operations

Learning Website with Next.js

Chapter 12 · Deployment, Testing & Operations

The last chapter is about the day after the code works: how to know a release is healthy before anyone sees it, how to get the accounts back after a mistake, how Apache and systemd run the sites, and how to undo a release. As in the other chapters, what could be run was run, and what could not is said plainly.

Run for real, and not run
Run: the production-settings check, the languages app built with LW_ENV=prod and started with next start and looked at over HTTP, a backup and a restore of the real accounts database, the generated Apache and systemd files, and 114 tests. Not run: Apache, systemd, cron, certificates, and deploy.sh (Windows has none of them; the script has only had bash -n).

One Node Process Per Site, Behind Apache

There were two shapes to choose from: a static export (just files, any web server) or a Node server (next start). The login API, the health address and the per-request checks need a server, so each site runs as its own Node process listening on 127.0.0.1 only, on the port it already has (portfolio 3000, languages 3001, and so on). Apache handles HTTPS and passes each host name to its port. A static export of the pages alone stays possible later if the accounts move elsewhere.

Both the virtual hosts and the systemd units are generated from the site map (node ops.mjs config), so they cannot disagree with it. Only sites that have an app get a virtual host; a site without one keeps answering from the old site until its turn.

<VirtualHost *:443> ServerName languages.osztromok.com SSLEngine on ... AddOutputFilterByType DEFLATE application/json application/xml text/plain ProxyPreserveHost On RequestHeader set X-Forwarded-Proto "https" ProxyPass / http://127.0.0.1:3001/ ProxyPassReverse / http://127.0.0.1:3001/ </VirtualHost>
The line that fails quietly
ProxyPreserveHost On. The login check compares the browser's Origin with the Host the app receives. Without that line the app sees 127.0.0.1:3001, every login is refused as cross-site, and every page still works, so nobody notices. The smoke test below looks at exactly this.

Compression, Measured

Next compresses its pages: a lesson page was 59,774 bytes plain and 12,406 bytes with gzip. What a route handler sends is not compressed: asked with Accept-Encoding: gzip, /search-index.json came back as 1,018,766 bytes, exactly the same as without, and so did the sitemap. Compressed with gzip it would be 326,259 bytes. So the virtual host asks Apache to compress JSON, XML and text. Untested: I could not run Apache, so I have not seen this work, nor seen it leave alone a response Next has already compressed (that is its documented behaviour).

Before the Release: Settings

node ops.mjs check stops the release on mistakes that would otherwise show up as a strange site: LW_ENV not prod (every link would point at localhost), no content folder, no data folder or one inside the release folder (the next release would not see the accounts), and the experiment pages switched on. With a good environment it says production settings: ok. With development settings, no data folder and the experiment pages on, it printed three errors and exited with 1.

While It Runs: Health and a Smoke Test

/healthz answers 200 only when the database answers and the site has pages ({"status":"ok","pages":561}), is never cached, and says nothing else. The smoke test then visits the running site as a visitor would, with the Host header Apache will send, and checks 11 things: the health address; the first, the middle and the last page of the site; that a missing address answers 404 and not 200; that a real PDF is served with nosniff; robots, sitemap and search index; and two logins, one from another site's page (must be 403) and one from our own (must get past the lock, 400). Against the running languages app: 11 checks passed. With the app stopped, all 11 failed with ECONNREFUSED instead of hanging.

Three mistakes in my own script that only the real run found
  1. Trailing slash. I asked for pages as /x/; the site answers with a 308 to /x. Three checks failed with “308”. Ask for the address the site really uses.
  2. The Host header. Node's fetch() does not send the Host you give it, so the own-origin login got 403 instead of 400. curl with the same header got 400. The script uses node:http. Notice that the check for a login from another site had passed all along, for the wrong reason.
  3. A non-ASCII address. One sampled page has し in its address and node:http refuses it unescaped, so addresses are percent-encoded, as a browser does.
The unit tests, which use a pretend site, could not have found any of the three.

Backups You Have Restored

The only data that cannot be rebuilt from the content folder is the accounts database. A plain file copy is the wrong tool: the database runs in WAL mode, so recent writes sit in accounts.sqlite-wal, and a copy made mid-write can be damaged. VACUUM INTO asks SQLite itself for a complete, consistent copy while the site runs. The order matters: the copy is made under a temporary name, checked, only then renamed, and only after that are backups beyond the newest 14 removed, so a broken backup can never push out a good one.

A restore checks the backup before touching anything, keeps the database it replaces (accounts.sqlite.before-restore), and moves the old -wal and -shm files away with it, because they belong to the old database. Run for real: two accounts, a backup, a third account, a restore; the third was gone and the three-account database was kept. A file that is not a database was refused: refusing to restore bad.sqlite: file is not a database, and the accounts were unchanged.

What a test cannot prove, and a test that proved nothing
(1) Removing the check on the new backup still fails a test, but only because that test also counts rows; a backup that opens but is damaged inside is not something any test makes, so that check is not proved. (2) My first restore demonstration used a third account called “cy”; the site rightly refuses names under three letters, so nothing changed after the backup and the restore had nothing to undo. A test that cannot fail is not a test: I redid it with a valid name.

A Release With a Way Back

deploy.sh release <ref> builds a new folder next to the old one, then: 1. backs up the database; 2. checks the settings; 3. installs, tests and builds; 4. switches the current symlink (atomic) and restarts the services; 5. runs the smoke test against every running site, and if one check fails, rolls back by itself. A rollback is moving the symlink and restarting. There are no migrations (tables are only created if missing), so old code normally still works with the database; if a release changed a table, restore the backup taken at its start.

Not run — and what I would expect to go wrong first
Apache (no virtual host, no mod_deflate, no certificate); systemd; cron; deploy.sh (only bash -n); a real rollback. The unit sets ProtectSystem=strict with only the data folder writable, and Next may want to write to its own .next/cache, so the first start on the server may fail until that folder is added to ReadWritePaths. Only the languages app exists, so the other seven sites have no unit and no virtual host yet.

A first rollout that cannot hurt

  1. Install on the server and run the settings check.
  2. Start the languages unit on its port and run the smoke test against it directly (curl -H "Host: ...").
  3. Only then add its virtual host (apachectl configtest, then reload) and run the smoke test over HTTPS from outside.
  4. Add the RedirectMatch line for its folders to the old site (Chapter 11), and watch the logs for 404s.
  5. Keep the old site's lines in place for a few weeks, so going back is removing one line.

Hands-On Exercises

Exercise 1

Write the production-settings check, a health address and a smoke test that visits a running site with a given Host header. Run them against the built app, and say what each of the mistakes you hit along the way was.

📄 View solution
Exercise 2

Back up the accounts database while the site is running, verify the copy before trusting it, prune old backups safely, and restore one. Show with a real run that an account made after the backup is gone after the restore.

📄 View solution
Exercise 3

Generate the Apache virtual hosts and systemd units from the site map, measure what is and is not compressed, and write a release script with an automatic rollback. List exactly what you could not run.

📄 View solution

Chapter 12 Quick Reference

  • One Node process per site (next start -p PORT -H 127.0.0.1); Apache does HTTPS and passes the host on; config and units generated from the site map
  • ProxyPreserveHost On is required, or every login is refused as cross-site while pages still work
  • Next compresses pages but not route-handler responses: search index 1,018,766 bytes either way; Apache DEFLATE added (untested)
  • node ops.mjs check: production settings; /healthz: 200 only with a database and pages, never cached
  • node ops.mjs smoke <url> <host>: 11 checks as a visitor; use node:http (fetch drops the Host header), no trailing slash, encode non-ASCII addresses
  • Back up with VACUUM INTO, verify, rename, then prune; never overwrite; restore verifies first and keeps the old database
  • deploy.sh: backup, check, build, switch symlink, restart, smoke test, automatic rollback; rollback = move the symlink back
  • Not run: Apache, systemd, cron, certificates, deploy.sh; likely first failure: the unit's write permission for .next/cache

Course Complete

You have built a Next.js workspace that serves several sites from shared packages, with the tools to move the old site across without breaking its addresses. What each chapter left you able to do:

ChapterYou can now…
1 · Monorepo for Many SitesLay out several apps and shared packages in one npm workspace
2 · Subdomains & MiddlewareChoose the site from the host name, strictly, in proxy.ts or as separate apps
3 · The Content Model in TypeScriptParse each file's banner into typed pages, shared by every site
4 · Reading Fragments at Build TimeGenerate every page statically and inject fragments safely, with their scripts running once
5 · The Shared Design SystemShare one theme with a per-site accent, keeping fragment CSS apart from module CSS
6 · Layout & NavigationDerive menus, breadcrumbs and links between sites from the data
7 · The Languages SiteGive one site its own front page, accent colour and answer blocks
8 · PDFs, Solutions & AssetsServe downloads as build artefacts with safe headers, and add a Copy button
9 · Search, Sitemaps & MetadataAdd a static client-side search index, sitemaps and canonical metadata
10 · Accounts & Progress (Optional)Add logins and progress without breaking the static pages
11 · Migrating from the Old SiteMeasure old addresses, review redirects and cut over one site at a time
12 · Deployment, Testing & OperationsCheck settings, monitor, back up, restore and release with a way back
What is still yours to do
Review the 55 suggested redirects; decide the 25 unmatched pages; run the Apache and systemd files on a real server (this course only generated and read them); build apps for the other six sites from the same packages. Then compare this course with Learning Website with Django: both end at the same numbers on the same content, and that agreement is what makes the choice between them one of preference, not of what works.