learning-website-nextjs1-12 Exercise 3: Apache and systemd Per Site, a Release with a Way Back, and What Was Not Run ======================================================================================================================= HOW THE SITES RUN. Each site is its own Node process ("next start"), listening on 127.0.0.1 only, on the port the site already has in DEV_PORTS (portfolio 3000, languages 3001, ...). Apache terminates HTTPS and passes each host to its port. The alternative, a static export, was ruled out by what the site does: accounts, the login API, the health address and per-request redirects need a server. (A static export of the pages alone is possible later, if the accounts are moved elsewhere.) Save as packages/operations/src/vhosts.ts: import { DEV_PORTS, DOMAIN, type SiteName } from "@lw/sites"; export interface HostOptions { readonly domain?: string; readonly certDir?: string; /** Which sites have an app running. A site with no app gets no virtual host, so its name keeps answering from the old site. */ readonly sites: readonly (SiteName | "portfolio")[]; readonly port?: (site: SiteName | "portfolio") => number; } const hostOf = (site: SiteName | "portfolio", domain: string) => (site === "portfolio" ? domain : `${site}.${domain}`); /** * One HTTPS virtual host that passes everything to the site's Node process. * ProxyPreserveHost On is not decoration: the login check compares the browser's Origin with the Host the app receives, so Apache * must pass the visitor's host on, not "127.0.0.1:3001", or every login would be refused. */ export function vhost(host: string, port: number, certDir: string): string { return ` ServerName ${host} SSLEngine on SSLCertificateFile ${certDir}/fullchain.pem SSLCertificateKeyFile ${certDir}/privkey.pem # Next compresses its pages but not what a route handler sends: the search index (1.0 MB) and the sitemap go out as they are. # Apache compresses those (it leaves alone a response that Next has already compressed). AddOutputFilterByType DEFLATE application/json application/xml text/plain ProxyPreserveHost On RequestHeader set X-Forwarded-Proto "https" ProxyPass / http://127.0.0.1:${port}/ ProxyPassReverse / http://127.0.0.1:${port}/ ServerName ${host} RedirectMatch 301 ^/(.*)$ https://${host}/$1 `; } /** The configuration for every site that has an app, generated from the site map so it cannot disagree with it. */ export function renderVhosts(options: HostOptions): string { const domain = options.domain ?? DOMAIN; const certDir = options.certDir ?? `/etc/letsencrypt/live/${domain}`; const port = options.port ?? ((site) => DEV_PORTS[site]); return options.sites.map((site) => vhost(hostOf(site, domain), port(site), certDir)).join("\n"); } /** A systemd unit that keeps one site's Node process running, and starts it again if it stops. */ export function renderUnit(site: SiteName | "portfolio", options: { root?: string; port?: number; user?: string } = {}): string { const root = options.root ?? "/srv/lw"; const port = options.port ?? DEV_PORTS[site]; return `[Unit] Description=Learning website: ${site} After=network.target [Service] User=${options.user ?? "lw"} WorkingDirectory=${root}/current/apps/${site} EnvironmentFile=/etc/lw/environment ExecStart=/usr/bin/npx next start -p ${port} -H 127.0.0.1 Restart=on-failure RestartSec=3 NoNewPrivileges=true ProtectSystem=strict ReadWritePaths=${root}/data PrivateTmp=true [Install] WantedBy=multi-user.target `; } GENERATED FROM THE SITE MAP so it cannot disagree with it. Only sites that have an app get a virtual host: a site with no app (today, most of them) keeps answering from the old site until its turn. The command: node ops.mjs config C:/lwdata/config portfolio languages wrote lw-sites.conf and 2 unit(s) to C:/lwdata/config: portfolio, languages lw-sites.conf (the first of its four virtual hosts): ServerName osztromok.com SSLEngine on SSLCertificateFile /etc/letsencrypt/live/osztromok.com/fullchain.pem SSLCertificateKeyFile /etc/letsencrypt/live/osztromok.com/privkey.pem # Next compresses its pages but not what a route handler sends: the search index (1.0 MB) and the sitemap go out as they are. # Apache compresses those (it leaves alone a response that Next has already compressed). AddOutputFilterByType DEFLATE application/json application/xml text/plain ProxyPreserveHost On RequestHeader set X-Forwarded-Proto "https" ProxyPass / http://127.0.0.1:3000/ ProxyPassReverse / http://127.0.0.1:3000/ ServerName osztromok.com RedirectMatch 301 ^/(.*)$ https://osztromok.com/$1 lw-languages.service: [Unit] Description=Learning website: languages After=network.target [Service] User=lw WorkingDirectory=/srv/lw/current/apps/languages EnvironmentFile=/etc/lw/environment ExecStart=/usr/bin/npx next start -p 3001 -H 127.0.0.1 Restart=on-failure RestartSec=3 NoNewPrivileges=true ProtectSystem=strict ReadWritePaths=/srv/lw/data PrivateTmp=true [Install] WantedBy=multi-user.target TWO LINES MATTER MORE THAN THEY LOOK. ProxyPreserveHost On: the login check compares the browser's Origin with the Host the app receives. Without that line the app would see "127.0.0.1:3001", every login would be refused as a cross-site request, and everything else would still work, so it could go unnoticed. (The smoke test checks exactly this, with a login from "our own page".) And the app is started with -H 127.0.0.1, so the only way in is through Apache. COMPRESSION, MEASURED. Pages are compressed by Next (a lesson page: 59,774 bytes plain, 12,406 with gzip). What a route handler sends is not: asked with Accept-Encoding: gzip, /search-index.json came back as 1,018,766 bytes, the same as without; so did the 74,907-character sitemap. Compressed by gzip -9 the search index is 326,259 bytes. So the virtual host adds "AddOutputFilterByType DEFLATE application/json application/xml text/plain". I did NOT run Apache, so I have not seen it work, nor seen it leave alone a response Next has already compressed (that is its documented behaviour). THE RELEASE SCRIPT: Save as deploy/deploy.sh: #!/usr/bin/env bash # Roll out a new release, and roll back if it is not healthy. # # deploy.sh release build a new release next to the old one and switch to it # deploy.sh rollback go back to the previous release # # Layout: /srv/lw/releases// one folder per release (code, node_modules and the built apps) # /srv/lw/current a symlink to the live release # /srv/lw/data/accounts.sqlite the accounts database (shared by all releases) # /srv/lw/backups/ database backups # /srv/lw/content/ the content folder (a checkout of the content repository) # /etc/lw/environment LW_ENV=prod, LW_CONTENT_ROOT, LW_DATA_DIR (the same file the services use) # # Not run on a real server while this was written: see the chapter. Everything below has at least been through `bash -n`. set -euo pipefail ROOT=${LW_ROOT:-/srv/lw} SITES=${LW_SITES:-"portfolio languages"} set -a # shellcheck disable=SC1091 . "${LW_ENV_FILE:-/etc/lw/environment}" set +a port_of() { # the same table as DEV_PORTS in packages/sites case "$1" in portfolio) echo 3000 ;; languages) echo 3001 ;; webdevelopment) echo 3002 ;; programming) echo 3003 ;; systems) echo 3004 ;; ai) echo 3005 ;; humanities) echo 3006 ;; lifeskills) echo 3007 ;; creative) echo 3008 ;; *) echo "unknown site $1" >&2; return 1 ;; esac } host_of() { if [ "$1" = portfolio ]; then echo osztromok.com; else echo "$1.osztromok.com"; fi } health_ok() { # real requests to every running site, with the Host header a visitor sends: the app, its database and its content all have to be right for site in $SITES; do node "$1/ops.mjs" smoke "http://127.0.0.1:$(port_of "$site")" "$(host_of "$site")" || return 1 done } restart_all() { for site in $SITES; do sudo systemctl restart "lw-$site"; done } release() { ref=${1:?usage: deploy.sh release } new=$ROOT/releases/$(date +%Y%m%d-%H%M%S) git -C "$ROOT/repo" fetch --quiet git -C "$ROOT/repo" worktree add --detach "$new" "$ref" cd "$new" # 1. a backup BEFORE anything changes: the database is shared, so a bad release can leave it in a state the old code does not expect node ops.mjs backup "$ROOT/backups" --keep 14 # 2. the new code's own checks, with the real settings, before it is built node ops.mjs check npm ci --silent npm test --silent # 3. build every app, with the production addresses baked in for site in $SITES; do npm run build --silent -w "@lw/$site"; done # 4. switch: a symlink change is atomic; the services are restarted from the new release previous=$(readlink -f "$ROOT/current" || true) echo "$previous" > "$ROOT/previous-release" ln -sfn "$new" "$ROOT/current" restart_all sleep 3 # 5. look at the running sites; if they are not healthy, go back at once if ! health_ok "$new"; then echo "The new release failed its smoke test: rolling back." >&2 rollback exit 1 fi echo "Released $new" } rollback() { previous=$(cat "$ROOT/previous-release") [ -d "$previous" ] || { echo "no previous release to go back to" >&2; exit 1; } ln -sfn "$previous" "$ROOT/current" restart_all echo "Back on $previous." echo "The database is only ever created if it is missing (there are no migrations), so the old code normally still matches it." echo "If the release you left changed a table, the old code may not: then restore the backup taken at its start." echo "If accounts misbehave, restore the newest backup (services stopped first):" echo " sudo systemctl stop lw-languages; node $previous/ops.mjs restore $ROOT/backups/; sudo systemctl start lw-languages" } case "${1:-}" in release) shift; release "$@" ;; rollback) rollback ;; *) echo "usage: deploy.sh release | rollback" >&2; exit 2 ;; esac The order is the point: a backup first; the settings checked; install, test and build; then the symlink is switched (atomic), the services restarted, and the smoke test run against EVERY running site; if one check fails the script rolls back by itself. Rolling back is moving the symlink and restarting. Because the database has no migrations (tables are created only when missing), the old code normally still works with it. If a release changed a table, restore the backup taken at its start. WHAT WAS AND WAS NOT RUN. Run, for real: the settings check; the built app started with "next start" and looked at over HTTP; backup and restore; the generated configuration written to disk and read; the tests. NOT run: Apache (no Apache here, so no virtual host, no mod_deflate, no certificate); systemd (this is Windows): the unit uses ProtectSystem=strict with only /srv/lw/data writable, and Next may want to write to its own .next/cache, so the first start on the server may fail until that folder is added to ReadWritePaths; deploy.sh (a Linux shell script with sudo and systemctl), which has only been through "bash -n" for syntax and read through twice; the cron line; the certificates; a real rollback. Only the languages app (and the portfolio's config lines) exist, so the other seven sites have no unit and no virtual host yet. A SMALL PLAN FOR THE FIRST REAL ROLLOUT (the same as the Django project's): 1. install on the server and run the settings check; 2. start the languages unit on its port and run the smoke test against it directly with curl -H "Host: ..."; 3. only then add its virtual host (apachectl configtest, then reload) and run the smoke test over HTTPS from outside; 4. add the RedirectMatch line for its folders to the OLD site (Chapter 11), and watch the logs for 404s; 5. keep the old site's lines in place for a few weeks, so going back is removing one line. WHY THIS WORKS AS AN ANSWER --------------------------- Everything that could be run was run and the output is shown; everything that could not is listed with the specific thing most likely to go wrong first, and the rollout starts with a step that cannot hurt.