Hosting and Deployment
Learning Website: Framework & Architecture
Chapter 8 · Hosting & Deployment
Design and content decide what the sites are. Hosting decides whether they stay up, whether a bad release can be undone, and whether anything private leaks out. The good news is that the multi-site version runs on the same kind of machine you already have: one Debian server running Apache, with a certificate from Certbot. What changes is the arrangement: one virtual host per site, a safer way to publish a build, and a backup plan that covers the things git does not.
website_checks.md.
How It Works Today
As you described it, the content reaches the server in two steps: your PC syncs to a folder under your home on
debserver, and that folder is synced to the web root, so you always hold several copies. The
server runs Apache with one virtual host, taken from the course Setting Up a Web Server on Debian:
Certificates come from Certbot, which proves ownership of a name over HTTP. That is a good foundation. Two weaknesses are worth fixing while you are changing it anyway: a sync that overwrites live files can show visitors a half-copied site and has no way back, and everything sits in one document root.
The Target Layout
Each site gets its own folder and its own virtual host. Inside each folder, the live site is not a normal
directory but a link called current that points at one release folder:
DNS and Virtual Hosts
Each site needs a DNS record pointing at the server (an A record per subdomain, or a wildcard
record; Chapter 1 shows the entries). Apache then chooses the site from the host name in the request, which is
called name-based virtual hosting: several <VirtualHost *:80> blocks share
one address, each with its own ServerName. A request for a name that matches none of them goes to
the first virtual host Apache loads for that address and port, so keep one deliberate default, and let
the existing 000-default.conf play that role, so a mistyped record never shows the wrong site.
Because the virtual hosts differ only by name, generate them from the site map (Exercise 1) rather than copying one by hand:
AllowOverride None is deliberate. The new builds need no .htaccess; every rule
lives in the virtual host, which is faster and cannot be changed by a stray file in the content.
Certificates
There are two sensible ways to cover the subdomains:
| Option | How | Trade-off |
|---|---|---|
| One certificate listing every name | Certbot with a -d for each name, using the same HTTP method you already use | Simple, with no DNS automation. Adding a site means re-running the command with the new name. |
| A wildcard certificate | *.osztromok.com, issued with the DNS challenge | New sites need no certificate change, but you must automate DNS record changes, and a wildcard covers one label only. |
For eight sites the first option is easier: one certificate with ten names (the root, www and
the eight sites). Check that renewal runs on its own (systemctl list-timers should list the
Certbot timer) and that it still works after you add names.
Atomic Deploys and Rollback
The deploy script copies a finished build into a new timestamped release folder, then switches
current to it with one atomic rename. A visitor sees the old site or the new site, never a mixture,
and the previous release stays on disk, so going back is one command:
rsync -a copies folder
times from the build, so it deleted the wrong release. Worse, two deploys in the same second shared a release
name, so a refused deploy emptied and then deleted the live site. The fixes: name releases with
nanoseconds, create them with mkdir without -p so an existing name fails, prune by
name, and never delete the live release. The test now deploys three times, rolls back, deploys again,
and tries an empty build. Test any script that deletes things by trying to make it hurt you.
Backups
You already keep several copies. The new layout changes what each copy has to contain, because two of the most important things are not in git:
| What | Where it lives | In git? |
|---|---|---|
| Content, rules, scripts | website-content and learning-website | Yes (private repos) |
| PDFs | The pdfs/ folders | No (ignored; build outputs) |
| The built sites | /var/www/<site>/releases/ | No (rebuildable) |
| Server configuration | /etc/apache2/sites-available/ | No |
| Certificates and account | /etc/letsencrypt/ | No (and private) |
| Any database (accounts, progress) | MySQL or similar | No |
PDFs are the one that deserves attention: they can be rebuilt from the HTML, but only with the builder scripts and some time, so your copies on the PC, server and backups are effectively their only safe store. Apply the simple 3-2-1 idea: three copies, on two kinds of storage, one of them away from the machine. Then test a restore at least once, because a backup you have never restored is a hope, not a plan.
Secrets Do Not Belong in Content or in the Web Root
Exercise 3 uses two scanners. The first looks for things that look like secrets and is noisy: it reported 66 candidates in 46 files, mostly example passwords in lessons, and it missed three of the five files that really hold a known password. The second searches for an exact string you already know is secret, and found that password in 5 content files, in 5 built pages and in one PHP file in the build. The lessons to take:
- Change the exposed credential first. Once a secret has been published, deleting the text does not make it private again.
- Use obvious placeholders in lessons (
your-db-password-here), even in a course about your own server. - Keep real credentials outside the content tree and the web root, in a file only the account that needs it can read, and outside every git repository.
- Make the scan part of the build so that it fails before publishing, not after.
A Rollout Checklist
- DNS record for the new subdomain, and wait for it to resolve.
- Virtual host generated, enabled,
configtestpassed, Apache reloaded. - Certificate includes the name; HTTPS works with no warning.
- The site's
robots.txtblocks crawling while it is a test site; it allows crawling when it goes live (Learning Website: Framework & Architecture 7). - First deploy through the release script;
currentpoints where expected. - Secrets scan passes; links and pages checked (Learning Website: Framework & Architecture 12).
- Redirects from the old paths switched on last, once the new site is verified.
- Rollback rehearsed once before it is needed.
Hands-On Exercises
Generate an Apache virtual host for each site from the site map, plus the single certbot command covering every name. Check for duplicate names or document roots, balanced tags, valid subdomain labels, and names missing from the certificate command. Say what the check cannot prove.
📄 View solutionWrite deploy.sh and rollback.sh: copy each build into a timestamped release, switch current atomically, refuse a build without index.html, prune old releases safely, and roll back in one command. Write a test that deploys three times, rolls back, deploys again and tries an empty build, and run it.
Write a pattern scanner and an exact-match scanner for secrets. Neither may print the values it finds. Run both on the content tree and the built site, compare what each catches, and write down the steps to take if a real credential is found.
📄 View solutionChapter 8 Quick Reference
- Same stack as today: Debian, Apache, Certbot; one virtual host per site, generated from the site map
- Layout:
/var/www/<site>/releases/<timestamp>/with acurrentlink as the document root - Name-based hosting: unknown names go to the first virtual host loaded, so keep a deliberate default
- Test and reload:
a2ensite,apache2ctl configtest,systemctl reload apache2 - Certificates: one certificate with every name (HTTP method) or a wildcard (DNS method, one label only)
- Deploy atomically:
ln -sfnthenmv -Tf; refuse a build with noindex.html; roll back by moving the link - Prune releases by name, never the live one, and give each release a unique name
- Back up what git does not hold: PDFs, server config, certificates, any database; test a restore
- Secrets: never in content or the web root; rotate anything that was published; scan before every deploy
- A pattern scan is a heuristic (it missed 3 of 5 known files); an exact-match scan finds what you already know