Search and SEO
Learning Website: Framework & Architecture
Chapter 7 ยท Search, Sitemaps & SEO
A site can be well organised and still be hard to find. Visitors need to search it, and search engines need to know what is on it and which address is the real one. When the content moves to several sites, three things become more important: each site must describe itself, the old addresses must lead to the new ones without losing ground, and no page should be listed under two addresses. This chapter looks at what the live site has today, then designs search, sitemaps, canonical addresses and redirects for the family of sites.
What Exists Today
| Feature | On the live site |
|---|---|
| Search | A course-name and tag filter on the Course Index page (client-side). No full-text search of lesson content |
| Sitemap | None: no sitemap.xml in the built output |
| robots.txt | None |
| Canonical link, meta description, social tags | None: the layout's head holds only the character set, viewport and title |
| Redirects | None: the .htaccess only turns off directory listings |
| Size | About 4,900 published pages in the content tree (5,014 directory index pages in the build) |
That is a clean slate rather than a problem: the new sites can include all of this from the start, and none of it has to be migrated.
Search
Search comes in three levels, and each one builds on the one before:
| Level | What it searches | How |
|---|---|---|
| 1. Course finder | Course names and tags in one site | The existing client-side filter, built per site from that site's own courses |
| 2. Full-text search per site | The text of every lesson on one site | A static index generated after the build and searched in the browser; no server needed |
| 3. Search across all sites | Every site at once | Merge the per-site indexes in one search page, usually on the root domain |
A static site search tool such as Pagefind fits the second level: it indexes the built
HTML files after the build and ships the index as static files, so the sites need no search server. For the
third level its documentation describes a multisite feature: a search page loads other sites' indexes with
mergeIndex, given each index's bundlePath, and an indexWeight option
ranks one index above another. Two caveats from the documentation matter here. Indexes on separate domains
must be served with CORS headers, which is the origin issue from Learning Website: Framework
& Architecture 1 again, and a merged index uses the main instance's language support, which matters for
a site with Hungarian, Japanese, German and French text.
Sitemaps
A sitemap is a list of the URLs you want search engines to know about, optionally with the date each page was last changed. The rules that matter, from Google's documentation:
- A single sitemap is limited to 50,000 URLs or 50 MB uncompressed. Larger sites use several files and a sitemap index.
- A sitemap's scope depends on where it lives: unless you submit it through Search Console, it affects only the directory it sits in and below. Put each one at the site root.
- Sitemaps can span several hosts, or a site's sitemap can be hosted centrally with each site's
robots.txtpointing to it. A simple one sitemap per site avoids needing either.
Exercise 1 builds the sitemaps from the real content tree. The result for the proposed sites:
| Site | Pages | With a banner date | Sitemap files |
|---|---|---|---|
| programming | 1,540 | 633 | 1 |
| humanities | 792 | 695 | 1 |
| systems | 739 | 205 | 1 |
| languages | 594 | 250 | 1 |
| webdevelopment | 583 | 53 | 1 |
| lifeskills | 284 | 227 | 1 |
| ai | 216 | 27 | 1 |
| creative | 182 | 20 | 1 |
Every site is far below the limit. The page counts add to slightly more than the 4,925 distinct pages,
because each of the six sites that receives sidebar pages has its own /sidebar/ index page.
Date Updated in their banner (the rule that added it applies
to new and rewritten files only). The sitemap script gives those pages a lastmod and leaves the
rest without one. Do not fill in a date you cannot trust, such as the file's modified time after a
bulk rewrite: a wrong date is worse than no date.
robots.txt, Canonical Links and Descriptions
robots.txt belongs to a host
Google's documentation states that a robots.txt file applies only to the host, protocol and port
that serve it, and that it must sit at the root. So languages.osztromok.com needs its own,
and it is also the natural place to name that site's sitemap:
A new site that is not ready to be public should disallow crawling, and the day it launches that line must change. It is the kind of thing that is forgotten, so it goes on the launch checklist (Learning Website: Framework & Architecture 12).
One address per page
A canonical link tells search engines which address is the real one when a page can be
reached in several ways (with and without www, over HTTP and HTTPS, with and without a trailing
slash, or from an old and a new site). Google recommends an absolute URL, and recommends that a page
include a canonical link to itself. Treat it as a strong signal, not a command. To keep the signals
consistent, make three things agree on the same string: the canonical link, the sitemap entry and the
internal links. That includes the decision about www from Learning Website: Framework &
Architecture 5, and the trailing slash that the site's configuration already requires.
The meta description does not need new writing: the banner's Topic: line is already a written
summary of each chapter, and can be shortened for the tag.
Moving Without Losing Ground
The part of SEO that can really hurt is the move itself. Google's site-move guidance gives a few rules that fit this migration well:
- Use permanent redirects, a 301 or 308, from each old URL to its new URL.
- Keep them a long time: generally at least a year, and indefinitely is better for visitors who have old links.
- A section-by-section move is acceptable for large sites, which matches moving languages first and the rest later.
- Expect weeks, not days: a small or medium site can take a few weeks for most pages to move.
- The Change of Address tool in Search Console applies to domain or subdomain changes (for example one subdomain to another), and is not needed for path-only changes. Whether it applies when a section moves from a path to a new subdomain is a question to check in Search Console at the time.
Redirect each old URL to its exact new URL in a single hop, not to the new site's home page, and not through a chain. Exercise 3 builds the rule for the languages area from the site map and tests it against every old page URL.
What Splitting Does to Search Visibility
Honest expectations help. Google has said it can handle subdomains and paths, but there is no guarantee that a moved section ranks the same afterwards, and for a while the old and new addresses are both known. Reduce the risk instead of trusting it: redirect every old URL, keep internal links consistent, submit each new site's sitemap, watch the Search Console reports for crawl and coverage errors, and move one area at a time so that a problem is found while it is small.
Hands-On Exercises
Write a script that walks content/, decides which files are published pages (skipping print variants, underscore files, pdfs/, solutions/ and the kanji archive), groups the URLs by site, and writes a sitemap per site with lastmod where the banner has a date. Report pages per site, date coverage and the number of sitemap files needed.
Write helpers that produce a page's title, meta description and canonical link, and a site's robots.txt with its Sitemap: line. Make the canonical the same string as the sitemap's <loc>, and reject paths that would break that match.
Build the Apache RedirectMatch rule for the languages area from the site map, and test it against every old page URL in that area: count the matches, confirm each lands on the same path on the new host, and confirm that another area is left alone. Say what this test cannot prove.
Chapter 7 Quick Reference
- Today: a course-name filter only; no full-text search, sitemap, robots.txt, canonical link or meta description
- Search levels: course finder, full-text per site (static index, built after the site), merged search across sites (needs CORS)
- Index the lesson body, not menus or the global bar
- Sitemap limits: 50,000 URLs or 50 MB each; one sitemap per site, at the site root
- Only 42% of pages have a banner date: give
lastmodto those and leave the rest out robots.txtapplies only to its own host, protocol and port: one per site- Canonical links: absolute, self-referential, and the same string as the sitemap and the internal links
- Moves: 301 or 308, one hop, exact URL to exact URL, kept at least a year; section-by-section is fine
- Never copy a page to two sites: link across sites instead
- Block crawling on a site that is not ready, and put reversing it on the launch checklist