Search and SEO

Learning Website: Framework & Architecture

Chapter 7 ยท Search, Sitemaps & SEO

A site can be well organised and still be hard to find. Visitors need to search it, and search engines need to know what is on it and which address is the real one. When the content moves to several sites, three things become more important: each site must describe itself, the old addresses must lead to the new ones without losing ground, and no page should be listed under two addresses. This chapter looks at what the live site has today, then designs search, sitemaps, canonical addresses and redirects for the family of sites.

Where the facts come from
Statements about the live site come from its source and built output. Statements about Google and Pagefind come from their documentation, read while writing this chapter. Search engines change, so treat the behaviour described as guidance and check the current documentation before launch.

What Exists Today

FeatureOn the live site
SearchA course-name and tag filter on the Course Index page (client-side). No full-text search of lesson content
SitemapNone: no sitemap.xml in the built output
robots.txtNone
Canonical link, meta description, social tagsNone: the layout's head holds only the character set, viewport and title
RedirectsNone: the .htaccess only turns off directory listings
SizeAbout 4,900 published pages in the content tree (5,014 directory index pages in the build)

That is a clean slate rather than a problem: the new sites can include all of this from the start, and none of it has to be migrated.

Search

Search comes in three levels, and each one builds on the one before:

LevelWhat it searchesHow
1. Course finderCourse names and tags in one siteThe existing client-side filter, built per site from that site's own courses
2. Full-text search per siteThe text of every lesson on one siteA static index generated after the build and searched in the browser; no server needed
3. Search across all sitesEvery site at onceMerge the per-site indexes in one search page, usually on the root domain

A static site search tool such as Pagefind fits the second level: it indexes the built HTML files after the build and ships the index as static files, so the sites need no search server. For the third level its documentation describes a multisite feature: a search page loads other sites' indexes with mergeIndex, given each index's bundlePath, and an indexWeight option ranks one index above another. Two caveats from the documentation matter here. Indexes on separate domains must be served with CORS headers, which is the origin issue from Learning Website: Framework & Architecture 1 again, and a merged index uses the main instance's language support, which matters for a site with Hungarian, Japanese, German and French text.

Index the content, not the chrome
Menus, footers and the global bar appear on every page. If they are indexed, every search matches every page. Whatever tool you use, mark the lesson body (the content area) as the part to index, and leave the navigation out. Also decide what you do not want searchable, such as the print variants of cheat sheets, which would appear as duplicates.

Sitemaps

A sitemap is a list of the URLs you want search engines to know about, optionally with the date each page was last changed. The rules that matter, from Google's documentation:

  • A single sitemap is limited to 50,000 URLs or 50 MB uncompressed. Larger sites use several files and a sitemap index.
  • A sitemap's scope depends on where it lives: unless you submit it through Search Console, it affects only the directory it sits in and below. Put each one at the site root.
  • Sitemaps can span several hosts, or a site's sitemap can be hosted centrally with each site's robots.txt pointing to it. A simple one sitemap per site avoids needing either.

Exercise 1 builds the sitemaps from the real content tree. The result for the proposed sites:

SitePagesWith a banner dateSitemap files
programming1,5406331
humanities7926951
systems7392051
languages5942501
webdevelopment583531
lifeskills2842271
ai216271
creative182201

Every site is far below the limit. The page counts add to slightly more than the 4,925 distinct pages, because each of the six sites that receives sidebar pages has its own /sidebar/ index page.

Only 42% of pages have a date
Just 2,110 of the 4,925 pages carry a Date Updated in their banner (the rule that added it applies to new and rewritten files only). The sitemap script gives those pages a lastmod and leaves the rest without one. Do not fill in a date you cannot trust, such as the file's modified time after a bulk rewrite: a wrong date is worse than no date.

robots.txt, Canonical Links and Descriptions

robots.txt belongs to a host

Google's documentation states that a robots.txt file applies only to the host, protocol and port that serve it, and that it must sit at the root. So languages.osztromok.com needs its own, and it is also the natural place to name that site's sitemap:

# https://languages.osztromok.com/robots.txt User-agent: * Allow: / Sitemap: https://languages.osztromok.com/sitemap.xml

A new site that is not ready to be public should disallow crawling, and the day it launches that line must change. It is the kind of thing that is forgotten, so it goes on the launch checklist (Learning Website: Framework & Architecture 12).

One address per page

A canonical link tells search engines which address is the real one when a page can be reached in several ways (with and without www, over HTTP and HTTPS, with and without a trailing slash, or from an old and a new site). Google recommends an absolute URL, and recommends that a page include a canonical link to itself. Treat it as a strong signal, not a command. To keep the signals consistent, make three things agree on the same string: the canonical link, the sitemap entry and the internal links. That includes the decision about www from Learning Website: Framework & Architecture 5, and the trailing slash that the site's configuration already requires.

<title>Buying Clothes: Sizes, Fit &amp; Returns</title> <meta name="description" content="Hungarian Basic Conversation 3, chapter 1: buying clothes, sizes, fit and returns."> <link rel="canonical" href="https://languages.osztromok.com/hungary/hungarian-basic-3/hungarian_basic_conversation_3_1/">

The meta description does not need new writing: the banner's Topic: line is already a written summary of each chapter, and can be shortened for the tag.

Moving Without Losing Ground

The part of SEO that can really hurt is the move itself. Google's site-move guidance gives a few rules that fit this migration well:

  • Use permanent redirects, a 301 or 308, from each old URL to its new URL.
  • Keep them a long time: generally at least a year, and indefinitely is better for visitors who have old links.
  • A section-by-section move is acceptable for large sites, which matches moving languages first and the rest later.
  • Expect weeks, not days: a small or medium site can take a few weeks for most pages to move.
  • The Change of Address tool in Search Console applies to domain or subdomain changes (for example one subdomain to another), and is not needed for path-only changes. Whether it applies when a section moves from a path to a new subdomain is a question to check in Search Console at the time.

Redirect each old URL to its exact new URL in a single hop, not to the new site's home page, and not through a chain. Exercise 3 builds the rule for the languages area from the site map and tests it against every old page URL.

Do not copy content to two sites
A page that exists on two sites at once competes with itself, and a canonical link then has to sort it out. The rule from Chapter 2 (one home per page) is also the best SEO rule: link across sites, never copy.

What Splitting Does to Search Visibility

Honest expectations help. Google has said it can handle subdomains and paths, but there is no guarantee that a moved section ranks the same afterwards, and for a while the old and new addresses are both known. Reduce the risk instead of trusting it: redirect every old URL, keep internal links consistent, submit each new site's sitemap, watch the Search Console reports for crawl and coverage errors, and move one area at a time so that a problem is found while it is small.

Hands-On Exercises

Exercise 1

Write a script that walks content/, decides which files are published pages (skipping print variants, underscore files, pdfs/, solutions/ and the kanji archive), groups the URLs by site, and writes a sitemap per site with lastmod where the banner has a date. Report pages per site, date coverage and the number of sitemap files needed.

๐Ÿ“„ View solution
Exercise 2

Write helpers that produce a page's title, meta description and canonical link, and a site's robots.txt with its Sitemap: line. Make the canonical the same string as the sitemap's <loc>, and reject paths that would break that match.

๐Ÿ“„ View solution
Exercise 3

Build the Apache RedirectMatch rule for the languages area from the site map, and test it against every old page URL in that area: count the matches, confirm each lands on the same path on the new host, and confirm that another area is left alone. Say what this test cannot prove.

๐Ÿ“„ View solution

Chapter 7 Quick Reference

  • Today: a course-name filter only; no full-text search, sitemap, robots.txt, canonical link or meta description
  • Search levels: course finder, full-text per site (static index, built after the site), merged search across sites (needs CORS)
  • Index the lesson body, not menus or the global bar
  • Sitemap limits: 50,000 URLs or 50 MB each; one sitemap per site, at the site root
  • Only 42% of pages have a banner date: give lastmod to those and leave the rest out
  • robots.txt applies only to its own host, protocol and port: one per site
  • Canonical links: absolute, self-referential, and the same string as the sitemap and the internal links
  • Moves: 301 or 308, one hop, exact URL to exact URL, kept at least a year; section-by-section is fine
  • Never copy a page to two sites: link across sites instead
  • Block crawling on a site that is not ready, and put reversing it on the launch checklist