Search, Sitemaps and Metadata

Learning Website with Next.js

Chapter 9 ยท Search, Sitemaps & Metadata

A static site has no server to ask, so search has to be a file, and a search engine has to be told, by files and tags, what the site contains. This chapter builds a search that runs in the visitor's browser, a sitemap, one canonical address for every page, and the Open Graph tags that describe a page when it is shared. As in earlier chapters, the results are compared with the Django project on the real content, and two of the differences turned out to be mistakes in the content.

Run for real, on 4,420 pages
76 tests pass; the text of all 4,420 pages is identical to what Django indexed; 35 searches find the same pages as Django's; the built languages site (561 pages) has a canonical address in its sitemap for every page; and a headless Chrome made 18 successful searches and clicks. The order of results is not identical to Django's, there are no highlighted snippets, and nothing was run behind Apache.

Search That Is a File

There is no search server and no database. When the site is built, an index of every page is written out as one JSON file, /search-index.json. The browser fetches it the first time somebody searches, then looks words up in it. Three small, tested jobs make it work:

  1. htmlToText: the words a person would read on a page (tags, scripts and styles removed, character references decoded).
  2. buildSearchIndex: for every word, which pages it is on and how often.
  3. search: a query against the index.
  • Words are lower-cased with their accents taken off, so “Szeretnék” and “szeretnek” are one word.
  • Japanese has no spaces, so it is stored as single characters, pairs and groups of three, and a query uses the biggest group that fits. (Taking accents off が would turn it into か, so Japanese is cut out first and never folded.)
  • A search needs every word; the last may be only the beginning of a word, so “polite cond” finds “conditional”.
  • Typed text is never a query language. It is split into units and looked up in tables, so no input can be an error or an injection.
  • Ranking is BM25 (a word counts more when it is rare and frequent, less on a long page) with the title weighted ten to one.
  • The index is compact: each page number is written as the difference from the one before, in base 36 (3,7,1:2 is pages 3, 10 and 11, the last with the word twice).

The package has two doors, because a browser must never load code that reads files: the main one (used at build time) and @lw/search/client (prepare and search only), which the page imports.

Is It the Same as Django's?

Step 1: the text

The Django index stores each page's plain text. A fingerprint of each was dumped and compared with the text made here.

RunPages with identical text
First4,413 of 4,420 (7 different)
After three fixes4,420 of 4,420

The seven pages came from three causes:

  • </> (a React fragment written in a code sample) stayed as text here; a parser ignores it.
  • &frac25; (two fifths) was a real HTML name my decoder did not know. (Earlier a survey of every page found 57 names that pages use and the shared decoder did not know; it was filled in, and a lower-case fallback was removed because &Delta; and &delta; are different letters.)
  • Two pages have a mistake in their content: a raw <style> block and a raw <script src> written in a sentence, where &lt;style&gt; was meant. An unclosed <style> or <script> swallows the rest of the page, in a browser and in Python's parser. So a browser hides the rest of both pages, and the Django index holds 524 of 54,617 characters of the Build Tooling page. Here an unclosed one swallows the rest too, so the two agree; the real fix is in the two content files (it is in the review notes).

Step 2: the searches

35 queries were put to the Django search and to this one: words, accents, beginnings, Japanese of 1, 2, 3 and 5 characters, all eight sites, and injection-like text.

What was comparedFirst versionFinal
The same set of pages31 of 3535 of 35
The same first three, in the same order5 of 35 (a title-first rule)23 of 35, which is 20 of the 32 queries that find anything (BM25)

The four differences in the sets were: “ありが” found one page too many (with pairs of characters a page that has the first pair in one place and the second in another counts, so groups of three were added to the index), and three queries that found the Build Tooling page, which Django could not (the swallowed text above). The order differs because SQLite and this code compute BM25 a little differently. I stopped there on purpose: matching another engine's order to the last result is not the goal; the same pages, in a sensible order, is.

SitePagesIndexGzipped
Languages561937 KB320 KB (273 KB with brotli)
Web development514966 KB338 KB
Programming1,3882.7 MB937 KB
Systems6531.5 MB513 KB
Humanities7071.5 MB561 KB

A search takes a median of 0.04 ms (at worst 1.7 ms) once the index is loaded. It is fetched once per visit, and only when somebody searches.

Two mistakes of my own, in the tests
I worked a base-36 number out in my head (“dx”; the real answer is “dl”): compute expected values with a tool, not by hand. And my test page for “pieces that are not together” had the exact characters written out, so it correctly matched: build test data so that the thing it is meant to prove is really absent. Both were mistakes in the tests, which the tests then found.

The Search Page

A client component reads ?q=, fetches the index once, searches it, and shows 20 results at a time. A plain form in the header goes to it, so it works without any script. The page is marked noindex (a list of results is not something to find in a search engine). In a headless Chrome, 18 checks passed:

  • The same counts as the Django search (18 for “szeretnek” and for “Szeretnék”, 1 for “brotchen”, 10 for “polite cond”, 25 for “ありがとう” and “ありが”, 5 for “水”, 0 for nonsense).
  • An injection-like query is a search with no results. HTML typed into the box (an <img onerror>) is shown as text and nothing runs.
  • “the” gives 348 results in pages of 20 (the last page shows 8; a page number past the end shows none, without a crash).
  • The header box with accented text finds pages; clicking a result moves to it without a reload; the index is fetched once.
  • From opening the page to seeing results: 64 to 136 ms after the first search, which took 439 ms because it loads the index.
Found on the way
The index is 995 KB on the wire here. next start compresses pages (a 107 KB page went out as 20 KB) but not this file, and a generated file has no ETag, so a browser would download it again on every visit. Two changes: Cache-Control: public, max-age=3600 for the file (after a rebuild a visitor can see an older index for up to an hour), and compression left to Apache, which would send about 320 KB. The Apache part is not tested (Chapter 12). A Turbopack build warning (“dynamic filesystem access causes tracing of the whole project”) came from Chapter 8's code that lists a folder's downloads; it only runs while the site is built, so it is marked to be ignored, and the warning is gone.

One Address for Every Page

The sitemap, the canonical link and the Open Graph address all come from one function, so they cannot disagree. Decision: the canonical form has no trailing slash (Next's way), where the Django project used one. Old addresses with a slash are redirected (308), so no link anyone made is lost, and each page has exactly one address:

/hungary/hungarian-basic-3/hungarian_basic_conversation_3_1/ 308 to /hungary/hungarian-basic-3/hungarian_basic_conversation_3_1 /hungary/hungarian-basic-3/ 308 to /hungary/hungarian-basic-3 /search/ to /search /sitemap.xml/ to /sitemap.xml

The built site was checked page by page (check-seo.mjs reads every built page):

Check on the 561 languages pagesResult
Addresses in the sitemap (all different), with a lastmod561, of which 250 have a date (the same 250 as Django; no date is invented)
An address with non-ASCII letterspercent-encoded: hiragana_%E3%81%82
The page's canonical link is in the sitemap561 of 561
Sitemap addresses that are not a page0
The description is identical to the one Django stored (the banner's Topic line, or the first 160 characters)561 of 561
Open Graph complete (title, description, url equal to the canonical link, type article, site name, card)561 of 561
robots.txt by default; the search pageDisallow: / (crawling is blocked until launch); noindex
Built for the real domain with crawling allowedaddresses begin https://languages.osztromok.com, all 561 still match; Allow: / and a Sitemap: line

Can the check fail? Because the sitemap and the canonical link normally use the same function, it takes a change to one of them to see it work: giving only the canonical link a trailing slash made all 561 pages fail. Note that robots.txt is decided when the site is built here (set LW_ALLOW_CRAWLING=1 for the real build); the Django project decides when it is asked.

What was not done
The results show a page's summary, not the matching words highlighted: the index holds no page text, to stay small (the Django search showed highlighted snippets). There is no og:image: no page has an image, so there is nothing to show (one picture per site would be your decision). Only the languages app has search and a sitemap yet. The order of results is not Django's. Nothing was run behind Apache, and only Chrome was used.

Hands-On Exercises

Exercise 1

Build a search that is a file: turn pages into plain text, build an index of words (and Japanese characters), and search it in the browser without ever treating typed text as a query language. Test the awkward cases and rank by BM25 with the title weighted.

๐Ÿ“„ View solution
Exercise 2

Prove the new search against the Django project's: first the plain text of every page, then 35 real queries. Explain every difference you find, including any that turn out to be errors in the content, and say what you chose not to match.

๐Ÿ“„ View solution
Exercise 3

Add the search page, a sitemap, robots.txt, canonical links and Open Graph tags, all from one address function. Check every built page, build for the real domain, and drive the search in a real headless Chrome.

๐Ÿ“„ View solution

Chapter 9 Quick Reference

  • Search is a file: /search-index.json is made at build time and fetched once per visit; @lw/search/client is the only part the browser loads
  • Words are lower-cased with accents removed; Japanese is stored as characters, pairs and groups of three; every word must match and the last may be a beginning
  • Typed text is split and looked up, never run as a query language; ranking is BM25 with the title weighted ten to one
  • The text is identical to Django's on 4,420 of 4,420 pages; 35 of 35 queries find the same pages; the first three agree on 20 of 32 non-empty queries
  • An unclosed <style> or <script> in the text of a page swallows the rest (two real pages): fix the content
  • Index sizes: languages 937 KB (320 KB gzipped); the largest site 2.7 MB (937 KB); search takes under 2 ms
  • One function gives the address for the sitemap, the canonical link and Open Graph; no trailing slash, and old addresses with one redirect (308)
  • 561 of 561 canonical links in the sitemap, descriptions identical to Django's, Open Graph complete; 250 pages with a lastmod
  • robots.txt blocks all until LW_ALLOW_CRAWLING=1 at build time; the search page is noindex
  • Not done: highlighted snippets, og:image, compression of the index (Apache), other sites, other browsers