Search, Sitemaps and Metadata
Learning Website with Next.js
Chapter 9 ยท Search, Sitemaps & Metadata
A static site has no server to ask, so search has to be a file, and a search engine has to be told, by files and tags, what the site contains. This chapter builds a search that runs in the visitor's browser, a sitemap, one canonical address for every page, and the Open Graph tags that describe a page when it is shared. As in earlier chapters, the results are compared with the Django project on the real content, and two of the differences turned out to be mistakes in the content.
Search That Is a File
There is no search server and no database. When the site is built, an index of every page is written out as one JSON file,
/search-index.json. The browser fetches it the first time somebody searches, then looks words up in it. Three small, tested jobs
make it work:
- htmlToText: the words a person would read on a page (tags, scripts and styles removed, character references decoded).
- buildSearchIndex: for every word, which pages it is on and how often.
- search: a query against the index.
- Words are lower-cased with their accents taken off, so “Szeretnék” and “szeretnek” are one word.
- Japanese has no spaces, so it is stored as single characters, pairs and groups of three, and a query uses the biggest group that fits. (Taking accents off が would turn it into か, so Japanese is cut out first and never folded.)
- A search needs every word; the last may be only the beginning of a word, so “polite cond” finds “conditional”.
- Typed text is never a query language. It is split into units and looked up in tables, so no input can be an error or an injection.
- Ranking is BM25 (a word counts more when it is rare and frequent, less on a long page) with the title weighted ten to one.
- The index is compact: each page number is written as the difference from the one before, in base 36 (
3,7,1:2is pages 3, 10 and 11, the last with the word twice).
The package has two doors, because a browser must never load code that reads files: the main one (used at build time) and
@lw/search/client (prepare and search only), which the page imports.
Is It the Same as Django's?
Step 1: the text
The Django index stores each page's plain text. A fingerprint of each was dumped and compared with the text made here.
| Run | Pages with identical text |
|---|---|
| First | 4,413 of 4,420 (7 different) |
| After three fixes | 4,420 of 4,420 |
The seven pages came from three causes:
</>(a React fragment written in a code sample) stayed as text here; a parser ignores it.⅖(two fifths) was a real HTML name my decoder did not know. (Earlier a survey of every page found 57 names that pages use and the shared decoder did not know; it was filled in, and a lower-case fallback was removed becauseΔandδare different letters.)- Two pages have a mistake in their content: a raw
<style> blockand a raw<script src>written in a sentence, where<style>was meant. An unclosed<style>or<script>swallows the rest of the page, in a browser and in Python's parser. So a browser hides the rest of both pages, and the Django index holds 524 of 54,617 characters of the Build Tooling page. Here an unclosed one swallows the rest too, so the two agree; the real fix is in the two content files (it is in the review notes).
Step 2: the searches
35 queries were put to the Django search and to this one: words, accents, beginnings, Japanese of 1, 2, 3 and 5 characters, all eight sites, and injection-like text.
| What was compared | First version | Final |
|---|---|---|
| The same set of pages | 31 of 35 | 35 of 35 |
| The same first three, in the same order | 5 of 35 (a title-first rule) | 23 of 35, which is 20 of the 32 queries that find anything (BM25) |
The four differences in the sets were: “ありが” found one page too many (with pairs of characters a page that has the first pair in one place and the second in another counts, so groups of three were added to the index), and three queries that found the Build Tooling page, which Django could not (the swallowed text above). The order differs because SQLite and this code compute BM25 a little differently. I stopped there on purpose: matching another engine's order to the last result is not the goal; the same pages, in a sensible order, is.
| Site | Pages | Index | Gzipped |
|---|---|---|---|
| Languages | 561 | 937 KB | 320 KB (273 KB with brotli) |
| Web development | 514 | 966 KB | 338 KB |
| Programming | 1,388 | 2.7 MB | 937 KB |
| Systems | 653 | 1.5 MB | 513 KB |
| Humanities | 707 | 1.5 MB | 561 KB |
A search takes a median of 0.04 ms (at worst 1.7 ms) once the index is loaded. It is fetched once per visit, and only when somebody searches.
The Search Page
A client component reads ?q=, fetches the index once, searches it, and shows 20 results at a time. A plain form in the header goes to it, so it works without any script. The
page is marked noindex (a list of results is not something to find in a search engine). In a headless Chrome, 18 checks passed:
- The same counts as the Django search (18 for “szeretnek” and for “Szeretnék”, 1 for “brotchen”, 10 for “polite cond”, 25 for “ありがとう” and “ありが”, 5 for “水”, 0 for nonsense).
- An injection-like query is a search with no results. HTML typed into the box (an
<img onerror>) is shown as text and nothing runs. - “the” gives 348 results in pages of 20 (the last page shows 8; a page number past the end shows none, without a crash).
- The header box with accented text finds pages; clicking a result moves to it without a reload; the index is fetched once.
- From opening the page to seeing results: 64 to 136 ms after the first search, which took 439 ms because it loads the index.
next start compresses pages (a 107 KB page went out as 20 KB) but not this file, and a generated file has no
ETag, so a browser would download it again on every visit. Two changes: Cache-Control: public, max-age=3600 for the file (after a rebuild a visitor can see an
older index for up to an hour), and compression left to Apache, which would send about 320 KB. The Apache part is not tested (Chapter 12).
A Turbopack build warning (“dynamic filesystem access causes tracing of the whole project”) came from Chapter 8's code that lists a folder's downloads; it only runs while the site is
built, so it is marked to be ignored, and the warning is gone.
One Address for Every Page
The sitemap, the canonical link and the Open Graph address all come from one function, so they cannot disagree. Decision: the canonical form has no trailing slash (Next's way), where the Django project used one. Old addresses with a slash are redirected (308), so no link anyone made is lost, and each page has exactly one address:
The built site was checked page by page (check-seo.mjs reads every built page):
| Check on the 561 languages pages | Result |
|---|---|
Addresses in the sitemap (all different), with a lastmod | 561, of which 250 have a date (the same 250 as Django; no date is invented) |
| An address with non-ASCII letters | percent-encoded: hiragana_%E3%81%82 |
| The page's canonical link is in the sitemap | 561 of 561 |
| Sitemap addresses that are not a page | 0 |
| The description is identical to the one Django stored (the banner's Topic line, or the first 160 characters) | 561 of 561 |
| Open Graph complete (title, description, url equal to the canonical link, type article, site name, card) | 561 of 561 |
robots.txt by default; the search page | Disallow: / (crawling is blocked until launch); noindex |
| Built for the real domain with crawling allowed | addresses begin https://languages.osztromok.com, all 561 still match; Allow: / and a Sitemap: line |
Can the check fail? Because the sitemap and the canonical link normally use the same function, it takes a change to one of them to see it work: giving only the canonical
link a trailing slash made all 561 pages fail. Note that robots.txt is decided when the site is built here (set LW_ALLOW_CRAWLING=1 for the real build); the Django
project decides when it is asked.
og:image: no page has an image, so there is nothing to show (one picture per site would be your decision). Only the languages app has search and a sitemap yet.
The order of results is not Django's. Nothing was run behind Apache, and only Chrome was used.
Hands-On Exercises
Build a search that is a file: turn pages into plain text, build an index of words (and Japanese characters), and search it in the browser without ever treating typed text as a query language. Test the awkward cases and rank by BM25 with the title weighted.
๐ View solutionProve the new search against the Django project's: first the plain text of every page, then 35 real queries. Explain every difference you find, including any that turn out to be errors in the content, and say what you chose not to match.
๐ View solutionAdd the search page, a sitemap, robots.txt, canonical links and Open Graph tags, all from one address function. Check every built page, build for the real domain, and drive the search in a real headless Chrome.
Chapter 9 Quick Reference
- Search is a file:
/search-index.jsonis made at build time and fetched once per visit;@lw/search/clientis the only part the browser loads - Words are lower-cased with accents removed; Japanese is stored as characters, pairs and groups of three; every word must match and the last may be a beginning
- Typed text is split and looked up, never run as a query language; ranking is BM25 with the title weighted ten to one
- The text is identical to Django's on 4,420 of 4,420 pages; 35 of 35 queries find the same pages; the first three agree on 20 of 32 non-empty queries
- An unclosed
<style>or<script>in the text of a page swallows the rest (two real pages): fix the content - Index sizes: languages 937 KB (320 KB gzipped); the largest site 2.7 MB (937 KB); search takes under 2 ms
- One function gives the address for the sitemap, the canonical link and Open Graph; no trailing slash, and old addresses with one redirect (308)
- 561 of 561 canonical links in the sitemap, descriptions identical to Django's, Open Graph complete; 250 pages with a
lastmod robots.txtblocks all untilLW_ALLOW_CRAWLING=1at build time; the search page isnoindex- Not done: highlighted snippets,
og:image, compression of the index (Apache), other sites, other browsers