learning-website-django1-9 Exercise 3: Sitemaps, Canonical Links and robots.txt ================================================================================= Tell search engines what each site contains, give every page one address, and keep crawlers out until launch. All three are per site, because each site has its own host name. Save as apps/seo/urls.py: """Addresses and the files search engines read: the canonical address of a page, the sitemap and robots.txt. The sitemap's address for a page and the page's canonical link are built by the SAME function, so they cannot disagree (Learning Website: Framework & Architecture 7).""" from urllib.parse import quote def page_location(url_path): """The path part as it must appear in a URL: non-ASCII characters percent-encoded.""" return quote(url_path) def canonical_url(request, url_path): """Absolute address of a page on the site being asked: scheme and host come from the request (behind Apache, SECURE_PROXY_SSL_HEADER makes the scheme https).""" return request.build_absolute_uri(page_location(url_path)) Save as apps/seo/sitemaps.py: from django.contrib.sitemaps import Sitemap from apps.content.models import Page from .urls import page_location class PageSitemap(Sitemap): """Every page of ONE site. Django's sitemap view adds the scheme and host of the request.""" changefreq = None # a guess search engines ignore, so leave it out priority = None limit = 50000 # the protocol's limit per file; Django splits into pages beyond it def __init__(self, site): self.site = site def items(self): return Page.objects.light().filter(site=self.site).order_by("path") def location(self, page): return page_location(page.url_path) def lastmod(self, page): return page.updated # only pages whose banner has a date get one: never invent a date Save as apps/seo/views.py: from django.conf import settings from django.http import HttpResponse def robots_txt(request, site): """robots.txt belongs to a host, so each site answers with its own. Crawling stays blocked until ALLOW_CRAWLING is switched on, which is a line on the launch checklist.""" if settings.ALLOW_CRAWLING: body = f"User-agent: *\nAllow: /\n\nSitemap: {request.build_absolute_uri('/sitemap.xml')}\n" else: body = "User-agent: *\nDisallow: /\n" return HttpResponse(body, content_type="text/plain; charset=utf-8") The site URL configuration adds three routes before the page route: ^search/$ search results ^sitemap\.xml$ django.contrib.sitemaps view, with {"pages": PageSitemap(site)} ^robots\.txt$ robots_txt The page template gets and from the page's own address and its banner Topic line (stored in Page.summary, or the first 160 characters of text when a banner has no Topic). Settings: ALLOW_CRAWLING = os.environ.get("LW_ALLOW_CRAWLING") == "1" (off by default) The one rule that matters: the sitemap address and the canonical address come from the same function, so they cannot disagree. A sitemap that lists one address while the page says another is worse than no sitemap. Save as tests/test_seo.py: import re import xml.etree.ElementTree as ET from urllib.parse import quote from django.test import Client, TestCase, override_settings from apps.content.models import Page NS = {"s": "http://www.sitemaps.org/schemas/sitemap/0.9"} def make(path, site="languages", title="T", updated=None, summary=""): return Page.objects.create(path=path, site=site, kind="lesson", title=title, fragment="

x

", updated=updated, summary=summary) class SitemapTests(TestCase): def setUp(self): from datetime import date make("hungary/a.html", updated=date(2026, 10, 5)) make("hungary/b.html") # no date in its banner make("japan/japanese-language/reference-materials/hiragana/hiragana_あ.html") make("linux/c.html", site="systems") self.client = Client(HTTP_HOST="languages.localhost") def urls(self): response = self.client.get("/sitemap.xml") self.assertEqual(response.status_code, 200) self.assertIn("xml", response["Content-Type"]) return ET.fromstring(response.content).findall("s:url", NS) def test_the_sitemap_lists_this_sites_pages_only(self): locs = [u.find("s:loc", NS).text for u in self.urls()] self.assertEqual(len(locs), 3) self.assertTrue(all("languages.localhost" in loc for loc in locs)) self.assertFalse(any("linux" in loc for loc in locs)) def test_a_date_appears_only_where_the_page_has_one(self): by_loc = {u.find("s:loc", NS).text: u.find("s:lastmod", NS) for u in self.urls()} dated = [loc for loc, node in by_loc.items() if node is not None] self.assertEqual(len(dated), 1) self.assertTrue(dated[0].endswith("/hungary/a/")) self.assertEqual(by_loc[dated[0]].text, "2026-10-05") def test_the_address_uses_the_pages_own_form_with_a_trailing_slash(self): self.assertIn("http://languages.localhost/hungary/b/", [u.find("s:loc", NS).text for u in self.urls()]) def test_non_ascii_addresses_are_percent_encoded(self): locs = [u.find("s:loc", NS).text for u in self.urls()] self.assertTrue(any(loc.endswith("hiragana_%E3%81%82/") for loc in locs)) self.assertFalse(any("あ" in loc for loc in locs)) def test_every_sitemap_address_is_the_pages_canonical_link(self): sitemap = {u.find("s:loc", NS).text for u in self.urls()} for page in Page.objects.filter(site="languages"): html = self.client.get(quote(page.url_path)).content.decode() canonical = re.search(r'', html) self.assertIn('', html) def test_a_page_without_a_summary_has_no_empty_description(self): make("hungary/x/b.html") html = Client(HTTP_HOST="languages.localhost").get("/hungary/x/b/").content.decode() self.assertNotIn('name="description"', html) def test_listings_and_the_front_page_have_no_canonical_link_yet(self): make("hungary/x/b.html") html = Client(HTTP_HOST="languages.localhost").get("/hungary/x/").content.decode() self.assertNotIn("canonical", html) Test run (whole project): Found 217 test(s). System check identified no issues (0 silenced). Creating test database for alias 'default'... ......................................................................................................................................................................................................................... ---------------------------------------------------------------------- Ran 217 tests in 4.326s OK Destroying test database for alias 'default'... Real sitemaps, canonical links and robots.txt, on the real database: languages sitemap 200 561 pages 561 unique 561 dated 250 44ms 69 KB sampled 75 bad 0 User-agent: * | Disallow: / | programming sitemap 200 1376 pages 1376 unique 1376 dated 654 238ms 189 KB sampled 60 bad 0 User-agent: * | Disallow: / | ai sitemap 200 187 pages 187 unique 187 dated 27 29ms 21 KB sampled 60 bad 0 User-agent: * | Disallow: / | ("sampled" pages were fetched, including pages with Japanese in the address, and each page's canonical link was checked to be in that site's sitemap. "dated" is how many entries have a lastmod: only pages whose banner has a date get one; none is invented.) What was not verified: nothing was submitted to a search engine, and robots.txt was only read with ALLOW_CRAWLING off in the real run (the on case is covered by tests). WHY THIS WORKS AS AN ANSWER --------------------------- Everything is checked on real data: sitemap counts equal page counts per site, no duplicates, non-ASCII addresses are percent-encoded, and the canonical link of every sampled page is in the sitemap.