learning-website-framework1-7 Exercise 2: Canonicals, Descriptions and robots.txt =================================================================================== Write helpers that produce, for any page, the head tags (title, description, canonical link) and, for any site, its robots.txt. Make the canonical URL the same string as the sitemap's for that page. Facts this relies on (Google Search Central documentation): - robots.txt rules apply only to the host, protocol and port that serve the file, so every subdomain needs its own robots.txt at its root. - A canonical link should be absolute, and a page should link to itself. Save as seo_tags.py (needs sitemap.py): import html from sitemap import site_for def origin(site): return f"https://{site}.osztromok.com" def canonical(site, url_path): """Absolute, self-referencing canonical URL; same string the sitemap uses.""" assert url_path.startswith("/") and url_path.endswith("/"), "paths are root-relative and end in /" return origin(site) + url_path def robots_txt(site): # robots.txt applies only to the host that serves it: every site needs its own return (f"User-agent: *\nAllow: /\n\n" f"Sitemap: {origin(site)}/sitemap.xml\n") def head_tags(site, url_path, title, description): url = canonical(site, url_path) return "\n".join([ f"{html.escape(title)}", f'', f'', ]) if __name__ == "__main__": path = "/hungary/hungarian-basic-3/hungarian_basic_conversation_3_1/" site = site_for(path) print(head_tags(site, path, "Buying Clothes: Sizes, Fit & Returns", "Hungarian Basic Conversation 3, chapter 1: buying clothes, sizes, fit and returns.")) print() print(robots_txt(site)) # the canonical must equal the URL the sitemap lists for the same page sitemap_loc = f"https://{site}.osztromok.com{path}" print("canonical matches sitemap :", canonical(site, path) == sitemap_loc) for bad in ("/hungary/page", "hungary/page/"): try: canonical(site, bad) except AssertionError as e: print(f"rejected {bad!r}: {e}") Run it: python seo_tags.py Output (checked by running it): Buying Clothes: Sizes, Fit & Returns User-agent: * Allow: / Sitemap: https://languages.osztromok.com/sitemap.xml canonical matches sitemap : True rejected '/hungary/page': paths are root-relative and end in / rejected 'hungary/page/': paths are root-relative and end in / Notes ----- - The tags must be HTML-escaped (the "&" in the title became &). - The helper refuses a path without a leading and trailing slash, because a canonical that differs from the sitemap URL (with or without the slash) is a classic source of mixed signals. - The live site's layout currently has no canonical link, no meta description and no robots.txt or sitemap (checked in the layout file and the built output). The new layouts should add them. - Until a site is ready to be public, its robots.txt should disallow crawling, and launching means remembering to change that. Put it on the launch checklist (Learning Website: Framework & Architecture 12). WHY THIS WORKS AS AN ANSWER --------------------------- Search tags are easy to get subtly wrong by hand. Generating them from the same function that builds the sitemap URL makes the canonical and the sitemap agree by construction, and the assertions catch the common slash mistake.