learning-website-framework1-7 Exercise 1: Generate a Sitemap per Site ========================================================================= Walk the content folder, decide which files are real published pages, group the page URLs by site, and write one sitemap per site. Use the Date Updated line from each banner as where one exists. Rules the script applies (all from the live site's own routing): - only .html files; skip _print variants and files starting with an underscore - skip pdfs/ and solutions/ folders - skip the kanji archive folder (it is excluded from the live routing) - a chapter URL is the path without .html, with a trailing slash - every folder above a published page is a real directory page too - /sidebar/ has no single owner: each site that receives sidebar pages has its own /sidebar/ index - Google's limits: 50,000 URLs or 50 MB per sitemap file Needs sitemap.py from Learning Website: Framework & Architecture 5. Save as sitemaps.py: import os, re, sys from collections import Counter, defaultdict from xml.sax.saxutils import escape from sitemap import site_for, SIDEBAR_ROUTES EXCLUDED = ("japan/japanese-language/reference-materials/kanji",) UPDATED = re.compile(r"Date Updated:\s*(\d{4}-\d{2}-\d{2})") LIMIT = 50_000 # URLs per sitemap file (Google's limit) def is_excluded(rel): return any(rel == p or rel.startswith(p + "/") for p in EXCLUDED) def collect(content): """Return {url_path: lastmod_or_None} for every page that should be published.""" pages = {} for dirpath, dirnames, filenames in os.walk(content): rel_dir = os.path.relpath(dirpath, content).replace(os.sep, "/") rel_dir = "" if rel_dir == "." else rel_dir if rel_dir and is_excluded(rel_dir): dirnames[:] = [] continue dirnames[:] = [d for d in dirnames if d not in ("pdfs", "solutions")] for name in filenames: if not name.endswith(".html") or name.endswith("_print.html") or name.startswith("_"): continue rel = f"{rel_dir}/{name}" if rel_dir else name with open(os.path.join(dirpath, name), "rb") as fh: head = fh.read(1500).decode("utf-8", "ignore") m = UPDATED.search(head) pages["/" + rel[:-5] + "/"] = m.group(1) if m else None # the directory pages above it are real pages too parts = rel.split("/")[:-1] for i in range(1, len(parts) + 1): pages.setdefault("/" + "/".join(parts[:i]) + "/", None) return pages def by_site(pages): groups = defaultdict(dict) for url, lastmod in pages.items(): if url == "/sidebar/": # the sidebar index has no single owner: each site that receives # sidebar pages gets its own /sidebar/ index page for site in set(SIDEBAR_ROUTES.values()): groups[site][url] = lastmod continue try: groups[site_for(url)][url] = lastmod except KeyError: groups["(unmapped)"][url] = lastmod return groups def sitemap_xml(site, urls): rows = [] for url, lastmod in sorted(urls.items()): loc = escape(f"https://{site}.osztromok.com{url}") lm = f"{lastmod}" if lastmod else "" rows.append(f" {loc}{lm}") return ('\n' '\n' + "\n".join(rows) + "\n\n") if __name__ == "__main__": pages = collect(sys.argv[1]) groups = by_site(pages) total = len(pages) with_date = sum(1 for v in pages.values() if v) print(f"pages: {total} with a banner Date Updated: {with_date} ({100*with_date//total}%)") print(f"{'site':<16}{'pages':>7}{'with date':>11}{'files needed':>14}") for site, urls in sorted(groups.items(), key=lambda kv: -len(kv[1])): d = sum(1 for v in urls.values() if v) print(f"{site:<16}{len(urls):>7}{d:>11}{-(-len(urls)//LIMIT):>14}") sample_site = "languages" sample = dict(list(sorted(groups[sample_site].items()))[:3]) print() print(sitemap_xml(sample_site, sample)) Run it: python sitemaps.py "/claude-projects/website-content/content" Output (checked by running it on the real folder): pages: 4925 with a banner Date Updated: 2110 (42%) site pages with date files needed programming 1540 633 1 humanities 792 695 1 systems 739 205 1 languages 594 250 1 webdevelopment 583 53 1 lifeskills 284 227 1 ai 216 27 1 creative 182 20 1 https://languages.osztromok.com/culture/ https://languages.osztromok.com/culture/japan/ https://languages.osztromok.com/culture/japan/japanese-film-and-television/ Reading the output ------------------ - 4,925 distinct page URLs. (The live build contains about 5,014 directory index pages; the small difference is pages this simple walk does not model, such as the course index, combined exports and utility pages. It was not traced further.) The per-site counts add up to more than the total because /sidebar/ is counted once for each of the six sites that receive sidebar pages. - Only 42% of pages (2,110) have a Date Updated in their banner. The rest get no at all. That is deliberate: leave a date out rather than invent one, because a wrong date is worse than none. - Every site is far below the 50,000-URL limit (the largest, programming, has 1,540), so each needs one file and no sitemap index. WHY THIS WORKS AS AN ANSWER --------------------------- A sitemap is generated from the same files and the same site map as everything else, so it cannot disagree with the build. Measuring the date coverage shows where the content metadata is thin, which is useful in its own right.