learning-website-django1-3 Exercise 3: A Dry Run Over the Whole Content Tree ============================================================================= Before importing anything, run the parser over every real page and see what it finds. Nothing is written to the database. The script reports the counts and anything that looks wrong. Save as scan_content.py in the project folder: """scan_content.py : a dry run of the importer. Parses every page and reports what it found and what looks wrong. Nothing is written to the database.""" import os import sys from collections import Counter, defaultdict os.environ.setdefault("DJANGO_SETTINGS_MODULE", "config.settings.dev") import django django.setup() from apps.content.banner import parse_banner from apps.content.parsing import parse_page from config.sites_config import NoSiteError EXCLUDED = ("japan/japanese-language/reference-materials/kanji",) # the archive copy, not routed SKIP_DIRS = {"pdfs", "solutions"} def pages(root): for dirpath, dirnames, filenames in os.walk(root): rel_dir = os.path.relpath(dirpath, root).replace(os.sep, "/") rel_dir = "" if rel_dir == "." else rel_dir if any(rel_dir == p or rel_dir.startswith(p + "/") for p in EXCLUDED): dirnames[:] = [] continue dirnames[:] = [d for d in dirnames if d not in SKIP_DIRS] for name in filenames: if name.endswith(".html") and not name.endswith("_print.html") and not name.startswith("_"): yield (rel_dir + "/" + name) if rel_dir else name root = sys.argv[1] kinds, sites = Counter(), Counter() no_site, weak_title, file_mismatch, others = [], [], [], [] course_names = defaultdict(set) chapters = defaultdict(list) total = 0 for rel in pages(root): total += 1 raw = open(os.path.join(root, *rel.split("/")), encoding="utf-8", errors="ignore").read() try: p = parse_page(raw, rel) except NoSiteError: no_site.append(rel) continue kinds[p.kind] += 1 if p.kind == "other": others.append(rel) sites[p.site] += 1 if p.title == " ".join(w.capitalize() for w in rel.rsplit("/", 1)[-1][:-5].replace("-", "_").split("_") if w): weak_title.append(rel) banner = parse_banner(raw[:6000]) if banner.get("file") and banner["file"] != rel.rsplit("/", 1)[-1]: file_mismatch.append(rel) if p.kind == "course_chapter": course_names[p.course_folder].add(p.course_name) chapters[p.course_folder].append(p.chapter_no) unnumbered = {f: c.count(None) for f, c in chapters.items() if None in c} gaps = {f: sorted(n for n in c if n is not None) for f, c in chapters.items() if sorted(n for n in c if n is not None) != list(range(1, len([n for n in c if n is not None]) + 1))} renamed = {f: n for f, n in course_names.items() if len(n) > 1} print(f"pages parsed: {total - len(no_site)} of {total}") print("by kind: " + ", ".join(f"{k} {v}" for k, v in kinds.most_common())) print("by site: " + ", ".join(f"{k} {v}" for k, v in sites.most_common())) print(f"courses found (folders of numbered chapters): {len(chapters)}") print(f"paths that belong to no site: {len(no_site)}") for x in no_site[:5]: print(" ", x) print(f"pages with no banner and no complete-page title (kind other): {len(others)}") for x in others[:7]: print(" ", x) print(f"pages whose title fell back to the file name: {len(weak_title)}") for x in weak_title[:3]: print(" ", x) print(f"banner File: line does not match the real file name: {len(file_mismatch)}") for x in file_mismatch[:3]: print(" ", x) print(f"course pages with no chapter number in the file name: {sum(unnumbered.values())} (in {len(unnumbered)} folders)") for f, n in list(unnumbered.items())[:3]: print(f" {f}: {n} page(s)") print(f"courses whose numbered chapters are not 1..N: {len(gaps)}") for f, c in list(gaps.items())[:4]: print(f" {f}: {c[:14]}{'...' if len(c) > 14 else ''}") print(f"course folders with more than one Course: name: {len(renamed)}") for f, n in list(renamed.items())[:3]: print(f" {f}: {sorted(n)}") Run it with the path to your content folder: python scan_content.py "/content" Output (checked by running it on the real content folder): pages parsed: 4402 of 4402 by kind: course_chapter 3895, full_page 410, sidebar 47, lesson 43, other 7 by site: programming 1370, humanities 707, systems 653, languages 561, webdevelopment 514, lifeskills 253, ai 187, creative 157 courses found (folders of numbered chapters): 395 paths that belong to no site: 0 pages with no banner and no complete-page title (kind other): 7 hungary/hungarian-lessons/hungarian_lesson_past_definite_tense.html hungary/hungarian-lessons/hungarian_lesson_past_indefinite_tense.html japan/japanese-language/japanese-lessons/japanese_lesson_family_members.html japan/japanese-language/japanese-lessons/japanese_lesson_food_likes.html japan/japanese-language/reference-materials/hiragana/hiragana_tiles_cms_content.html japan/japanese-language/reference-materials/katakana/katakana_tiles_cms_content.html web-development/scripting-and-backend/php/fundamentals/php1-1.html pages whose title fell back to the file name: 1 hungary/hungarian-lessons/hungarian_alphabet.html banner File: line does not match the real file name: 7 projects/web-server-on-debian/setting_up_a_web_server_on_debian_01.html projects/web-server-on-debian/setting_up_a_web_server_on_debian_02.html projects/web-server-on-debian/setting_up_a_web_server_on_debian_03.html course pages with no chapter number in the file name: 3 (in 2 folders) linux/shell-and-scripting/vim/learning-vim: 1 page(s) linux/system-administration/linux-installation-and-configuration: 2 page(s) courses whose numbered chapters are not 1..N: 1 web-development/scripting-and-backend/php/fundamentals: [2, 3, 4, 5, 6, 7, 8, 9, 10] course folders with more than one Course: name: 0 What the dry run found ---------------------- - All 4,402 pages parse, and none belongs to no site: the site map from Chapter 2 of the framework course is complete. (The kanji archive folder, _print variants, underscore files, pdfs and solutions are skipped, as on the live site.) - The kinds: 3,895 course chapters, 410 complete pages (the kanji pages), 47 sidebar pages, 43 standalone lessons, and 7 with no banner at all. - 395 course folders contain numbered chapters. No folder uses two different Course: names. - The first version of the parser failed on 21 pages and mis-numbered the HTML course, because the file names have several generations. Adding the _lesson_NN_C form (chapter first) and the hyphen form (psp-croll-1-2) took it to 3 un-numbered pages (appendices). - Seven pages have no banner: four older lessons, the hiragana and katakana tiles fragments, and php1-1.html (which is why PHP Fundamentals appears to start at chapter 2). They are content problems to decide on, not parser problems: add a banner to each. - Seven banners name a different file in their File: line than the real file (all in the Setting Up a Web Server on Debian course). - One page falls back to its file name for a title: hungarian_alphabet.html (no banner title and no heading). WHY THIS WORKS AS AN ANSWER --------------------------- A dry run turns "the importer might hit surprises" into a list. It found a whole family of file names the first parser did not know, and it listed seven real content problems, all before anything was imported.