learning-website-django1-4 Exercise 3: Import the Real Content and Measure It ============================================================================ Run the importer on the real content folder, measure it, and look at the result. Point the project at the content (or pass --root): set LW_CONTENT_ROOT=\content (Linux: export LW_CONTENT_ROOT=...) Then run, on a fresh database: python manage.py migrate python manage.py import_content --dry-run python manage.py import_content python manage.py import_content (nothing changed) python manage.py import_content --force python manage.py import_content --site languages Output (checked by running it on the real content folder; the times are for a folder synced by OneDrive on an ordinary PC, so yours will differ): migrate (content only) Applying content.0001_initial... OK Applying content.0002_page_content_hash_page_fragment_page_imported_at... OK --dry-run DRY RUN (nothing written): created 4403, updated 0, unchanged 0, errors 0 (43.2s) first import created 4403, updated 0, unchanged 0, errors 0 (19.7s) second import, nothing changed created 0, updated 0, unchanged 4403, errors 0 (11.5s) --force created 0, updated 4403, unchanged 0, errors 0 (22.8s) --site languages (everything already imported) created 0, updated 0, unchanged 561, errors 0 (1.2s) database file size: 99 MB Save as import_stats.py in the project folder and run it: """import_stats.py: look at what the import stored. Run from the project folder.""" import os os.environ.setdefault("DJANGO_SETTINGS_MODULE", "config.settings.dev") import django django.setup() from django.db.models import Count, Sum from django.db.models.functions import Length from apps.content.models import Course, Page print(f"pages: {Page.objects.count()} courses: {Course.objects.count()}") for row in Page.objects.values("site").annotate(n=Count("id"), size=Sum(Length("fragment"))).order_by("-n"): print(f" {row['site']:<15}{row['n']:>6} pages {row['size'] / 1e6:>6.1f} MB of HTML") page = Page.objects.light().get(path__endswith="hiragana_あ.html") print("a non-ASCII file name:", page.url_path, "|", page.site, "|", page.kind) course = Course.objects.get(folder="hungary/hungarian-basic-3") print(course.name, "-> chapters", [p.chapter_no for p in course.pages.light()]) print("pages whose stored body still starts with the banner comment:", Page.objects.filter(fragment__startswith="