Content Pipeline
Learning Website: Framework & Architecture
Chapter 6 ยท The Content Pipeline
Chapter 3 described the files. This chapter follows one file on its journey to a finished page: how the build reads it, cleans it, finds its title, fixes its links, wraps it in a layout and publishes the assets next to it. The good news is that this pipeline already exists and works. The live Astro site implements it in a few small files. The job here is to describe it clearly enough that the Django and Next.js versions can behave identically, because the content is shared and must look the same whichever framework builds it.
transform-html.ts,
content-fs.ts, site-tree.ts, the page and layout files and the asset-copy script.
The exercises re-implement the stages in Python and run them on real files.
The Stages
| # | Stage | What it does |
|---|---|---|
| 1 | Walk and classify | Visit every file under content/ and label it: chapter HTML, PDF, solution .txt, or other |
| 2 | Detect the shape | A file with a <body> tag is a complete document; anything else is a fragment |
| 3 | Extract the fragment | Complete document: take the body contents and keep the <style> blocks from the head. Fragment: remove the leading banner comment |
| 4 | Rewrite solution links | Point every .txt link at <course>/solutions/<file> |
| 5 | Find the title | Banner Chapter:, else the <title> tag, else the first <h1> or <h2>, else the file name made readable |
| 6 | Decide the heading | Add a synthetic <h1> only if the fragment has none of its own |
| 7 | Wrap in the layout | Inject the fragment into the page layout, with breadcrumb and previous/next links from the sorted sibling chapters |
| 8 | Activate scripts | Re-create any <script> inside the injected content so that it runs; provide the code-copy function once, globally |
| 9 | Publish assets | Mirror every pdfs/ and solutions/ folder into the output at the same relative path |
| 10 | Validate | Check the assumptions against every real file before the templates rely on them |
Shape and Extraction
Nearly all files are fragments, and the kanji pages are complete documents (Learning Website: Framework &
Architecture 3). The pipeline must handle both, so it tests for a <body> tag and
extracts accordingly. For a complete document it keeps the head's <style> blocks and
discards the rest of the head, so the styles are not lost. A document with an opening but no closing body tag is
marked malformed so the build can complain.
Titles and Headings
The page title comes from the banner's Chapter: line first. This is deliberate: numbered-course
chapters have no real heading of their own (they show a styled course name and a breadcrumb line instead), so
the banner is the only place that holds the snappy chapter title. When there is no such line, the pipeline
falls back to the document's own <title>, then to the first heading, then to the filename.
The fallback order matters. A kanji page's <title> ends in “| osztromok.com”,
which is right for the browser tab but wrong for a page heading.
A synthetic <h1> is added only when the fragment has no <h1>, so
nothing is ever titled twice.
Solution Links
Old chapters link to their solution files in many stale ways: absolute paths from before the folders were
reorganised, relative variants, a reference to an old backup location, or a bare filename. Instead of
listing every variant, the pipeline matches any .txt link and rebuilds it from its file name:
<course path>/solutions/<file name>. That is robust, with one condition: the file
names inside one course's solutions/ folder must be unique and must exist there.
solutions/ folder and 10,975 solution links,
seven courses have a problem. JavaScript Fundamentals has 60 links but only 3 solution files.
TypeScript Fundamentals and the three Practical Scripting projects have solution files that no chapter links
to. OWASP Top 10 and Excel Fundamentals each have a few links with no file. They are listed in
website_checks.md; none has been fixed.
Scripts, and the Copy Button
Fragments are injected into the page as HTML text. Browsers do not run <script> tags that
arrive that way, so the pipeline has two measures, both in the layout:
- The code-copy function exists once, globally. Each code block carries
onclick="copyCodeBlock(this)", and the function lives in a normal script in the layout. New chapters must not embed their own copy. - Other scripts are re-created. When the page loads, every
<script>inside the content area is replaced by a freshly created one with the same attributes and text, which does run. That is what makes the interactive My Tools pages work.
Two details are easy to miss. The function relies on the code block being the element
immediately after the button, so the markup order is part of the contract, and any new framework must
keep it. And navigator.clipboard only works in a secure context, meaning HTTPS or
localhost. Test copying on the real HTTPS subdomains, not just on a development address.
Ordering, Neighbours and Assets
The live site sorts chapters with a natural sort, so chapter 2 comes before chapter 10 (the trap from Learning Website: Framework & Architecture 3). Previous and next links are the neighbours in that sorted list. The new builds must use the same ordering, or the chapter links will disagree between sites.
PDFs and solution files are not part of the page. A small step before the build mirrors every
pdfs/ and solutions/ folder into the published output at the same relative
path, which is why the rewritten links work. It copies a file only when it is missing or newer, so the
second run does almost nothing (Exercise 3).
PDFs are build outputs
Each PDF is generated from the chapter HTML by a builder script, and the repository ignores
*.pdf. That keeps git small, but it has a consequence: the asset-copy step reads PDFs from the
working folder, not from git. On a new machine or a build server they will be missing until you regenerate
them or sync them from where they are kept. Decide which, document it, and add a check that every course
which links to a PDF actually has it.
Validate Before the Templates Rely on It
The live site has a separate script that runs the transform over every real file and reports shapes, malformed documents and unresolved solution links before any template uses them. Keep that habit. A pipeline check that fails the build is cheap, and the alternative is a visitor finding the broken link for you. In the multi-site version run the same validation per site, and once more across all sites together.
Hands-On Exercises
Port the transform stages (detect shape, extract fragment, rewrite solution links, find the title, decide the synthetic heading) to Python without any framework. Run them on a course chapter, a chapter with solution links, a complete kanji page and a cheat sheet, and explain each result.
๐ View solutionWrite an audit that checks every .txt link in a course against that course's solutions/ folder, and reports links with no file and files no chapter links to. Run it on four courses, then on every course with a solutions/ folder, and summarise the problems.
Write an idempotent asset-copy step that mirrors every pdfs/ and solutions/ folder to an output folder at the same relative path, copying only when a file is missing or newer. Run it twice and then after touching one file, and explain what it shows. Say what you would do about PDFs on a machine that has no copies.
Chapter 6 Quick Reference
- Stages: walk, shape, extract, solution links, title, heading, layout, scripts, assets, validate
- A file with
<body>is a complete document; everything else is a fragment - Title order: banner
Chapter:, then<title>, then first h1/h2, then the file name - Add a synthetic
<h1>only if the fragment has none - Any
.txtlink is rewritten to<course>/solutions/<file name>; the file must exist there - The audit found 7 of 268 courses with solution problems (10,975 links checked)
copyCodeBlocklives once in the layout and needs the code block right after the button; clipboard needs HTTPS or localhost- Scripts inside content are re-created on page load so that they run
- Sort chapters naturally (2 before 10); previous/next are the neighbours in that order
- Mirror
pdfs/andsolutions/at the same relative path, copying only if missing or newer - PDFs are build outputs, ignored by git: regenerate or sync them, and check every linked PDF exists