learning-website-framework1-6 Exercise 1: Port the Fragment Pipeline
=========================================================================
The live site turns a raw content file into a page in a few small steps (in
astro-site/src/lib/transform-html.ts). Re-implement those steps in a language-
neutral way, so the same logic can be reused by Django or Next.js:
1. detect_shape: is the file a complete document (has
) or a fragment?
2. extract_fragment: for a complete document take the contents and keep
any ", re.I | re.S)
LEADING_COMMENT = re.compile(r"^\s*\s*", re.S)
TXT_HREF = re.compile(r'href="([^"]*?)([^"/]+\.txt)"', re.I)
TITLE_TAG = re.compile(r"]*>(.*?)", re.I | re.S)
HEADING = re.compile(r"]*>(.*?)", re.I | re.S)
H1 = re.compile(r"]*>", re.I)
TAGS = re.compile(r"<[^>]+>")
BANNER_CHAPTER = re.compile(r"^[ \t]*Chapter:[ \t]*(.+?)[ \t]*$", re.M)
def detect_shape(raw):
return "wrapped" if BODY_OPEN.search(raw) else "fragment"
def extract_fragment(raw, shape):
if shape == "fragment":
return LEADING_COMMENT.sub("", raw, count=1), False
o, c = BODY_OPEN.search(raw), BODY_CLOSE.search(raw)
if not o or not c or c.start() < o.start():
return raw, True
inner = raw[o.end():c.start()]
head = HEAD.search(raw)
styles = "\n".join(STYLE.findall(head.group(1))) + "\n" if head and STYLE.search(head.group(1)) else ""
return styles + inner, False
def rewrite_solution_links(fragment, course_url_path):
prefix = course_url_path if course_url_path.startswith("/") else "/" + course_url_path
def fix(m):
return f'href="{prefix}/solutions/{unquote(m.group(2))}"'
new, n = TXT_HREF.subn(fix, fragment)
return new, n
def banner_chapter(raw):
m = re.match(r"\s*", raw, re.S)
if not m:
return None
f = BANNER_CHAPTER.search(m.group(1))
return html.unescape(f.group(1)) if f else None
def text_of(m):
if not m:
return None
t = html.unescape(TAGS.sub("", m.group(1)).strip())
return t or None
def extract_title(raw):
return banner_chapter(raw) or text_of(TITLE_TAG.search(raw)) or text_of(HEADING.search(raw))
def process(raw, course_url_path):
shape = detect_shape(raw)
fragment, malformed = extract_fragment(raw, shape)
fragment, n_txt = rewrite_solution_links(fragment, course_url_path)
return {
"shape": shape, "malformed": malformed,
"title": extract_title(raw),
"show_synthetic_h1": not H1.search(fragment),
"txt_links_rewritten": n_txt,
"fragment_bytes": len(fragment.encode("utf-8")),
}
if __name__ == "__main__":
root = sys.argv[1]
for rel in sys.argv[2:]:
raw = open(os.path.join(root, rel), encoding="utf-8").read()
course = "/" + "/".join(rel.split("/")[:-1])
print(rel)
print(" " + str(process(raw, course)))
Run it on four different kinds of file:
python pipeline.py "/claude-projects/website-content/content" hungary/hungarian-basic-3/hungarian_basic_conversation_3_1.html linux/system-administration/debian-development-machine-setup/debian_development_machine_setup_1_1.html japan/japanese-language/reference-materials/kanji/kanji_上.html sidebar/programming/cheat_sheet_git_command_line.html
Output (checked by running it):
hungary/hungarian-basic-3/hungarian_basic_conversation_3_1.html
{'shape': 'fragment', 'malformed': False, 'title': 'Buying Clothes: Sizes, Fit & Returns', 'show_synthetic_h1': True, 'txt_links_rewritten': 0, 'fragment_bytes': 34042}
linux/system-administration/debian-development-machine-setup/debian_development_machine_setup_1_1.html
{'shape': 'fragment', 'malformed': False, 'title': 'Planning & Installing', 'show_synthetic_h1': True, 'txt_links_rewritten': 3, 'fragment_bytes': 23846}
japan/japanese-language/reference-materials/kanji/kanji_上.html
{'shape': 'wrapped', 'malformed': False, 'title': '上 | Kanji | osztromok.com', 'show_synthetic_h1': False, 'txt_links_rewritten': 0, 'fragment_bytes': 15458}
sidebar/programming/cheat_sheet_git_command_line.html
{'shape': 'fragment', 'malformed': False, 'title': 'Git Command Line Cheat Sheet', 'show_synthetic_h1': True, 'txt_links_rewritten': 0, 'fragment_bytes': 27462}
Reading the output
------------------
- The course chapters and the cheat sheet are fragments; the kanji page is a
complete document, so it is unwrapped from its .
- The Debian chapter's three exercise links (.txt) were rewritten; the others
have no solution links.
- The kanji title keeps its " | osztromok.com" suffix. That is fine for the
browser tab but not for an on-page heading, which is why the live site uses
the banner's Chapter: line first.
- Every file gets a synthetic except the kanji page, which already has
its own .
WHY THIS WORKS AS AN ANSWER
---------------------------
The pipeline is only a handful of small, testable functions. Writing them out
without the framework shows that the build logic does not belong to Astro, and
that the behaviour you care about (shape, title, links, headings) can be
checked directly on real files.