Building a Template Engine: Parsing, Compiling & Rendering

Building a Web Framework

Chapter 4 ยท Building a Template Engine: Parsing, Compiling & Rendering

Chapter 3 closed on a real handler shape this course hadn't built yet: Response(render('page.html', title='Home')). Every handler written so far builds its own HTML with an f-string, which works for one line and falls apart the moment a real page needs a shared layout, a loop over real data, or a value nobody has separately checked for safety. This chapter builds that missing piece — a real, file-based TemplateEngine class this framework keeps and reuses from here on — and runs directly into the one design decision a template engine can't get wrong: what happens to a value the instant it lands inside real HTML.

The Basic Job: Substituting Into a Real Template File

A template engine's simplest possible form finds placeholders in a string and replaces them with real values — the same regex-substitution shape Chapter 2's own router used for dynamic path segments, applied to a whole document loaded from disk instead of a URL:

templates/basic.html
Hello, {{ name }}!
import re def render_basic(path, context): with open(path, 'r', encoding='utf-8') as f: template = f.read() return re.sub( r'\{\{\s*(\w+)\s*\}\}', lambda m: str(context.get(m.group(1), '')), template, ) print(render_basic('templates/basic.html', {'name': 'World'})) # Hello, World!

Genuinely working, genuinely unsafe the moment a real, user-controlled value reaches it.

The Real Security Problem: Naive Interpolation

malicious = {'name': '<script>alert(1)</script>'} print(render_basic('templates/basic.html', malicious)) # Hello, <script>alert(1)</script>!
Verified directly — a real, live <script> tag reaches the output intact
Loaded straight from a real .html file on disk and rendered with a real malicious context value, the output contains a genuine, executable <script> tag — not the text describing one. A profile display name, a comment, a search query echoed back — anything reaching render_basic() unescaped becomes a real, working cross-site scripting vector the instant it's rendered into an actual page.

Auto-Escaping by Default, With an Explicit Opt-Out

Every value gets HTML-escaped automatically unless it's deliberately marked as already safe — a plain str subclass is enough to carry that marker:

import html class SafeString(str): # its only job is to be recognizable as "already safe" -- nothing else changes pass def escape(value): if isinstance(value, SafeString): return str(value) return html.escape(str(value))
Verified directly — identical text, two genuinely different outcomes
Rendering the exact same malicious string through escape() unwrapped returns &lt;script&gt;alert(1)&lt;/script&gt; — inert, visible text. Wrapping the identical characters as SafeString(...) first returns them completely untouched. Safety isn't a property of the value itself; it's a decision made once, deliberately, exactly at the one point a developer genuinely knows a value is already safe — never trusted implicitly by default.

A Real Node Tree, With Nested if/for

Real templates need loops and conditionals, and both can nest inside each other — an if inside a for, say. That needs a real parsed structure, not one flat regex pass. Four small node classes, each responsible only for rendering itself:

class TextNode: def __init__(self, text): self.text = text def render(self, ctx): return self.text class VarNode: def __init__(self, name): self.name = name def render(self, ctx): return escape(ctx.get(self.name, '')) class IfNode: def __init__(self, cond_name, body_nodes): self.cond_name = cond_name self.body_nodes = body_nodes def render(self, ctx): if ctx.get(self.cond_name): return ''.join(node.render(ctx) for node in self.body_nodes) return '' class ForNode: def __init__(self, var, coll_name, body_nodes): self.var = var self.coll_name = coll_name self.body_nodes = body_nodes def render(self, ctx): out = [] for item in ctx.get(self.coll_name, []): local = dict(ctx) local[self.var] = item for node in self.body_nodes: out.append(node.render(local)) return ''.join(out)

A tokenizer splits the raw source on {{ ... }} and {% ... %} boundaries; a small recursive-descent parser walks the tokens once, recursing into a fresh call for each if/for body and returning control the moment it hits the matching endif/endfor:

TOKEN_RE = re.compile(r'(\{\{.*?\}\}|\{%.*?%\})', re.S) def tokenize(template): return [t for t in TOKEN_RE.split(template) if t] def parse(tokens, index=0): nodes = [] while index < len(tokens): token = tokens[index] if token.startswith('{{'): nodes.append(VarNode(token[2:-2].strip())) index += 1 elif token.startswith('{%'): tag = token[2:-2].strip() if tag.startswith('if '): cond_name = tag[3:].strip() body, index = parse(tokens, index + 1) nodes.append(IfNode(cond_name, body)) elif tag.startswith('for '): m = re.match(r'for (\w+) in (\w+)', tag) var, coll_name = m.group(1), m.group(2) body, index = parse(tokens, index + 1) nodes.append(ForNode(var, coll_name, body)) elif tag in ('endif', 'endfor'): return nodes, index + 1 else: index += 1 else: nodes.append(TextNode(token)) index += 1 return nodes, index def compile_template(source): nodes, _ = parse(tokenize(source)) return nodes
Verified directly — a real nested if-inside-for, from a real file
A template of {% for item in items %}{% if item %}<li>{{ item }}</li>{% endif %}{% endfor %}, compiled once and rendered against {'items': ['a', '', 'b', None, 'c']}, produces exactly <li>a</li><li>b</li><li>c</li> — the two falsy items silently skipped, with the nested IfNode correctly evaluated fresh against each loop iteration's own local context.

Template Inheritance: extends & block, From Real Files

A shared layout lives in one parent file; a child file extends it and overrides only the regions it names:

templates/base.html
<html><head><title>{% block title %}My Site{% endblock %}</title></head><body>{% block content %}{% endblock %}</body></html>
templates/child.html
{% extends "base.html" %} {% block content %}Welcome, {{ user }}!{% endblock %}

child.html deliberately never mentions title at all — only content.

Resolving a child means reading both real files, extracting the child's own named blocks with one regex, and substituting each into the parent's matching block — before the merged result ever reaches compile_template(), so anything inside a block still gets the full real {{ }}/{% %} treatment:

def extract_blocks(text): blocks = {} for m in re.finditer(r'\{%\s*block\s+(\w+)\s*%\}(.*?)\{%\s*endblock\s*%\}', text, re.S): blocks[m.group(1)] = m.group(2) return blocks def substitute_blocks(parent_text, child_blocks): def replace(m): return child_blocks.get(m.group(1), m.group(2)) # child wins; else keep parent default return re.sub(r'\{%\s*block\s+(\w+)\s*%\}(.*?)\{%\s*endblock\s*%\}', replace, parent_text, flags=re.S)

Chaining the pieces manually, exactly the way a class will soon do internally:

def read_file(path): with open(path, 'r', encoding='utf-8') as f: return f.read() child_raw = read_file('templates/child.html') parent_raw = read_file('templates/base.html') merged = substitute_blocks(parent_raw, extract_blocks(child_raw)) nodes = compile_template(merged) print(''.join(n.render({'user': 'Ada'}) for n in nodes)) # <html><head><title>My Site</title></head><body>Welcome, Ada!</body></html>
Verified directly — title falls back to the parent's own real default
child.html only ever defines content; its real, rendered output nonetheless shows <title>My Site</title>, taken straight from base.html's own default text, since extract_blocks() never found a title block in the child at all, and substitute_blocks()'s own .get(m.group(1), m.group(2)) falls back to the parent's original content the instant a key is genuinely missing. A child only has to say what's genuinely different about it.

Building the Real TemplateEngine Class

Everything above — reading a real file, resolving extends, compiling to a node tree — becomes one class that owns a template directory and caches the compiled result, parsed once per template name and reused on every later render, exactly the discipline Web Framework Internals' own Chapter 4 verified Django's and Jinja2's real engines both follow:

import os class TemplateEngine: def __init__(self, template_dir): self.template_dir = template_dir self._cache = {} def _read_file(self, name): path = os.path.join(self.template_dir, name) with open(path, 'r', encoding='utf-8') as f: return f.read() def _resolve_source(self, template_name): raw = self._read_file(template_name) m = re.match(r'\s*\{%\s*extends\s+"([^"]+)"\s*%\}', raw) if m: parent_raw = self._read_file(m.group(1)) return substitute_blocks(parent_raw, extract_blocks(raw)) return raw def get_template(self, template_name): if template_name not in self._cache: self._cache[template_name] = compile_template(self._resolve_source(template_name)) return self._cache[template_name] def render(self, template_name, **context): nodes = self.get_template(template_name) return ''.join(node.render(context) for node in nodes)
A real, genuine bug found while writing this exact class
The first version of render() named its own template-name parameter name, matching Chapter 2's own Router.add_route(..., name=...) convention. It broke immediately, on the very first real call: engine.render('basic.html', name='World') raised TypeError: render() got multiple values for argument 'name' — 'basic.html' filled the positional name parameter, and the keyword name='World' (a completely ordinary, plausible context value) collided with it directly. Renaming the parameter to template_name fixed it outright. This is the exact reason Flask's own real render_template() never calls its own first parameter name either — name is too common a real context key for a template engine's own API to claim for itself.
Verified directly — compile-once caching adds zero extra file reads
Instrumenting _read_file() to count real calls, then rendering the same already-cached template five times in a row, adds exactly zero further file reads. The file is opened and parsed on the very first render only; every render after that walks the same in-memory node tree Chapter 4's own compiler already built.

A Real, Live-Verified Limitation: Stale Caching

Caching by name has an honest cost: once a template is compiled, editing its file on disk mid-run changes nothing a running server actually serves.

# a real, live server, already running against templates/greeting.html: # "<p>Hello, {{ name }}! (v1)</p>" first_response = get('/') # <p>Hello, Ada! (v1)</p> # the real file on disk is edited WHILE the server keeps running -- # greeting.html now reads "<p>Hi there, {{ name }}! (v2, updated on disk)</p>" second_response = get('/') # <p>Hello, Ada! (v1)</p> -- still the OLD content
Verified directly against a real, live running server
Two genuine GET requests, sent to the same live wsgiref server, with the real template file edited on disk in between, both return the identical v1 output — confirmed by actually running the edit and the second request against a process that never restarted. This isn't a bug so much as an unavoidable consequence of caching correctly by name: nothing in TemplateEngine has any way of knowing the file underneath a cached name has changed. Fixing this — detecting a changed file and recompiling automatically — is exactly the real job of the live-reload development server this course builds in Chapter 8.

Wiring TemplateEngine Into Response

A handler now renders a real file, through the real engine, and wraps the result in a Response exactly the way Chapter 3's own closing line promised. A third, separate real template exercises everything at once — extending the same base.html, but this time overriding both of its blocks:

templates/users.html
{% extends "base.html" %} {% block title %}Users{% endblock %} {% block content %} <h1>Users</h1> {% if empty %}<p>No users found.</p>{% endif %} <ul>{% for user in users %}<li>{{ user }}</li>{% endfor %}</ul> {% endblock %}
templates = TemplateEngine('templates') router.add_route('GET', '/users', lambda req: Response(templates.render( 'users.html', users=['Ada', 'Grace', '<b>Rogue</b>'], empty=False, ))) router.add_route('GET', '/users/empty', lambda req: Response(templates.render( 'users.html', users=[], empty=True, )))

Served through the real wsgiref app and hit with genuine urllib.request calls:

GET /users → <html><head><title>Users</title></head><body><h1>Users</h1> <ul><li>Ada</li><li>Grace</li><li>&lt;b&gt;Rogue&lt;/b&gt;</li></ul></body></html> GET /users/empty → <html><head><title>Users</title></head><body><h1>Users</h1> <p>No users found.</p><ul></ul></body></html>
Verified directly — extends, block, if, for, and escaping, all correct together
Every mechanism this chapter built is exercised in one real response: the layout comes from base.html, the title and content blocks come from users.html, the empty-state {% if %} correctly fires only on the second route, the {% for %} loop correctly renders three real list items, and the third item's own literal <b>Rogue</b> string comes back fully escaped — confirming auto-escaping survives the full route-to-template pipeline, not just the isolated examples earlier in this chapter.

Where This Course Is Headed

Chapter 5 builds a real middleware pipeline — an onion-model chain wrapping every request and response, the mechanism a real framework would use to add authentication, logging, or CSRF protection around every route this chapter's own Router already dispatches to.

Hands-On Exercises

Exercise 1

Extend this chapter's own IfNode and parse() function to support a real {% else %} branch โ€” {% if flag %}...{% else %}...{% endif %} โ€” and verify it against both a truthy and a falsy value for flag, plus a genuinely nested case with an if/else living inside a for loop.

๐Ÿ“„ View solution
Exercise 2

Write a real render_bio_markdown(raw_text) function that escapes its input completely first, then converts **bold** markers on the now-safe text into real <strong> tags, returning the result as a SafeString โ€” and verify against a single real input containing both a **bold** marker and a literal <script> tag that only the bold marker becomes real markup while the script tag stays fully inert.

๐Ÿ“„ View solution
Exercise 3

Reproduce this chapter's own stale-cache finding yourself against a real, live wsgiref server serving a route rendered through TemplateEngine: make one real request, edit the template file on disk while the server keeps running, make a second real request, and confirm both responses are identical โ€” then explain in your own words why get_template()'s own caching is what causes this, not a bug in compile_template() itself.

๐Ÿ“„ View solution

Chapter 4 Quick Reference

  • The basic job โ€” read a real .html file, substitute {{ }} placeholders from a context dict
  • The real security problem โ€” unescaped interpolation puts a real, live <script> tag straight into rendered output
  • Auto-escaping by default โ€” escape() escapes everything except values explicitly wrapped as SafeString
  • The node tree โ€” TextNode/VarNode/IfNode/ForNode, parsed once via a real recursive-descent tokenizer/parser supporting nested if/for
  • extends/block inheritance โ€” child blocks substituted into the parent's raw text before compilation, verified falling back to the parent's own default for any block a child leaves untouched
  • TemplateEngine โ€” real file loading, inheritance resolution, and compile-once caching in one class; a genuine name/template_name naming-collision bug found and fixed along the way
  • Stale caching โ€” verified live: editing a cached template's file on disk doesn't change what a running server serves, until Chapter 8's own live-reload dev server fixes it
  • Next chapter: A real onion-model middleware pipeline