A Routing Engine From Scratch: Path Matching & Dynamic Parameters

Building a Web Framework

Chapter 2 ยท A Routing Engine From Scratch: Path Matching & Dynamic Parameters

Chapter 1 verified that a real request arrives as a plain dict — PATH_INFO, REQUEST_METHOD, and the rest, sitting right there in environ. Nothing yet decides which piece of your own code should actually run for a given path and method; every request would currently have to be handled by one giant function with its own if/elif chain reading environ by hand. This chapter builds the piece that removes that need — a real, reusable Router class — and it's not a one-off demonstration: this is the exact object every later chapter in this course imports and builds on, starting with Chapter 3's request object next.

Mirroring, not repeating, Web Framework Internals' own Chapter 2
That course built and verified a standalone router as a way of explaining how routing works, then moved on to comparing five real frameworks' own syntax for it. This chapter builds the identical underlying mechanics — static/dynamic segments, typed converters, registration order, named routes, method dispatch — as a real, permanent class this course's own framework keeps and reuses, wired into a genuine running WSGI app by the end of the chapter rather than left as isolated examples.

The Basic Job: A Lookup Table From Path to Handler

At its simplest, a router is nothing more than a dict: a path string maps to a handler, and dispatch is a dict lookup with a default for "not found."

routes = {} def add_route(path, handler): routes[path] = handler def dispatch(path): handler = routes.get(path) if handler is None: return '404 Not Found' return handler() add_route('/about', lambda: 'About page') print(dispatch('/about')) # About page print(dispatch('/missing')) # 404 Not Found

Genuinely correct, genuinely useless past a handful of static pages — there's no way to write one route that matches every /users/<some id> path at once.

Dynamic Segments, and the String Trap They Leave Behind

Regular expressions fix the matching problem directly — a path like /users/<id> can be turned into a real regex with a named capture group:

import re def compile_pattern(path): regex = re.sub(r'<(\w+)>', r'(?P<\1>[^/]+)', path) return re.compile(f'^{regex}$') pattern = compile_pattern('/users/<id>') match = pattern.match('/users/42') print(match.groupdict()) # {'id': '42'} print(type(match.groupdict()['id'])) # <class 'str'>
Verified directly — the captured value is always a plain string
match.groupdict() returns {'id': '42'} — the character string '42', confirmed by type() as a real Python str, never the number 42. A regex has no idea a segment was "supposed to" be numeric; it just matched characters. A handler that does id + 1 on that value crashes with a real TypeError — not because the router is broken, but because nothing has told it a segment should be anything other than text.

Typed Converters: Validate and Cast, Once, Before the Handler Sees It

The fix is to let a route segment declare its own type, and have the router validate and convert it in one place rather than leaving every handler to do it separately:

CONVERTERS = { 'int': (r'\d+', int), 'str': (r'[^/]+', str), } def compile_typed_pattern(path): parts = [] types = {} for segment in path.split('/'): m = re.match(r'<(\w+):(\w+)>', segment) if m: conv_name, param_name = m.group(1), m.group(2) regex, cast = CONVERTERS[conv_name] parts.append(f'(?P<{param_name}>{regex})') types[param_name] = cast else: parts.append(re.escape(segment)) return re.compile('^' + '/'.join(parts) + '$'), types pattern, types = compile_typed_pattern('/users/<int:id>') for test_path in ['/users/42', '/users/abc']: m = pattern.match(test_path) if m: params = {k: types[k](v) for k, v in m.groupdict().items()} print(test_path, '→', params) else: print(test_path, '→ no match') # /users/42 → {'id': 42} # /users/abc → no match
Verified directly — a real int, and a real, correct rejection
/users/42 prints {'id': 42}, a genuine Python int cast once by the router before any handler ever runs. /users/abc prints no match outright — the \d+ pattern behind int never matches non-digit characters, so the request falls through to a real 404 instead of reaching a handler that expected a number and would have crashed on "abc".

Building the Real Router Class

Everything above works as standalone functions. This framework needs something that can hold many routes at once, across many HTTP methods, and be queried repeatedly — so it becomes a real class, Router, the actual file this course keeps building on:

import re CONVERTERS = { 'int': (r'\d+', int), 'str': (r'[^/]+', str), } class Router: """A real, minimal router this framework reuses in every later chapter.""" def __init__(self): self._routes = [] # (method, compiled_regex, types, handler) self._named = {} # name -> raw path template, for url_for() def add_route(self, method, path, handler, name=None): regex, types = self._compile(path) self._routes.append((method.upper(), regex, types, handler)) if name: self._named[name] = path def _compile(self, path): parts, types = [], {} for segment in path.split('/'): m = re.match(r'<(\w+):(\w+)>', segment) if m: conv_name, param_name = m.group(1), m.group(2) seg_regex, cast = CONVERTERS[conv_name] parts.append(f'(?P<{param_name}>{seg_regex})') types[param_name] = cast else: parts.append(re.escape(segment)) pattern = re.compile('^' + '/'.join(parts) + '$') return pattern, types def resolve(self, method, path): method = method.upper() allowed = set() for route_method, regex, types, handler in self._routes: m = regex.match(path) if m: allowed.add(route_method) if route_method == method: params = {k: types[k](v) for k, v in m.groupdict().items()} return 'MATCH', handler, params, None if allowed: return '405', None, {}, allowed return '404', None, {}, None def url_for(self, name, **params): template = self._named[name] for key, value in params.items(): template = re.sub(rf'<\w+:{key}>', str(value), template) return template

resolve() is the one method everything else in this chapter, and this course, calls. It walks every registered route once, remembers every HTTP method that matched the path even when it doesn't match the method, and returns one of exactly three outcomes: a real match with its typed params, a 405 with the set of methods that would have worked, or a plain 404.

Registration Order: A Real, Reproducible Bug on This Course's Own Router

Two registered routes can genuinely both match the same path. Router.resolve() scans its own _routes list top to bottom and returns the first one whose method also matches — which means, exactly as with Django's own real URL resolver, registration order is load-bearing.

router = Router() router.add_route('GET', '/users/<str:id>', lambda **p: f"user detail {p['id']}", name='user_detail') router.add_route('GET', '/users/new', lambda **p: 'user create form', name='user_create') status, handler, params, allowed = router.resolve('GET', '/users/new') print(status, handler(**params)) # MATCH user detail new
Verified directly — /users/new never reaches its own handler
Registered in this order, resolve('GET', '/users/new') returns MATCH against the general /users/<str:id> route, with id captured as the literal string "new" — the dedicated /users/new route is never even checked, because the loop already returned on the first match it found. Registering the exact same two routes with /users/new listed first fixes it directly, verified against the same input: MATCH user create form, with the general route only ever reached by a path that genuinely isn't "new".
Typed converters narrow this risk without eliminating it
Swapping in <int:id> instead of <str:id> happens to reject "new" on its own, since it contains no digits — verified separately: registered in either order, /users/<int:id> and /users/new resolve correctly, because the general route's own regex is now narrow enough to never accidentally swallow the literal path. That's a real, useful side effect of typing a segment, not a substitute for specific-before-general ordering: any two routes whose patterns can both genuinely match the same path still need the more specific one registered first.

Method Dispatch: A Real, Distinct 405, Not Just a Second 404

HTTP defines 405 Method Not Allowed for exactly the case resolve() tracks with its own allowed set: a path that matches something registered, on a method that isn't. It's a genuinely different situation from no match at all, and a real client can tell the two apart — including, correctly implemented, via the response's own Allow header naming which methods would have worked.

router = Router() router.add_route('GET', '/users/<int:id>', lambda **p: f"get user {p['id']}", name='user_get') router.add_route('DELETE', '/users/<int:id>', lambda **p: f"delete user {p['id']}", name='user_delete') print(router.resolve('GET', '/users/42')) print(router.resolve('DELETE', '/users/42')) print(router.resolve('PATCH', '/users/42')) # path matches, method doesn't print(router.resolve('GET', '/nonexistent')) # no path match at all
Verified directly, wired to a real WSGI app in the next section
GET /users/42 and DELETE /users/42 both return MATCH against the same path with their own correct handler and params. PATCH /users/42 returns ('405', None, {}, {'GET', 'DELETE'}) — the path itself genuinely matched, so resolve() reports every method that would have worked instead of pretending nothing was found there. GET /nonexistent returns ('404', None, {}, None) — no route's own pattern matched this path under any method at all, a real, distinct third outcome the code below turns into two genuinely different HTTP status lines.

Named & Reverse Routing

url_for() goes the other direction — given a route's own name and its param values, it generates the real path string, so nothing downstream ever has to hardcode /users/42 as a literal.

router = Router() router.add_route('GET', '/users/<int:id>', lambda **p: None, name='user_detail') print(router.url_for('user_detail', id=42)) # /users/42
A real, previously-found bug this exists specifically to avoid
Web Framework Internals' own Personal Catalogue (Django & PostgreSQL) and Premier League Predictor: Django & MySQL courses each found the identical real consequence of skipping this: a handler that hardcodes redirect('/users/42') as a literal string breaks silently the instant the app is mounted somewhere other than the site root, since nothing in that literal knows about the new prefix. A handler that calls router.url_for('user_detail', id=42) instead can be told about a mount-point change in exactly one place — inside Router itself — and have every generated path update automatically. This framework's own future deployment chapter revisits this directly.

Wiring It Into a Real, Running WSGI App

Every piece above is verified in isolation. Put together, Router becomes the actual dispatch layer of a genuine WSGI application — built directly on the environ dict Chapter 1 already verified, and tested against real HTTP requests exactly the way Chapter 1's own examples were:

router = Router() router.add_route('GET', '/', lambda **p: 'Home page', name='home') router.add_route('GET', '/users/<int:id>', lambda **p: f"User {p['id']}", name='user_detail') router.add_route('DELETE', '/users/<int:id>', lambda **p: f"Deleted user {p['id']}", name='user_delete') router.add_route('GET', '/users/new', lambda **p: 'New user form', name='user_create') def app(environ, start_response): method = environ['REQUEST_METHOD'] path = environ.get('PATH_INFO', '/') status, handler, params, allowed = router.resolve(method, path) if status == 'MATCH': body = handler(**params).encode('utf-8') start_response('200 OK', [('Content-type', 'text/plain')]) return [body] elif status == '405': allow_header = ', '.join(sorted(allowed)) start_response('405 Method Not Allowed', [ ('Content-type', 'text/plain'), ('Allow', allow_header), ]) return [f"405 Method Not Allowed (allowed: {allow_header})".encode('utf-8')] else: start_response('404 Not Found', [('Content-type', 'text/plain')]) return [b"404 Not Found"]

Served through Chapter 1's own wsgiref.simple_server and hit with genuine urllib.request calls, including a real DELETE and a real PATCH:

GET / → 200 b'Home page' GET /users/42 → 200 b'User 42' GET /users/new → 200 b'New user form' DELETE /users/42 → 200 b'Deleted user 42' PATCH /users/42 → 405 b'405 Method Not Allowed (allowed: DELETE, GET)' Allow: DELETE, GET GET /nonexistent → 404 b'404 Not Found'
Verified directly — a real, correct Allow header on a real 405 response
Every line above is a genuine round trip through a live wsgiref server: a real request goes out, Router.resolve() decides its outcome, and app() turns that outcome into a real HTTP status line and body. The PATCH /users/42 response's own Allow header, read back from the actual response object, correctly lists DELETE, GET — exactly the two methods registered against that path, in sorted order, with nothing hardcoded anywhere in app() itself. That header is computed fresh from whatever Router actually has registered at the moment of the request.

Where This Course Is Headed

Chapter 3 wraps the raw environ dict and this chapter's own typed params dict into a single, ergonomic request object — and a matching response object, so a handler never has to call start_response directly again.

Hands-On Exercises

Exercise 1

Add a real 'slug' entry to this chapter's own CONVERTERS dictionary (matching lowercase letters, digits, and hyphens only), register a route /posts/<slug:title> on this course's own Router class, and verify it correctly matches "hello-world-42" while rejecting "Hello World!".

๐Ÿ“„ View solution
Exercise 2

Using this chapter's own untyped <str:id> router, register /users/<str:id> before /users/new and demonstrate concretely, via a real Router instance and a printed resolve() call, that GET /users/new resolves to the wrong handler โ€” then fix it by reordering the two registrations and confirm the correct handler now resolves.

๐Ÿ“„ View solution
Exercise 3

Register a single POST-only route on a fresh Router instance, then call resolve() with GET against that same path and against a completely unregistered path, printing both results โ€” and explain in your own words why the first genuinely deserves a 405 with an Allow header while the second deserves a plain 404.

๐Ÿ“„ View solution

Chapter 2 Quick Reference

  • The basic job โ€” match an HTTP method + path to a registered handler, or report that nothing matched
  • Dynamic segments โ€” regex capture groups; verified always coming back as plain strings, never numbers, without a converter
  • Typed converters โ€” CONVERTERS maps a name (int, str, โ€ฆ) to a (regex, cast) pair, validating and casting a segment once before any handler sees it
  • Registration order โ€” resolve() returns the first matching route it finds; a general pattern registered before a more specific one really does swallow it, verified directly on this course's own Router
  • 404 vs. 405 โ€” resolve() tracks every method whose path matched, returning a real 405 (with the Allow header this chapter's own app() builds from it) instead of a plain 404 when only the method was wrong
  • Named/reverse routing โ€” url_for() generates a path from a route's own name, the exact real fix this site's own Django courses needed for a mount-point path bug
  • Router is now a real, permanent file โ€” every later chapter in this course imports and extends it, starting with Chapter 3's request/response objects
  • Next chapter: Request & response objects โ€” wrapping the raw WSGI environ