Middleware: A Real Onion-Model Pipeline

Building a Web Framework

Chapter 5 ยท Middleware: A Real Onion-Model Pipeline

Every chapter so far has hand-wired its own app(environ, start_response) function, manually calling Router.resolve() and manually turning the result into a Response. That works for one route file, but it has no place to put code that should run around every request — authentication, logging, timing, error handling — without editing every single handler to add it. This chapter builds that missing layer: a real middleware pipeline, wired into a new, persistent App class that finally becomes the actual WSGI callable Chapter 1 first introduced, replacing every hand-rolled app() function this course has written since.

The Basic Job: A Chain of Functions Wrapping "Next"

A middleware is a function taking the request and a reference to whatever comes next — another middleware, or the real handler at the center. Composing a list means nesting each one inside the next:

def build_chain(middlewares, handler): chain = handler for middleware in reversed(middlewares): chain = (lambda mw, nxt: lambda request: mw(request, nxt))(middleware, chain) return chain def middleware_a(request, next_): trace.append('before A') response = next_(request) trace.append('after A') return response # middleware_b, middleware_c follow the identical shape def toy_handler(request): trace.append('handler') return Response('OK') chain = build_chain([middleware_a, middleware_b, middleware_c], toy_handler) chain(Request({'REQUEST_METHOD': 'GET', 'PATH_INFO': '/'})) print(trace)
Verified directly — real, genuinely nested execution, not sequential
trace comes out as ['before A', 'before B', 'before C', 'handler', 'after C', 'after B', 'after A']. Each middleware's own "after" code runs in the exact reverse of registration order — A registers first and finishes last, since A's own call to next_(request) is what invokes everything nested inside it. Real frameworks name this precisely — Django's own documentation: "You can think of it like an onion: each middleware class is a 'layer' that wraps the view, which is in the core of the onion."

Building the Real App Class: A Router, a Middleware List, and a WSGI Callable

Every piece from Chapters 2–4 — Router, Request, Response — folds into one persistent class that owns a middleware list too, and implements __call__(environ, start_response) directly, the exact real interface Chapter 1 verified against wsgiref's own reference server:

class App: def __init__(self): self.router = Router() self.middlewares = [] self._chain = None def add_route(self, method, path, handler, name=None): self.router.add_route(method, path, handler, name=name) self._chain = None # a new route can change what the innermost handler does def use(self, middleware): self.middlewares.append(middleware) self._chain = None # force a rebuild on the next real request def _dispatch(self, request): status, handler, params, allowed = self.router.resolve(request.method, request.path) if status == 'MATCH': request.params = params return handler(request) elif status == '405': return Response('405 Method Not Allowed', status=405, headers=[('Allow', ', '.join(sorted(allowed)))]) else: return Response('404 Not Found', status=404) def _build_chain(self): chain = self._dispatch # the innermost "handler" IS route dispatch itself for mw in reversed(self.middlewares): chain = (lambda m, nxt: lambda req: m(req, nxt))(mw, chain) return chain def __call__(self, environ, start_response): if self._chain is None: self._chain = self._build_chain() request = Request(environ) response = self._chain(request) return response.to_wsgi(start_response)

The middleware chain is now built once, on first use, and cached — the exact same "build once, reuse" discipline Chapter 4's own TemplateEngine already established, invalidated only when a new route or middleware is actually registered.

Verified directly — an App instance is a real WSGI application, with nothing wrapped around it
make_server('127.0.0.1', 8861, app) — passing the App instance itself, not a function — served a real request correctly. Python's own real wsgiref server only ever cared that its third argument is callable with (environ, start_response); a class instance with a real __call__ method satisfies that identically to the plain functions every prior chapter used. This is the exact same real interface Chapter 1 verified against PEP 3333, now implemented by a genuine, stateful object instead of a bare function.

Short-Circuiting: A Real Authentication Middleware

Nothing forces a middleware to call next_() at all. An auth check is the classic real reason not to — an unauthenticated request has no business reaching the real handler:

def auth_middleware(request, next_): if request.environ.get('HTTP_X_API_KEY') != 'secret123': return Response('401 Unauthorized', status=401) # next_() is never called return next_(request) app = App() app.use(auth_middleware) app.add_route('GET', '/secret', lambda req: Response('secret data'))
Verified directly against a real, live server
Three genuine requests, sent to the same running App: no X-Api-Key header returns a real 401; a wrong key also returns 401; the correct key, secret123, returns 200 and the real handler's own body. The rejected requests never reach the handler at all — verified by a second, chained logging_middleware that never records anything for either of the two rejected requests.

The Real Cost of Forgetting to Call next_()

Short-circuiting on purpose is a feature. Forgetting to call next_() at all, and not returning anything either, is a real, easy mistake — Express's own documentation names the toy-version consequence directly: "Otherwise, the request will be left hanging." A real WSGI app built on this chapter's own App class has a genuinely worse, more concrete version of that same mistake:

def broken_middleware(request, next_): print('ran, but forgot to call next_() or return anything') # no call to next_(request); no return statement at all app = App() app.use(broken_middleware) app.add_route('GET', '/', lambda req: Response('should never be reached'))
Verified directly — a real crash, not a silent None
Hit with a genuine live request, the server's own console shows a real Python traceback: AttributeError: 'NoneType' object has no attribute 'to_wsgi', raised from App.__call__'s own response.to_wsgi(start_response) line — self._chain(request) genuinely returned None, and nothing downstream of that was ever written to be None-safe. wsgiref itself catches that exception and produces a real HTTP 500, with a body of exactly "A server error occurred. Please contact the administrator." — the client receives a completely real, working 500 response, just one carrying zero information about what actually broke.

Mutating the Request and the Response

Code running before next_() can attach data every downstream layer can read — this chapter's own Request class gets a plain state dict specifically for this:

def inject_user_middleware(request, next_): request.state['user'] = {'id': 42, 'name': 'Sam'} return next_(request)

Code running after next_() returns can inspect or modify the real response on its way back out — the exact real example FastAPI's own documentation uses to introduce middleware at all is a timing header:

FastAPI's own documentation, quoted directly
"You can add code to be run with the request, before any path operation receives it. And also after the response is generated, before returning it." FastAPI's own real example measures elapsed time around call_next(request) and sets response.headers["X-Process-Time"] with the result.
def timing_middleware(request, next_): start = time.perf_counter() response = next_(request) elapsed = time.perf_counter() - start response.headers.append(('X-Process-Time', f"{elapsed:.6f}")) return response def slow_handler(request): time.sleep(0.05) user = request.state.get('user') return Response(f"Hello, {user['name']} (id={user['id']})") app = App() app.use(inject_user_middleware) app.use(timing_middleware) app.add_route('GET', '/', slow_handler)
Verified directly — a real, accurately-measured elapsed time from a real live request
slow_handler genuinely sleeps for 0.05 real seconds; the live response body reads Hello, Sam (id=42) — confirming inject_user_middleware's own write to request.state really was visible two layers downstream — and the real X-Process-Time header measured 0.050328, accurate to a fraction of a millisecond of real overhead, independently reproducing FastAPI's own documented example on this framework's own code.

Error-Handling Middleware: With It, and Genuinely Without It

Since next_() is an ordinary function call, an exception raised deep in the chain propagates upward through every enclosing layer exactly like any other Python exception — a middleware can catch it with a plain try/except around that one call:

def error_handling_middleware(request, next_): try: return next_(request) except ValueError as e: return Response(f"500 Internal Server Error: {e}", status=500) def crashing_handler(request): raise ValueError('something genuinely went wrong')

Registering the identical crashing_handler two ways — with the middleware, and without it — on two separate live servers:

ConfigurationReal, live client result
error_handling_middleware present500, body exactly 500 Internal Server Error: something genuinely went wrong
no error-handling middleware at all500, body exactly A server error occurred. Please contact the administrator.
Verified directly, live, both configurations — the real difference is what the client actually learns
Both configurations genuinely return a real 500 — the raw ValueError, left uncaught, propagates all the way out of App.__call__ itself and is caught by wsgiref's own internal handler exactly the same way Chapter 5's own broken_middleware case was, producing the identical generic message. With error_handling_middleware in place, the client instead receives this framework's own specific, real message — "something genuinely went wrong", the exact real exception text, not a generic placeholder. Nothing here is special-cased for errors specifically; a real try/except around one ordinary function call is the entire mechanism, and skipping it doesn't lose the request — it loses the information a real developer or user would actually need.

Where This Course Is Headed

Chapter 6 builds real sessions and cookies — state that survives across genuinely separate, stateless HTTP requests — almost certainly implemented as a middleware itself, reading a session cookie before next_() and writing an updated one after it returns, using exactly the two mutation points this chapter just verified.

Hands-On Exercises

Exercise 1

Add a cors_middleware to this chapter's own App system that adds an Access-Control-Allow-Origin: * header to every response after next_() returns, register it on a real App with app.use(), and verify against a real live request that the header genuinely arrives on the client's own response.

๐Ÿ“„ View solution
Exercise 2

Using this chapter's own auth_middleware and a new logging_middleware that appends to request.state before and after calling next_(), build two real Apps that register the two middlewares in opposite orders (auth-then-logging vs. logging-then-auth), send an unauthenticated real request to each, and verify from request.state whether logging_middleware's own "before" code actually ran before the rejection in each case.

๐Ÿ“„ View solution
Exercise 3

Define a real NotFoundError exception class, extend this chapter's own error_handling_middleware to catch it separately from ValueError and return a distinct real 404 response for it, and verify live against two separate routes โ€” one whose handler raises ValueError, one whose handler raises NotFoundError โ€” that each produces its own distinct, correct real status code and body.

๐Ÿ“„ View solution

Chapter 5 Quick Reference

  • The basic job โ€” each middleware wraps "whatever comes next," composed into one nested call chain
  • The real onion order โ€” verified as before-A/before-B/before-C/handler/after-C/after-B/after-A, matching Django's own documented "onion" model
  • App โ€” a real, persistent class owning Router + a middleware list, implementing __call__(environ, start_response) directly; verified as a genuine WSGI app handed straight to wsgiref, no wrapper function needed
  • Short-circuiting โ€” a middleware can return a response without calling next_(), verified stopping every inner layer, including a chained logging middleware, from running at all
  • Forgetting to call next_() โ€” verified as a real crash in this chapter's own App, caught by wsgiref and surfaced as its own generic 500, not a silent None
  • Request/response mutation โ€” request.state before next_(); response.headers after next_() returns โ€” verified with a real, measured X-Process-Time header
  • Error-handling middleware โ€” an ordinary try/except around one call to next_(); verified live that it's the difference between the client receiving this framework's own real error message and wsgiref's own generic one
  • Next chapter: Sessions and cookies โ€” real state across stateless requests