Request & Response Objects: Wrapping the Raw WSGI Environ

Building a Web Framework

Chapter 3 ยท Request & Response Objects: Wrapping the Raw WSGI Environ

A handler in Chapter 2's own demo app already reads req.params['id'] instead of parsing PATH_INFO by hand, but it still touches environ directly for the method and path, and every handler builds its own status line and header list by hand. This chapter wraps both directions of that raw interface into two real classes — Request and Response — so a handler only ever has to read a clean object and return a clean object, never call start_response or reach into environ itself again.

Building the Request Class

Request takes the raw environ dict Chapter 1 verified, plus the typed params dict Chapter 2's own Router.resolve() already produces, and exposes them as plain attributes:

import urllib.parse class Request: def __init__(self, environ, params=None): self.environ = environ self.method = environ.get('REQUEST_METHOD', 'GET').upper() self.path = environ.get('PATH_INFO', '/') self.params = params or {} self._query_string = environ.get('QUERY_STRING', '') self._query = None self._body = None

query and body are both real Python properties, computed once on first access and cached — a genuine, deliberate design decision, not an arbitrary one, for reasons this chapter's own middle section verifies directly.

@property def query(self): if self._query is None: parsed = urllib.parse.parse_qs(self._query_string) self._query = {k: v[0] for k, v in parsed.items()} return self._query
Verified directly — a real query string, parsed into a plain dict
A live request to /users/42?active=true, routed through Chapter 2's own Router, reaches a handler whose req.query returns {'active': 'true'} — urllib.parse.parse_qs, part of Python's own standard library, doing the real work; this class only wraps it lazily so a route that never reads req.query never pays the parsing cost at all.

Headers: A Real, Documented Gap in the WSGI Spec

Per PEP 3333, every real HTTP request header a client sends arrives in environ with an HTTP_ prefix and underscores in place of hyphens — X-Api-Key becomes the key HTTP_X_API_KEY. A headers property can reverse that mechanically:

@property def headers(self): result = {} for key, value in self.environ.items(): if key.startswith('HTTP_'): result[key[5:].replace('_', '-').title()] = value return result
Verified directly — Content-Type never shows up in this dict, on purpose
A real request sent with both a custom X-Api-Key header and a real Content-Type: application/json header produces req.headers equal to {'X-Api-Key': 'secret-123', 'Host': ..., 'User-Agent': ..., 'Accept-Encoding': ..., 'Connection': ...} — genuinely missing Content-Type entirely, even though it was really sent. This isn't a bug in the property above: PEP 3333 itself deliberately excludes CONTENT_TYPE and CONTENT_LENGTH from the HTTP_-prefixed set, storing them as their own separate, unprefixed environ keys instead. A headers property that only scans for HTTP_* keys will silently miss both, every time, on every real WSGI server — not just this one.

Reading CONTENT_TYPE back out correctly just means asking for the unprefixed key directly, exactly as this chapter's own content_length property (below) already does for its own equally unprefixed sibling.

The Request Body: A Live Socket Is Not a BytesIO Object

content_length reads the same real, unprefixed key, guarding against the genuine case where it's missing entirely — a plain GET request with no body sends no CONTENT_LENGTH at all, and int('') raises a real ValueError if that's not handled:

@property def content_length(self): raw = self.environ.get('CONTENT_LENGTH', '') try: return int(raw) except ValueError: return 0 @property def body(self): if self._body is None: self._body = self.environ['wsgi.input'].read(self.content_length) return self._body
Verified directly — a real GET request, CONTENT_LENGTH genuinely absent
A real live GET request through this exact code returns content_length=0 body=b'' — no crash, no exception, despite there being no CONTENT_LENGTH key in environ for it at all.

body reads exactly content_length bytes from environ['wsgi.input'] — the real request-body stream PEP 3333 itself warns about directly: "the application should not attempt to read more data than is specified by the CONTENT_LENGTH variable," and "the server is not required to read past the client's specified Content-Length." Reading a bounded, correct amount looks safe. Whether it's actually safe to read twice turns out to depend entirely on what kind of stream is behind wsgi.input — verified two genuinely different ways below.

import io # a fake, in-memory environ -- the kind a quick unit test might use fake_stream = io.BytesIO(b'name=test') print(fake_stream.read(9)) # b'name=test' print(fake_stream.read(9)) # b'' -- looks perfectly safe
Verified directly — a second read on BytesIO just returns empty, instantly
Reading past the end of an in-memory BytesIO object behaves exactly the way most Python programmers already expect a stream to behave: it hits a real, immediate end-of-file and returns b''. A unit test built on a hand-constructed environ dict like this one would show Request.body accessed twice returning the identical cached bytes both times, with nothing to suggest any danger exists.

A real, live wsgi.input — the one a genuine wsgiref server actually hands a handler — is backed by a live TCP socket, not a bounded in-memory buffer. Three scenarios, run against a real server and a real client, each timed:

ScenarioWhat it doesReal, measured result
single_cachedOne handler reads req.body exactly once200 OK in 0.02s
double_cachedOne handler reads req.body twice; caching means the stream is only ever touched by the first access200 OK in 0.03s — both reads return the same cached object (first is second → True)
raw_peek_then_cachedA handler reads environ['wsgi.input'] directly first (a correctly-bounded read, using the real Content-Length), then constructs a Request and reads req.body afterwardClient times out — verified hanging past a forced 3-second cutoff, with the request never completing
Verified directly — a correctly-bounded second read still hangs forever on a real connection
The dangerous scenario above isn't calling .read() with no size argument at all — it's calling a properly bounded .read(content_length) a second time, after the real body bytes have already been consumed once. On a real, persistent HTTP connection, the client has finished sending exactly those bytes and is now waiting for a response; there is no more data coming, and a live socket has no built-in concept of "logically done with the HTTP body" the way a bounded BytesIO object does — it just keeps waiting for more bytes to arrive. The very first version of this experiment, using a genuinely unbounded .read() with no argument at all, was left running for over two real minutes with no response before it had to be forcibly terminated.
This is exactly why Request.body caches, and why it matters more than it looks
Request.body's own if self._body is None: check is the entire reason double_cached above completes in 0.03 seconds instead of hanging: the underlying stream is read from environ['wsgi.input'] exactly once, no matter how many times any code later asks for request.body. Any code path — a handler and a piece of middleware both wanting to inspect the body, for instance, a real scenario this course's own Chapter 5 builds toward — that reaches the raw stream a second time, bypassing this cache, reintroduces the exact hang verified above. And because an in-memory BytesIO-backed test can't reproduce this at all, a test suite built entirely on fake environs can pass cleanly while this exact bug ships to a real server.

Form Data: Reusing body()

A URL-encoded POST body is parsed the same way a query string is — form() reuses body (and therefore its caching) directly:

def form(self): parsed = urllib.parse.parse_qs(self.body.decode('utf-8')) return {k: v[0] for k, v in parsed.items()}
Verified directly against a real POST request
A real request sending name=Ada&age=30 as its body reaches a handler whose req.form() returns exactly {'name': 'Ada', 'age': '30'}.

Building the Response Class

Chapter 2's own app() built every status line and header list by hand, in every branch. Response moves that into one place, and computes Content-Length automatically — a handler never has to get it right or wrong:

import json class Response: STATUS_MESSAGES = { 200: 'OK', 201: 'Created', 204: 'No Content', 301: 'Moved Permanently', 302: 'Found', 400: 'Bad Request', 404: 'Not Found', 405: 'Method Not Allowed', 500: 'Internal Server Error', } def __init__(self, body='', status=200, headers=None, content_type='text/plain'): self.body = body self.status = status self.headers = list(headers) if headers else [] self.content_type = content_type def status_line(self): return f"{self.status} {self.STATUS_MESSAGES.get(self.status, 'Unknown')}" def to_wsgi(self, start_response): body_bytes = self.body.encode('utf-8') if isinstance(self.body, str) else self.body header_list = [ ('Content-type', self.content_type), ('Content-Length', str(len(body_bytes))), ] + self.headers start_response(self.status_line(), header_list) return [body_bytes]
Verified directly — correct Content-Length, every time, with zero manual work
A handler that returns Response('a handler never has to compute Content-Length by hand') and nothing else produces a real response whose measured body length and whose Content-Length header both come back as 53 — identical, verified from the client side, without the handler itself ever mentioning the number 53 anywhere.

Two Real Convenience Constructors: redirect() and json()

@classmethod def redirect(cls, location, status=302): return cls(body='', status=status, headers=[('Location', location)]) @classmethod def json(cls, data, status=200): return cls(body=json.dumps(data), status=status, content_type='application/json')
Verified directly — both real redirect behaviors, and a real JSON body
Response.redirect('/new-page') produces a real 302 Found with a real Location: /new-page header — confirmed two ways: urllib's own default redirect-following behavior transparently follows it (landing on a genuine 404 there, since /new-page was never actually registered in this demo — a real, honest confirmation that the redirect is genuinely followed end to end, not just returned), and a second request made through a custom opener with redirect-following disabled reads the raw 302 status and the exact /new-page Location header directly. Response.json({'id': 7, 'name': 'Ada'}) produces the real body b'{"id": 7, "name": "Ada"}' with Content-type: application/json, verified byte for byte.

Wiring Request & Response Into the Router

app() shrinks to three real cases — match, 405, 404 — and every handler below returns a Response instead of touching start_response:

router = Router() router.add_route('GET', '/', lambda req: Response('Home page'), name='home') router.add_route('GET', '/users/<int:id>', lambda req: Response(f"User {req.params['id']}, query={req.query}"), name='user_detail') router.add_route('POST', '/submit', lambda req: Response(f"Got form: {req.form()}"), name='submit') router.add_route('GET', '/old-page', lambda req: Response.redirect('/new-page'), name='old_page') router.add_route('GET', '/api/users/<int:id>', lambda req: Response.json({'id': req.params['id'], 'name': 'Ada'}), name='api_user') def app(environ, start_response): method = environ['REQUEST_METHOD'] path = environ.get('PATH_INFO', '/') status, handler, params, allowed = router.resolve(method, path) if status == 'MATCH': request = Request(environ, params) response = handler(request) return response.to_wsgi(start_response) elif status == '405': response = Response('405 Method Not Allowed', status=405, headers=[('Allow', ', '.join(sorted(allowed)))]) return response.to_wsgi(start_response) else: response = Response('404 Not Found', status=404) return response.to_wsgi(start_response)

Served through the real wsgiref server and hit with genuine urllib.request calls:

GET / → 200 b'Home page' GET /users/42?active=true → 200 b"User 42, query={'active': 'true'}" POST /submit → 200 b"Got form: {'name': 'Ada', 'age': '30'}" GET /api/users/7 → 200 b'{"id": 7, "name": "Ada"}'
Nothing in Chapter 2's own Router changed at all
Router.resolve() is completely unaware that its own third return value is about to become request.params, or that its handler is now expected to return a Response instead of a plain string. This is deliberate — Chapter 2's own class is a real, stable dependency this chapter builds on top of, not something that had to change to support it.

Where This Course Is Headed

Chapter 4 builds a real template engine — parsing, compiling, rendering — so a handler can return Response(render('page.html', title='Home')) instead of hand-assembling an HTML string with f-strings.

Hands-On Exercises

Exercise 1

Add a content_type property to this chapter's own Request class that reads the real, unprefixed CONTENT_TYPE key directly from environ, and verify it against a real live request that sends both a custom header and a real Content-Type header โ€” confirming content_type correctly returns the real value while headers (the HTTP_*-only property) still doesn't contain it.

๐Ÿ“„ View solution
Exercise 2

Using a real, live wsgiref server (not an in-memory BytesIO-backed environ), write a handler that reads environ['wsgi.input'] directly with a correctly-bounded read before ever constructing a Request, then calls request.body afterward โ€” using a short client-side timeout (e.g. 3 seconds) to confirm the request hangs rather than waiting indefinitely โ€” and explain concretely why the identical scenario tested against a BytesIO-backed fake environ would NOT reveal this bug at all.

๐Ÿ“„ View solution
Exercise 3

Add a Response.html() classmethod that behaves like Response.json() but sets content_type='text/html' instead, register a route returning Response.html('<h1>Hello</h1>'), and verify the real Content-type header a live client receives is exactly text/html rather than the default text/plain.

๐Ÿ“„ View solution

Chapter 3 Quick Reference

  • Request โ€” wraps environ + Router's own params dict; method, path, params as plain attributes
  • query โ€” a lazy, cached property parsing QUERY_STRING via urllib.parse.parse_qs
  • headers โ€” built from HTTP_* environ keys only; verified to genuinely exclude Content-Type/Content-Length, which PEP 3333 deliberately stores unprefixed
  • content_length / body โ€” content_length safely defaults to 0 on a real missing-header GET request; body caches its single real read from wsgi.input
  • The double-read danger โ€” verified directly: an in-memory BytesIO-backed test returns b'' instantly on a second read and looks completely safe; the identical code against a real live socket-backed server hangs indefinitely instead, confirmed with a forced timeout
  • Response โ€” status_line()/to_wsgi() compute Content-Length automatically; verified matching the real body length exactly, every time
  • redirect() / json() โ€” two real convenience constructors, verified against genuine 302/Location and application/json responses
  • Next chapter: A real template engine โ€” parsing, compiling & rendering