Deployment

Premier League Predictor: FastAPI & Redis

Chapter 11 · Deployment

Deploying the PostgreSQL sibling meant running a web server and configuring a database it connects to. Deploying this variant is the same plus one thing that changes the picture: Redis is a running server with its own defaults, and Chapter 9 already showed that those defaults lose data. This chapter covers the app side briefly — it follows the earlier siblings — and spends its time on the Redis side, where five things needed checking. Every result below was run: the app under gunicorn with four workers in a Linux container, talking to a real Redis 8.10.1 container.

The App Side

Gunicorn doesn't run on Windows, so it was run inside a python:3.13-slim container with Redis on the same Docker network. The command and the settings that change between environments:

# environment REDIS_URL=redis://plredis:6379/0 REDIS_MAX_CONNECTIONS=10 gunicorn app_prod:app -k uvicorn.workers.UvicornWorker -w 4 -b 0.0.0.0:8000

The static page is still served by the same app (StaticFiles, as in Chapter 1), so the page and the API share an origin and no CORS configuration exists to get wrong. The interactive API documentation is switched off unless a debug flag is set:

DEBUG = os.environ.get("APP_DEBUG") == "1" app = FastAPI(lifespan=lifespan, docs_url="/docs" if DEBUG else None, redoc_url=None, openapi_url="/openapi.json" if DEBUG else None)

Verified in the running container: /docs and /openapi.json both return 404. (Switching off /docs alone would have left the schema itself readable at /openapi.json.)

Connections: The Default Pool Fails Under Load

The PostgreSQL sibling's headline deployment finding was connection arithmetic: four workers, each with its own pool, opened up to 60 connections. Redis has a version of the same problem, and this one is specific to how this app uses Redis. Chapter 4's WATCH transaction — and the claim_both function from Chapter 10 — holds one connection exclusively from the WATCH until the EXEC. Each request doing that occupies a connection for the whole check-then-write gap.

redis-py 8.1.0, the version tested, gives every client a pool with a default max_connections of 100. To see what that does, 200 concurrent claims were fired, each holding its WATCH for 50 ms to mimic a realistic gap, while a monitor counted the server's connected clients:

Verified: Three Pools, 200 Concurrent WATCH Claims
pool peak clients succeeded failed default (max_connections 100) 101 100 100 MaxConnectionsError: Too many connections ConnectionPool(max_connections=20) 21 20 180 MaxConnectionsError: Too many connections BlockingConnectionPool(max=20) 21 200 0
(The peaks include the one monitoring connection.) A plain pool refuses the request that would exceed it, immediately. A BlockingConnectionPool makes the request wait for a free connection instead, and all 200 completed within the pool's limit of 20. For a request that briefly holds a connection, waiting is what is wanted.

The app therefore builds its pool explicitly, with a wait timeout so that a stuck pool becomes an error rather than a hang:

from redis.asyncio import BlockingConnectionPool @asynccontextmanager async def lifespan(app): pool = BlockingConnectionPool.from_url(REDIS_URL, max_connections=MAX_CONN, timeout=10, decode_responses=True) app.state.redis = redis.Redis(connection_pool=pool) yield await app.state.redis.aclose(close_connection_pool=True)
A Shutdown Detail That Cost a Wrong Number
A client given a pool of its own doesn't close that pool when it is closed. An earlier version of the test above called plain aclose() between runs and reported a peak of 41 instead of 21, because the previous run's connections were still open. Passing close_connection_pool=True fixed it, and the app's shutdown uses it for the same reason.

With REDIS_MAX_CONNECTIONS=10 and four workers, the ceiling is 40 connections. Under real gunicorn, 300 simultaneous fixture-creation requests (each doing the WATCH claim, the INCR and the transaction) all succeeded in 3.2 seconds, with a peak of 35 application connections — within the 40. Redis's own maxclients is 10,000 by default, so Redis isn't the limit here; the pool size is a control on how much waiting and memory the app tolerates, and it should be workers × per-worker maximum, kept well below maxclients.

Persistence, Set Deliberately

Chapter 9's finding is the reason this section exists: the default server lost everything written since its last snapshot on a hard kill, while the append-only file with appendfsync always survived. For a deployment where the fixtures and predictions are the only copy, the configuration is a decision to be made and written down, not left to defaults. The settings were exercised as command-line flags (--appendonly yes --appendfsync always); the same names go in redis.conf:

appendonly yes appendfsync always # survived a hard kill in Chapter 9; 'everysec' is the cheaper setting in between (not tested)

Memory Policy: Choose to Fail Loudly

The default is maxmemory 0 (no limit) with maxmemory-policy noeviction. Someone hardening the server, or copying a caching tutorial, might set a limit and an eviction policy like allkeys-lru. For a cache that is right; for source data it is not. With a deliberately tiny 2 MB limit and 4,000 writes of 500-byte values:

Verified: One Policy Refuses, the Other Deletes
noeviction: wrote 752 keys, then "command not allowed when used memory > 'maxmemory'" dbsize 752 fixture:0 still there allkeys-lru: wrote 4000 keys, no error at all dbsize 659 fixture:0 gone
With allkeys-lru every write reported success and 3,341 keys were silently discarded to make room, including the oldest fixture. With noeviction the server stopped accepting writes and said so. For this app, choose noeviction and, if a limit is set at all, leave plenty of headroom: ten seasons of data measured about 1.7 MiB in Chapter 9, so even a limit in the hundreds of megabytes is generous.

Security: The Defaults Are Open

Checked against the official image: bind is * -::* (all interfaces), protected-mode is no, and there is no password. With the port published, any client that can reach it can run any command:

Verified: Anyone Who Can Connect Can Wipe It
await r.flushall() # from the host, no credentials -> True

Two layers of defence, both worth having. First, don't expose Redis at all: in the gunicorn test the app reached it by container name over a private Docker network, and only the app's port needs to be published. Second, give the app its own account instead of the all-powerful default:

ACL SETUSER app on >apppw ~* &* +@all -@dangerous CONFIG SET requirepass adminpw
Verified: The App Still Works; the Dangerous Commands Don't
Running as that user, every command the app uses succeeded — INCR, HSET, SADD, ZADD, ZRANGE, a WATCH/MULTI/EXEC transaction, a Lua script (the Chapter 7 result script), PUBLISH, and SCAN. FLUSHALL, CONFIG GET and KEYS were each refused with "User app has no permissions to run the '...' command". After requirepass, a client with no credentials was refused outright (redis-py reports it as a failed HELLO, because it tries to authenticate as it connects).

Keys can be restricted too. A user limited to ~season:* and ~gameweek:* could SADD to a gameweek set and ZADD to a season table, but HSET fixture:1 and INCR fixture:next_id were refused, and a Lua script that touched an out-of-pattern key was refused partway through. The refusal did not undo the script's earlier, permitted write: the permitted key existed afterwards and the forbidden one didn't — Redis doesn't roll back a script any more than it rolls back a failed transaction (Chapter 3's finding; Exercise 2). The app's full key set is broad, so ~* is the practical setting here; the narrow form matters if a second service ever gets its own account. These were set with runtime ACL SETUSER commands; keeping them across restarts needs them in redis.conf or an ACL file, which wasn't tested.

Two Things That Turned Out Not to Be Problems

The Lua script cache. Every Lua script in this course is loaded into Redis and then called by its hash. A Redis restart or SCRIPT FLUSH empties that cache, which could break a running app with "no matching script" errors. It doesn't: after SCRIPT FLUSH, the same register_script object ran again and returned the right result (verified) — redis-py reloads the script when needed.

A Redis restart under a running app. With Redis stopped, a health route that does a PING returned 503 {"status": "redis unavailable"}; after Redis was started again, the route returned to 200 without restarting the app. The first check a few seconds after the restart was still 503 (Redis hadn't finished starting) and later checks were 200 — so a health check should be retried, not trusted once.

@app.get("/healthz") async def healthz(): try: await app.state.redis.ping() except RedisError: return JSONResponse({"status": "redis unavailable"}, status_code=503) return {"status": "ok"}

Backups

Persistence protects against a crash; it doesn't protect against a mistake or a lost disk. A snapshot file can be copied off the machine. The procedure was tested end to end: write data, run BGSAVE, wait for rdb_bgsave_in_progress to reach 0, copy /data/dump.rdb out, place it in a brand-new container's /data, and start it.

Verified: A Snapshot Restores a Whole Database
source: 202 keys (200 fixtures, a season table, the current-season pointer), dump.rdb 10,771 bytes restored: dbsize 202 | fixture:200 {'home_team_id': '1', 'away_team_id': '2'} | table top ['020'] | current season 1
A Backup Is Only As Recent As Its Snapshot
The snapshot holds what existed when BGSAVE ran. Fifty fixtures written afterwards were absent from the copied file: the source had 150 keys, the restore had 100 (Exercise 3). A backup job should run BGSAVE, wait for it, then copy — not just copy whatever file is there.
ConcernPostgreSQL siblingThis course (Redis)
Connection sizing4 workers × default pool = up to 60; sized down to 28Default pool is 100 per worker and fails at the limit; a blocking pool of 10 per worker peaked at 35
Crash safetyDefaultOff by default; needs appendonly yes
Out-of-memory behaviourErrorsDepends on policy; allkeys-lru deletes silently
Access controlDedicated roleDefault is open; an ACL user with -@dangerous
Backuppg_dumpBGSAVE + copy dump.rdb

Hands-On Exercises

Exercise 1

Run 200 concurrent WATCH-guarded claims (with a 50 ms hold) through a default pool, a ConnectionPool with max_connections=20, and a BlockingConnectionPool with max_connections=20. Report how many succeed, how many fail and with what error, and the peak client count for each.

📄 View solution
Exercise 2

Create an ACL user limited to keys matching season:* and gameweek:*. Try SADD on a gameweek key, ZADD on a season key, HSET on a fixture key, INCR on fixture:next_id, and a Lua script that writes one permitted and one forbidden key. Report each result, and whether the permitted write inside the script took effect.

📄 View solution
Exercise 3

Write 100 fixtures, run BGSAVE and wait for it, write 50 more, then copy dump.rdb into a fresh container and start it. Report the live key count and the restored key count, and explain what a correct backup job does differently.

📄 View solution

Chapter 11 Quick Reference

  • gunicorn + UvicornWorker, settings from the environment — page and API on one origin; /docs and /openapi.json both off (verified 404)
  • Verified: the default pool fails at its limit — 100 of 200 concurrent WATCH claims got MaxConnectionsError; a BlockingConnectionPool completed all 200
  • Pool = workers × per-worker max — 4 × 10 held 300 concurrent requests at a peak of 35 connections
  • Close the pool you built — aclose(close_connection_pool=True)
  • Persistence is a decision — appendonly yes (Chapter 9: survived a hard kill)
  • Verified: allkeys-lru silently deletes source data — keep noeviction so full memory is a loud error
  • Verified: the defaults are open — no password, all interfaces; keep it off public networks and give the app an ACL user with -@dangerous
  • Lua scripts and Redis restarts are fine — the script cache reloads itself; health checks should be retried after a restart
  • Backups: BGSAVE, wait, copy — verified restore; a stale snapshot restored 100 keys of 150