Looking Ahead: A Redis Kanji Cache

Romaji to Kana Converter: React & Next.js

Chapter 6 · Looking Ahead: A Redis Kanji Cache

Scope Note — Read This First
This chapter does not add kanji conversion to the app. Real kanji conversion is a genuinely different, harder problem than the gojūon/yōon table this course has built so far — it needs a large vocabulary dictionary, and resolving which kanji a given reading actually means depends on real surrounding context (many words share an identical kana reading but write to different kanji), the same kind of disambiguation a real Japanese input-method editor (IME) performs, not a simple lookup table. Building that is a real project of its own, deliberately left for later. What this chapter does do is answer a narrower, concrete question honestly, with real measured numbers: once a vocabulary-scale dictionary exists, where should it live, and how fast can it realistically be looked up from a Next.js API route?

Why KANA_MAP's Own Approach Won't Scale

Chapter 2's own KANA_MAP works as a plain, hardcoded TypeScript object because it's small — roughly a hundred entries covering the gojūon grid, dakuten/handakuten rows, and yōon combinations. It ships inside the app's own JavaScript bundle and costs nothing extra to look up.

A real kanji/vocabulary dictionary is a different scale of problem entirely. JMdict — the real, widely-used open Japanese-English dictionary project — holds well over 100,000 entries. Even a deliberately trimmed, common-words-only subset would realistically run into the tens of thousands. A structure that size can't reasonably be hardcoded into a TypeScript module the way KANA_MAP was; it needs to live somewhere external, and it needs a real, fast way to look a word up by its reading.

The Real Architectural Question

A Next.js API route deployed the way Chapter 7 will cover doesn't guarantee that an in-memory JavaScript object survives, or is shared, across every request it handles. Serverless function instances can be created and destroyed per request, and even where an instance does stay warm for a while, a production deployment typically runs several instances behind a load balancer — each with its own separate memory, none of them seeing what another one cached. An in-process Map, which was genuinely free and fast for KANA_MAP, isn't a reliable place to hold a large shared dictionary in that environment.

Redis is the standard real answer to exactly this problem — a single external store every function instance can reach over the network, so a lookup cached by one request is genuinely available to the next one, regardless of which instance handles it. This isn't a hypothetical fit either: Redis's own real, current documentation for its Vercel integration lists "Persistent session storage across serverless functions" as one of its stated real use cases, and both Redis Cloud (the official managed offering) and Upstash (used directly in Next.js's own official with-redis example) are real, verifiable managed providers built specifically for this deployment shape.

Where This Would Actually Sit
src/ ├── lib/ │ ├── convert.ts // existing: gojūon/yōon table, tokenizer │ └── kanjiCache.ts // future: Redis client + lookup helpers └── app/ └── api/ ├── convert/ │ └── route.ts // existing (Ch.4/5): kana-only conversion └── convert-kanji/ └── route.ts // future: calls kanjiCache.ts, not convert.ts
A future kanji-aware route would sit alongside the existing one, not replace it — the kana-only /api/convert route this course already built stays exactly as it is.

A Real Benchmark: File Scan vs. In-Process Map vs. Real Redis

Rather than guess at the numbers, a real, containerized Redis 7 instance was started locally (via Docker) and benchmarked directly against a fake but realistically-sized 50,000-entry vocabulary dictionary — genuinely tested with real code, not assumed. Three approaches were compared, each performing 200 real lookups:

// Method 1: naive cold scan -- re-read + re-parse the whole file, then Array.find(), per lookup for (const w of targetWords) { const arr = JSON.parse(fs.readFileSync(dictFile, 'utf8')); const found = arr.find(e => e.word === w); } // Method 2: warm in-process Map, built once const map = new Map(entries.map(e => [e.word, e])); for (const w of targetWords) { const found = map.get(w); } // Method 3: real Redis GET over real TCP, one round trip per lookup for (const w of targetWords) { const raw = await redis.get(`vocab:${w}`); }

Two runs, re-verified for consistency:

MethodPer-call cost (run 1)Per-call cost (run 2)
Naive re-scan of the 50,000-entry file41,134 µs42,476 µs
Warm in-process Map0.66 µs0.63 µs
Real Redis GET (localhost, containerized)962 µs926 µs
Real, Measured Findings
Redis beat the naive file re-scan by roughly 43-46× — a real, decisive win once the dictionary is genuinely large. But Redis lost to the in-process Map by roughly 1,400-1,500× — a real network round trip, even to a container running on the same machine, costs vastly more than a function call that never leaves the process. This is the exact same shape of finding Chapter 4 already made about convert() vs. fetch(): a network hop is the expensive part, not the actual work being done on the other end of it.

Why Reach for Redis Anyway, Given That Result

The in-process Map wins the benchmark, but it isn't actually available for free the way it looks in this single-process test. It only wins because this benchmark ran in one warm process that built the Map once and reused it — exactly the condition the previous section explained a real serverless deployment doesn't reliably provide. Redis isn't being chosen because it's fast in an absolute sense; it's chosen because it's the one option in this comparison that's actually still there on the next request, no matter which function instance handles it.

The Round Trip Is (Almost) the Whole Cost

A further real test makes that concrete. Fetching the same 200 keys as 200 separate GET calls was compared against fetching all 200 in a single MGET call — one round trip carrying every key at once:

200 individual GETs total: 169,381.5 µs (846.91 µs/call) 1 MGET for all 200 keys: 2,010.4 µs total speedup: 84.3×

Batching the 200 lookups into one round trip cut the total time by 84×. Redis itself didn't get faster at doing the actual lookups — what changed is that the app paid the real network round-trip cost once instead of 200 times. That's a direct, practical echo of the same round-trip-dominates lesson this course already found for HTTP in Chapter 4, now confirmed for Redis specifically.

The Real Design Implication
A future kanji-aware route shouldn't call redis.get() once per candidate substring while scanning a sentence. It should identify every kanji-eligible substring first, then issue one redis.mget() call carrying all of them at once — turning what could be a dozen or more round trips per conversion request into exactly one.
An Honest Caveat on the Absolute Numbers
This benchmark's Redis container ran under Docker Desktop on Windows, which routes container networking through a lightweight virtualization layer rather than talking to Redis directly on the host's own network stack. The real ~0.85-0.96ms per-call figure measured here likely includes some genuine overhead from that layer specifically, and a managed provider like Redis Cloud or Upstash, or a native Linux deployment, could measure differently in absolute terms. The proportional finding — that batching into one round trip beats N individual round trips by roughly two orders of magnitude — is a structural property of "one network hop vs. many," not an artifact of this specific setup, and should hold regardless of exactly which numbers a different environment produces.

Hands-On Exercises

Exercise 1

Re-run this chapter's own three-way benchmark (file re-scan / in-process Map / real Redis) at a smaller scale — 5,000 dictionary entries instead of 50,000 — and explain, using the real numbers you measure, which method's own timing changes the most and which barely changes at all.

📄 View solution
Exercise 2

This chapter's own MGET test batched 200 keys into a single round trip. Real production sentences would realistically need far fewer lookups per request — closer to 5-15 kanji-eligible substrings per sentence. Re-run the individual-GET-vs-MGET comparison at a more realistic batch size of 10 keys, and explain whether the speedup still matters at that smaller scale.

📄 View solution
Exercise 3

Explain, in your own words, why this chapter's own benchmark result (Redis loses to an in-process Map) is not actually an argument against using Redis in the real deployment this course is building toward. What does the in-process Map's own win depend on that a real serverless deployment doesn't reliably provide?

📄 View solution

Chapter 6 Quick Reference

  • Kanji conversion itself is still deferred — this chapter is architecture and benchmarking only, not a new feature
  • Real measured result — Redis beats a naive file re-scan by ~43-46×, but loses to an in-process Map by ~1,400-1,500×
  • Why Redis anyway — a serverless deployment doesn't reliably share one process's own memory across every request; Redis Cloud and Upstash are both real, verified managed options built for exactly this
  • Batching wins — one real MGET for 200 keys measured 84× faster than 200 individual GET calls, the same round-trip-dominates lesson from Chapter 4