The Conversion Engine
Romaji to Kana Converter: React & Next.js
Chapter 2 · The Conversion Engine
Before this variant touches a single Route Handler or a single React component, it needs the same thing both sibling courses started with: a real, complete mapping table from romaji to kana, and a tokenizer that actually walks a string correctly. This module is deliberately kept plain, framework-agnostic TypeScript with zero Next.js imports anywhere in it — Chapter 4's own real decision about whether this code runs in the browser or behind a Route Handler shouldn't have to touch this file at all, whichever way it goes.
A Real Gojūon-Plus-Yōon Mapping Table
src/lib/convert.ts starts with a real lookup table covering the standard gojūon
grid, the voiced (dakuten) and semi-voiced (handakuten) rows, and the yōon combination sounds —
each romaji key mapped to a real [hiragana, katakana] pair:
The full table (roughly 100 entries once every row, every dakuten variant, and every yōon combination is filled in) shares the exact same key set both sibling courses already verified, since the underlying gojūon system doesn't change from one JavaScript framework to another. What does need re-verifying, independently, is the tokenizer that walks a typed string against this table — and it's worth doing that with real, executed code rather than assuming the same two bugs the Astro and Angular & Express courses each found will show up here identically.
A First Attempt: Never Trying a Three-Character Token
The table above has keys of three different lengths — one character (a,
n), two characters (ka, shi), and three characters
(kya, cha). A tokenizer has to check for all three, in the right
order, at every position in the string. A first, understandable attempt only checks two lengths:
Run against a genuine 3-character yōon word, the failure is immediate — verified directly in Node rather than assumed:
"ky" isn't a valid 2-character key, and "k" alone isn't a valid 1-character key either, so the
loop falls through to the literal-character fallback and emits a raw "k" — then correctly
matches "ya" on the next pass. But convertNaiveV1 doesn't only fail on words
containing an actual yōon combination. A word with no yōon in it at all can still trigger the
identical bug, because "shi" and "chi" are themselves real 3-character
keys in this table — not yōon combinations, just ordinary gojūon syllables that happen
to need three romaji letters:
A Second Attempt: Trying All Three Lengths, in the Wrong Order
The obvious fix is to try all three lengths. A second attempt does exactly that — but checks them shortest-first:
This resolves every case convertNaiveV1 got wrong — "kya," "sushi," and
"konnichiwa" all convert correctly now, since a 3-character match is genuinely attempted once
the shorter lengths fail. But checking length 1 first introduces a real, different bug of its
own: the standalone syllabic n token is a valid 1-character key and a
genuine prefix of every real 2-character token starting with "n" — na,
ni, nu, ne, no. Ascending order means the
shorter, wrong match wins every single time one of those tokens is actually intended:
convertNaiveV2 against "konnichiwa" — the exact word that broke
convertNaiveV1 a moment ago — shows the bug reaches well past one simple two-letter
example:
ko + n (the syllabic n) +
ni + chi + wa. Once the first "n" is correctly consumed
as the syllabic n, the very next character is the second "n" — and ascending order
matches that second "n" as its own standalone syllabic-n token too, before ever trying the real
2-character "ni" that should have started there. The result silently fragments a genuine kana
syllable (に) into a stray ん followed by a bare い. This isn't a variant of the "na" bug — it's
the identical root cause reaching into the middle of a real multi-syllable word.
The Fix: Check Longest-to-Shortest
Both bugs share one real fix — check 3-character keys first, then 2, then 1, so a longer valid match is always taken over a shorter one that happens to also exist:
Every case that broke either earlier version is re-verified correct against this one function:
convert('konnichiwa') now produces こんにちわ — the common, casual spelling — rather
than the traditionally "correct" こんにちは, where は is pronounced "wa" specifically because it's
functioning as a grammatical particle rather than the ordinary syllable は. A plain lookup table
has no way to know that from the letters alone; it's exactly the kind of は/へ/を particle
exception Chapter 3 is scoped to handle, alongside long vowels and sokuon — the same real
cliffhanger both sibling courses ended their own Chapter 2 on.
Hands-On Exercises
Run convertNaiveV1 against "chotto" and "gyoza." Predict the output before running it, then verify with real code and explain exactly where each one breaks.
Find a real word (not "na" and not "konnichiwa") that triggers convertNaiveV2's own ascending-order bug, verify it with real code, and explain why it triggers the bug.
Explain, in your own words, why convertNaiveV1's bug and convertNaiveV2's bug are genuinely different failures — not just two symptoms of "the tokenizer is broken" — and why the single longest-to-shortest fix resolves both at once regardless.
Chapter 2 Quick Reference
- KANA_MAP — a plain, framework-agnostic
Record<string, [string, string]>, gojūon + dakuten/handakuten + yōon - Bug 1 (never tries 3-char) — breaks on real yōon words ("kya") AND ordinary 3-letter gojūon syllables ("sushi," "konnichiwa")
- Bug 2 (ascending 1,2,3) — the standalone "n" token wins over any longer token starting with "n," with real downstream damage inside multi-syllable words
- Fix — check lengths longest-to-shortest (3, 2, 1), resolving both bugs with one change
- Forward reference — こんにちわ vs. こんにちは is a real particle exception, not yet handled; Chapter 3's own job