The Conversion Engine
Romaji to Kana Converter: Astro
Chapter 2 · The Conversion Engine
Chapter 1 deliberately left the client-vs-server question open for Chapter 5 to answer with real evidence.
Until then, nothing in this project needs to know where the conversion logic will eventually run — it just
needs to be correct. This chapter builds src/lib/convert.ts as a plain, framework-agnostic
module, tested directly with node, with no Astro page, component, or endpoint touching it yet.
The Gojūon Table: Japan's Real 5×10 Sound Grid
Japanese kana are built around a real, systematic structure — five vowels (a, i, u, e, o) crossed with nine consonant rows, plus one standalone sound (ん/ン, "n"). A handful of cells are irregular under Hepburn romanization — し is "shi" not "si," ち is "chi" not "ti," つ is "tsu" not "tu," ふ is "fu" not "hu" — genuine historical exceptions baked into the standard system itself, not something this course invented:
Voiced and Semi-Voiced Sounds
A dakuten mark (゛) or handakuten mark (゜) shifts a base consonant into a real, distinct sound —
G/Z/D/B rows, and a P row from H. There's a real, genuine ambiguity worth naming honestly rather than
hiding: standard Hepburn romanizes both じ and ぢ as "ji", and both ず and づ as
"zu" — a flat romaji-keyed table simply cannot tell these apart. This table maps
"ji"/"zu" to their far more common Z-row readings (じ/ず); the rarer ぢ/づ, mostly
confined to specific compound-word contexts, aren't reachable through plain romaji input at all.
Combination Sounds (Yōon)
A consonant fused with a small ゃ/ゅ/ょ produces a real third category — genuinely three-character romaji
tokens like kya, sha, cha — with one more real Hepburn simplification:
the sounds spelled じゃ/じゅ/じょ (from the Z row) romanize as ja/ju/jo,
just two characters, not "jya."
A Naive Two-Character Tokenizer — And a Real Bug
Romaji tokens genuinely vary in length — one character (a, n), two
(ka, shi... except shi is actually three, which is exactly the trap
below), or three (kya). A first, reasonable-looking tokenizer only ever tries two characters,
then falls back to one:
Run for real against 'kyaku' — the ordinary word 客 (customer) — this produces
kやく, verified directly against real Node output, not just reasoned about in
the abstract. Position 0 tries "ky" (not a valid token) and then "k" alone (also
invalid, since only vowels and n are genuine one-character tokens), falls all the way through
to emitting the raw Latin letter k, and only then picks up "ya" and
"ku" correctly.
The base gojūon grid has its own, second, sneakier way to trigger the exact same class of bug — no yōon
needed at all. Run against 'sushi', the same real function produces
すsひ: "su" matches correctly, but "sh" then
"s" both fail, the raw letter s leaks through, and only then does
"hi" pick up the remainder — because shi itself is one of the base grid's own
genuine three-character tokens, this bug was never really about yōon specifically. It's about any token
three characters long, wherever it happens to sit in the table.
Adding a Third Length Isn't Enough on Its Own
The obvious fix adds a length-3 check into the same loop — but the order those lengths are tried in matters just as much as trying them at all:
convertNaiveV2('kyaku') now actually works — real, verified output きゃく —
since "k" and "ky" both genuinely fail to match, so the loop reaches length 3 and
finds "kya". But convertNaiveV2('na') — the single, ordinary syllable な — fails
in a real, different way, verified producing んあ instead of な. At position
0, length 1 is tried first, and "n" is a genuinely valid one-character token on its
own (ん). The loop matches it immediately and stops, never trying the longer, correct "na" at
that same position at all.
n is unusual in this table — it's one of the only single-character tokens that's also a
genuine prefix of several longer, different tokens (na, ni, nu,
ne, no). Checking shortest-first means the algorithm can find a real, valid match
and stop — before ever discovering a longer, equally real match starting at the exact same position. Trying
"kya" never hit this problem only because no 1- or 2-character prefix of it happens to be a
valid token in its own right.
The Real Fix: Always Check Longest-to-Shortest
Trying lengths in descending order guarantees the longest possible match is always found first — and trying a longer chunk that turns out not to match costs nothing, since the loop simply falls through to a shorter length next:
Both real bugs are resolved by the identical, single change — checking length 3, then 2, then 1, rather
than the other way around. Re-run against every case that broke earlier, this one function correctly
produces きゃく for "kyaku," すし (not すsひ) for "sushi," and
な (not んあ) for "na."
hiragana and katakana are built up together, from the exact same tokenization
decisions, in a single loop — not by calling two separate conversion functions. That guarantees the two
outputs can never disagree about where one syllable ends and the next begins, since they're never tokenized
independently in the first place.
Trying the Engine in Isolation
A quick script — not yet imported into any .astro page or component — confirms the engine on
real words, run directly with node before anything else in this project touches it:
That last one is worth pausing on. The real, formal greeting is written こんにちは — using the character は (normally "ha"), pronounced "wa" specifically in this one fossilized grammatical role. Typed literally as "konnichiwa," this engine — correctly, given only what it knows so far — produces こんにちわ instead, a genuinely common casual spelling that plenty of native speakers use interchangeably, but not the traditionally "correct" one. Recognizing when a trailing "wa" really means は is exactly the kind of particle exception Chapter 3 builds real handling for.
Where This Engine Will Actually Live
Nothing built in this chapter depends on Astro at all — KANA_MAP and convert() are
plain TypeScript, sitting in src/lib/convert.ts exactly where Chapter 1's own project structure
already reserved a place for them, testable with nothing more than node or a browser console.
Chapter 4 imports this exact, unmodified module directly into the real input/output UI; Chapter 5 is where
it gets a genuine server-side alternative to be measured against, rather than being assumed to belong on one
side or the other.
Hands-On Exercises
Build the complete KANA_MAP (vowels through yōon) and the final, correct convert() function in src/lib/convert.ts, then run it against "sushi", "kyaku", and "gakusei" (学生, student), confirming correct hiragana and katakana output for all three against real node output.
📄 View solutionBuild convertNaiveV1 (the fixed 2-then-1-character tokenizer) exactly as this chapter describes. Run it against "kyaku" and confirm the real "kやく" output, then run it against "kaki" (a word with no 3-character tokens anywhere in it) and confirm it produces the correct result there — showing the bug specifically depends on a 3-character token being present at that position, not a universal failure.
📄 View solutionBuild convertNaiveV2 (checking lengths 1, 2, 3 in that order) and confirm it produces the correct result for "kyaku" but the wrong result ("んあ") for "na". Then explain in your own words why "n" specifically is the token that exposes this bug, when most other single-character or short tokens wouldn't.
📄 View solutionChapter 2 Quick Reference
- KANA_MAP — the real gojūon grid, plus voiced/semi-voiced rows and yōon combinations, keyed by romaji, in src/lib/convert.ts
- Known limitation — "ji"/"zu" can't distinguish じ/ぢ or ず/づ from romaji alone; this table resolves to the more common reading
- Deliberately deferred — long vowels and sokuon (っ) both need dedicated handling, built in Chapter 3
- Bug 1 — a tokenizer that never tries 3 characters breaks on any 3-character token, yōon or not — verified on both "kyaku" (kやく) and the base-grid case "sushi" (すsひ)
- Bug 2 — even trying all three lengths, checking shortest-first can match a valid short token ("n") before ever considering a longer, correct one ("na")
- The fix — always check lengths longest-to-shortest; a failed longer match costs nothing and simply falls through
- Still framework-agnostic — this module has no Astro dependency yet, by design, keeping Chapter 5's own client-vs-server comparison genuinely open
- Next chapter: Handling Real Edge Cases — long vowels, sokuon, and particle exceptions like は→"wa"