The Conversion Engine

Romaji to Kana Converter: Astro

Chapter 2 · The Conversion Engine

Chapter 1 deliberately left the client-vs-server question open for Chapter 5 to answer with real evidence. Until then, nothing in this project needs to know where the conversion logic will eventually run — it just needs to be correct. This chapter builds src/lib/convert.ts as a plain, framework-agnostic module, tested directly with node, with no Astro page, component, or endpoint touching it yet.

The Gojūon Table: Japan's Real 5×10 Sound Grid

Japanese kana are built around a real, systematic structure — five vowels (a, i, u, e, o) crossed with nine consonant rows, plus one standalone sound (ん/ン, "n"). A handful of cells are irregular under Hepburn romanization — し is "shi" not "si," ち is "chi" not "ti," つ is "tsu" not "tu," ふ is "fu" not "hu" — genuine historical exceptions baked into the standard system itself, not something this course invented:

// src/lib/convert.ts — the base gojūon grid export const KANA_MAP: Record<string, { hiragana: string; katakana: string }> = { // Vowels a: { hiragana: 'あ', katakana: 'ア' }, i: { hiragana: 'い', katakana: 'イ' }, u: { hiragana: 'う', katakana: 'ウ' }, e: { hiragana: 'え', katakana: 'エ' }, o: { hiragana: 'お', katakana: 'オ' }, // K row ka: { hiragana: 'か', katakana: 'カ' }, ki: { hiragana: 'き', katakana: 'キ' }, ku: { hiragana: 'く', katakana: 'ク' }, ke: { hiragana: 'け', katakana: 'ケ' }, ko: { hiragana: 'こ', katakana: 'コ' }, // S row — shi is the real Hepburn exception, not "si" sa: { hiragana: 'さ', katakana: 'サ' }, shi: { hiragana: 'し', katakana: 'シ' }, su: { hiragana: 'す', katakana: 'ス' }, se: { hiragana: 'せ', katakana: 'セ' }, so: { hiragana: 'そ', katakana: 'ソ' }, // T row — chi and tsu are the real exceptions ta: { hiragana: 'た', katakana: 'タ' }, chi: { hiragana: 'ち', katakana: 'チ' }, tsu: { hiragana: 'つ', katakana: 'ツ' }, te: { hiragana: 'て', katakana: 'テ' }, to: { hiragana: 'と', katakana: 'ト' }, // N row na: { hiragana: 'な', katakana: 'ナ' }, ni: { hiragana: 'に', katakana: 'ニ' }, nu: { hiragana: 'ぬ', katakana: 'ヌ' }, ne: { hiragana: 'ね', katakana: 'ネ' }, no: { hiragana: 'の', katakana: 'ノ' }, // H row — fu is the real exception, not "hu" ha: { hiragana: 'は', katakana: 'ハ' }, hi: { hiragana: 'ひ', katakana: 'ヒ' }, fu: { hiragana: 'ふ', katakana: 'フ' }, he: { hiragana: 'へ', katakana: 'ヘ' }, ho: { hiragana: 'ほ', katakana: 'ホ' }, // M row ma: { hiragana: 'ま', katakana: 'マ' }, mi: { hiragana: 'み', katakana: 'ミ' }, mu: { hiragana: 'む', katakana: 'ム' }, me: { hiragana: 'め', katakana: 'メ' }, mo: { hiragana: 'も', katakana: 'モ' }, // Y row — only three real cells exist (ya/yu/yo) ya: { hiragana: 'や', katakana: 'ヤ' }, yu: { hiragana: 'ゆ', katakana: 'ユ' }, yo: { hiragana: 'よ', katakana: 'ヨ' }, // R row ra: { hiragana: 'ら', katakana: 'ラ' }, ri: { hiragana: 'り', katakana: 'リ' }, ru: { hiragana: 'る', katakana: 'ル' }, re: { hiragana: 'れ', katakana: 'レ' }, ro: { hiragana: 'ろ', katakana: 'ロ' }, // W row — only two real cells; "wo" is the object-particle character を wa: { hiragana: 'わ', katakana: 'ワ' }, wo: { hiragana: 'を', katakana: 'ヲ' }, // The standalone syllabic n — its own real, single-character sound n: { hiragana: 'ん', katakana: 'ン' }, };

Voiced and Semi-Voiced Sounds

A dakuten mark (゛) or handakuten mark (゜) shifts a base consonant into a real, distinct sound — G/Z/D/B rows, and a P row from H. There's a real, genuine ambiguity worth naming honestly rather than hiding: standard Hepburn romanizes both じ and ぢ as "ji", and both ず and づ as "zu" — a flat romaji-keyed table simply cannot tell these apart. This table maps "ji"/"zu" to their far more common Z-row readings (じ/ず); the rarer ぢ/づ, mostly confined to specific compound-word contexts, aren't reachable through plain romaji input at all.

// src/lib/convert.ts — added to the same KANA_MAP object // G row ga: { hiragana: 'が', katakana: 'ガ' }, gi: { hiragana: 'ぎ', katakana: 'ギ' }, gu: { hiragana: 'ぐ', katakana: 'グ' }, ge: { hiragana: 'げ', katakana: 'ゲ' }, go: { hiragana: 'ご', katakana: 'ゴ' }, // Z row — "ji"/"zu" deliberately resolve here, per the note above za: { hiragana: 'ざ', katakana: 'ザ' }, ji: { hiragana: 'じ', katakana: 'ジ' }, zu: { hiragana: 'ず', katakana: 'ズ' }, ze: { hiragana: 'ぜ', katakana: 'ゼ' }, zo: { hiragana: 'ぞ', katakana: 'ゾ' }, // D row — only de/do have their own unambiguous romaji da: { hiragana: 'だ', katakana: 'ダ' }, de: { hiragana: 'で', katakana: 'デ' }, do: { hiragana: 'ど', katakana: 'ド' }, // B row ba: { hiragana: 'ば', katakana: 'バ' }, bi: { hiragana: 'び', katakana: 'ビ' }, bu: { hiragana: 'ぶ', katakana: 'ブ' }, be: { hiragana: 'べ', katakana: 'ベ' }, bo: { hiragana: 'ぼ', katakana: 'ボ' }, // P row (semi-voiced) pa: { hiragana: 'ぱ', katakana: 'パ' }, pi: { hiragana: 'ぴ', katakana: 'ピ' }, pu: { hiragana: 'ぷ', katakana: 'プ' }, pe: { hiragana: 'ぺ', katakana: 'ペ' }, po: { hiragana: 'ぽ', katakana: 'ポ' },

Combination Sounds (Yōon)

A consonant fused with a small ゃ/ゅ/ょ produces a real third category — genuinely three-character romaji tokens like kya, sha, cha — with one more real Hepburn simplification: the sounds spelled じゃ/じゅ/じょ (from the Z row) romanize as ja/ju/jo, just two characters, not "jya."

// src/lib/convert.ts — added to the same KANA_MAP object kya: { hiragana: 'きゃ', katakana: 'キャ' }, kyu: { hiragana: 'きゅ', katakana: 'キュ' }, kyo: { hiragana: 'きょ', katakana: 'キョ' }, sha: { hiragana: 'しゃ', katakana: 'シャ' }, shu: { hiragana: 'しゅ', katakana: 'シュ' }, sho: { hiragana: 'しょ', katakana: 'ショ' }, cha: { hiragana: 'ちゃ', katakana: 'チャ' }, chu: { hiragana: 'ちゅ', katakana: 'チュ' }, cho: { hiragana: 'ちょ', katakana: 'チョ' }, nya: { hiragana: 'にゃ', katakana: 'ニャ' }, nyu: { hiragana: 'にゅ', katakana: 'ニュ' }, nyo: { hiragana: 'にょ', katakana: 'ニョ' }, hya: { hiragana: 'ひゃ', katakana: 'ヒャ' }, hyu: { hiragana: 'ひゅ', katakana: 'ヒュ' }, hyo: { hiragana: 'ひょ', katakana: 'ヒョ' }, mya: { hiragana: 'みゃ', katakana: 'ミャ' }, myu: { hiragana: 'みゅ', katakana: 'ミュ' }, myo: { hiragana: 'みょ', katakana: 'ミョ' }, rya: { hiragana: 'りゃ', katakana: 'リャ' }, ryu: { hiragana: 'りゅ', katakana: 'リュ' }, ryo: { hiragana: 'りょ', katakana: 'リョ' }, gya: { hiragana: 'ぎゃ', katakana: 'ギャ' }, gyu: { hiragana: 'ぎゅ', katakana: 'ギュ' }, gyo: { hiragana: 'ぎょ', katakana: 'ギョ' }, ja: { hiragana: 'じゃ', katakana: 'ジャ' }, ju: { hiragana: 'じゅ', katakana: 'ジュ' }, jo: { hiragana: 'じょ', katakana: 'ジョ' }, bya: { hiragana: 'びゃ', katakana: 'ビャ' }, byu: { hiragana: 'びゅ', katakana: 'ビュ' }, byo: { hiragana: 'びょ', katakana: 'ビョ' }, pya: { hiragana: 'ぴゃ', katakana: 'ピャ' }, pyu: { hiragana: 'ぴゅ', katakana: 'ピュ' }, pyo: { hiragana: 'ぴょ', katakana: 'ピョ' },
Deliberately Not Handled Yet
Long vowels (chōon — real double-length vowel sounds) and sokuon (the small っ doubling the following consonant) both need real, dedicated handling of their own — Chapter 3's job specifically. An input containing either isn't converted correctly by anything built in this chapter; that's a scope boundary, not an oversight.

A Naive Two-Character Tokenizer — And a Real Bug

Romaji tokens genuinely vary in length — one character (a, n), two (ka, shi... except shi is actually three, which is exactly the trap below), or three (kya). A first, reasonable-looking tokenizer only ever tries two characters, then falls back to one:

// A naive tokenizer that never tries a 3-character token at all function convertNaiveV1(romaji: string): string { let hiragana = ''; let i = 0; while (i < romaji.length) { const two = romaji.slice(i, i + 2); const one = romaji.slice(i, i + 1); if (KANA_MAP[two]) { hiragana += KANA_MAP[two].hiragana; i += 2; } else if (KANA_MAP[one]) { hiragana += KANA_MAP[one].hiragana; i += 1; } else { hiragana += romaji[i]; i += 1; } } return hiragana; }

Run for real against 'kyaku' — the ordinary word 客 (customer) — this produces kやく, verified directly against real Node output, not just reasoned about in the abstract. Position 0 tries "ky" (not a valid token) and then "k" alone (also invalid, since only vowels and n are genuine one-character tokens), falls all the way through to emitting the raw Latin letter k, and only then picks up "ya" and "ku" correctly.

The base gojūon grid has its own, second, sneakier way to trigger the exact same class of bug — no yōon needed at all. Run against 'sushi', the same real function produces すsひ: "su" matches correctly, but "sh" then "s" both fail, the raw letter s leaks through, and only then does "hi" pick up the remainder — because shi itself is one of the base grid's own genuine three-character tokens, this bug was never really about yōon specifically. It's about any token three characters long, wherever it happens to sit in the table.

Adding a Third Length Isn't Enough on Its Own

The obvious fix adds a length-3 check into the same loop — but the order those lengths are tried in matters just as much as trying them at all:

// Tries all three lengths — but in ascending order function convertNaiveV2(romaji: string): string { let hiragana = ''; let i = 0; while (i < romaji.length) { let matched = false; for (let len = 1; len <= 3; len++) { const chunk = romaji.slice(i, i + len); if (KANA_MAP[chunk]) { hiragana += KANA_MAP[chunk].hiragana; i += len; matched = true; break; } } if (!matched) { hiragana += romaji[i]; i += 1; } } return hiragana; }

convertNaiveV2('kyaku') now actually works — real, verified output きゃく — since "k" and "ky" both genuinely fail to match, so the loop reaches length 3 and finds "kya". But convertNaiveV2('na') — the single, ordinary syllable な — fails in a real, different way, verified producing んあ instead of な. At position 0, length 1 is tried first, and "n" is a genuinely valid one-character token on its own (ん). The loop matches it immediately and stops, never trying the longer, correct "na" at that same position at all.

The Real Cause: A Valid Short Token That's Also a Prefix of a Longer One
n is unusual in this table — it's one of the only single-character tokens that's also a genuine prefix of several longer, different tokens (na, ni, nu, ne, no). Checking shortest-first means the algorithm can find a real, valid match and stop — before ever discovering a longer, equally real match starting at the exact same position. Trying "kya" never hit this problem only because no 1- or 2-character prefix of it happens to be a valid token in its own right.

The Real Fix: Always Check Longest-to-Shortest

Trying lengths in descending order guarantees the longest possible match is always found first — and trying a longer chunk that turns out not to match costs nothing, since the loop simply falls through to a shorter length next:

// src/lib/convert.ts — the real, correct tokenizer const MAX_TOKEN_LENGTH = 3; export function convert(romaji: string): { hiragana: string; katakana: string } { let hiragana = ''; let katakana = ''; let i = 0; while (i < romaji.length) { let matched = false; for (let len = MAX_TOKEN_LENGTH; len >= 1; len--) { const chunk = romaji.slice(i, i + len).toLowerCase(); const entry = KANA_MAP[chunk]; if (entry) { hiragana += entry.hiragana; katakana += entry.katakana; i += len; matched = true; break; } } if (!matched) { hiragana += romaji[i]; katakana += romaji[i]; i += 1; } } return { hiragana, katakana }; }

Both real bugs are resolved by the identical, single change — checking length 3, then 2, then 1, rather than the other way around. Re-run against every case that broke earlier, this one function correctly produces きゃく for "kyaku," すし (not すsひ) for "sushi," and な (not んあ) for "na."

Computing Both Scripts in One Pass
hiragana and katakana are built up together, from the exact same tokenization decisions, in a single loop — not by calling two separate conversion functions. That guarantees the two outputs can never disagree about where one syllable ends and the next begins, since they're never tokenized independently in the first place.

Trying the Engine in Isolation

A quick script — not yet imported into any .astro page or component — confirms the engine on real words, run directly with node before anything else in this project touches it:

console.log(convert('sushi')); // { hiragana: 'すし', katakana: 'スシ' } console.log(convert('gakusei')); // 学生 (student) — { hiragana: 'がくせい', katakana: 'ガクセイ' } console.log(convert('konnichiwa')); // { hiragana: 'こんにちわ', katakana: 'コンニチワ' }

That last one is worth pausing on. The real, formal greeting is written こんにちは — using the character は (normally "ha"), pronounced "wa" specifically in this one fossilized grammatical role. Typed literally as "konnichiwa," this engine — correctly, given only what it knows so far — produces こんにちわ instead, a genuinely common casual spelling that plenty of native speakers use interchangeably, but not the traditionally "correct" one. Recognizing when a trailing "wa" really means は is exactly the kind of particle exception Chapter 3 builds real handling for.

Where This Engine Will Actually Live

Nothing built in this chapter depends on Astro at all — KANA_MAP and convert() are plain TypeScript, sitting in src/lib/convert.ts exactly where Chapter 1's own project structure already reserved a place for them, testable with nothing more than node or a browser console. Chapter 4 imports this exact, unmodified module directly into the real input/output UI; Chapter 5 is where it gets a genuine server-side alternative to be measured against, rather than being assumed to belong on one side or the other.

Hands-On Exercises

Exercise 1

Build the complete KANA_MAP (vowels through yōon) and the final, correct convert() function in src/lib/convert.ts, then run it against "sushi", "kyaku", and "gakusei" (学生, student), confirming correct hiragana and katakana output for all three against real node output.

📄 View solution
Exercise 2

Build convertNaiveV1 (the fixed 2-then-1-character tokenizer) exactly as this chapter describes. Run it against "kyaku" and confirm the real "kやく" output, then run it against "kaki" (a word with no 3-character tokens anywhere in it) and confirm it produces the correct result there — showing the bug specifically depends on a 3-character token being present at that position, not a universal failure.

📄 View solution
Exercise 3

Build convertNaiveV2 (checking lengths 1, 2, 3 in that order) and confirm it produces the correct result for "kyaku" but the wrong result ("んあ") for "na". Then explain in your own words why "n" specifically is the token that exposes this bug, when most other single-character or short tokens wouldn't.

📄 View solution

Chapter 2 Quick Reference

  • KANA_MAP — the real gojūon grid, plus voiced/semi-voiced rows and yōon combinations, keyed by romaji, in src/lib/convert.ts
  • Known limitation — "ji"/"zu" can't distinguish じ/ぢ or ず/づ from romaji alone; this table resolves to the more common reading
  • Deliberately deferred — long vowels and sokuon (っ) both need dedicated handling, built in Chapter 3
  • Bug 1 — a tokenizer that never tries 3 characters breaks on any 3-character token, yōon or not — verified on both "kyaku" (kやく) and the base-grid case "sushi" (すsひ)
  • Bug 2 — even trying all three lengths, checking shortest-first can match a valid short token ("n") before ever considering a longer, correct one ("na")
  • The fix — always check lengths longest-to-shortest; a failed longer match costs nothing and simply falls through
  • Still framework-agnostic — this module has no Astro dependency yet, by design, keeping Chapter 5's own client-vs-server comparison genuinely open
  • Next chapter: Handling Real Edge Cases — long vowels, sokuon, and particle exceptions like は→"wa"