The Conversion Engine

Romaji to Kana Converter: Angular & Express

Chapter 2 · The Conversion Engine

Chapter 1 decided the conversion algorithm will eventually live on a real Express API. This chapter deliberately doesn't wire anything up to Angular or Express yet — it builds the engine itself as a plain, framework-agnostic module, tested against real input in isolation first. Chapter 4 decides how Angular's own side organizes around it; Chapter 5 moves the canonical copy into Express. Getting the actual algorithm correct comes first.

The Gojūon Table: Japan's Real 5×10 Sound Grid

Japanese kana are built around a real, systematic structure — five vowels (a, i, u, e, o) crossed with nine consonant rows, plus one standalone sound (ん/ン, "n"). A handful of cells are irregular under Hepburn romanization — し is "shi" not "si," ち is "chi" not "ti," つ is "tsu" not "tu," ふ is "fu" not "hu" — genuine historical exceptions baked into the standard system itself, not something this course invented:

// conversion-engine.ts — the base gojūon grid export const KANA_MAP: Record<string, { hiragana: string; katakana: string }> = { // Vowels a: { hiragana: 'あ', katakana: 'ア' }, i: { hiragana: 'い', katakana: 'イ' }, u: { hiragana: 'う', katakana: 'ウ' }, e: { hiragana: 'え', katakana: 'エ' }, o: { hiragana: 'お', katakana: 'オ' }, // K row ka: { hiragana: 'か', katakana: 'カ' }, ki: { hiragana: 'き', katakana: 'キ' }, ku: { hiragana: 'く', katakana: 'ク' }, ke: { hiragana: 'け', katakana: 'ケ' }, ko: { hiragana: 'こ', katakana: 'コ' }, // S row — shi is the real Hepburn exception, not "si" sa: { hiragana: 'さ', katakana: 'サ' }, shi: { hiragana: 'し', katakana: 'シ' }, su: { hiragana: 'す', katakana: 'ス' }, se: { hiragana: 'せ', katakana: 'セ' }, so: { hiragana: 'そ', katakana: 'ソ' }, // T row — chi and tsu are the real exceptions ta: { hiragana: 'た', katakana: 'タ' }, chi: { hiragana: 'ち', katakana: 'チ' }, tsu: { hiragana: 'つ', katakana: 'ツ' }, te: { hiragana: 'て', katakana: 'テ' }, to: { hiragana: 'と', katakana: 'ト' }, // N row na: { hiragana: 'な', katakana: 'ナ' }, ni: { hiragana: 'に', katakana: 'ニ' }, nu: { hiragana: 'ぬ', katakana: 'ヌ' }, ne: { hiragana: 'ね', katakana: 'ネ' }, no: { hiragana: 'の', katakana: 'ノ' }, // H row — fu is the real exception, not "hu" ha: { hiragana: 'は', katakana: 'ハ' }, hi: { hiragana: 'ひ', katakana: 'ヒ' }, fu: { hiragana: 'ふ', katakana: 'フ' }, he: { hiragana: 'へ', katakana: 'ヘ' }, ho: { hiragana: 'ほ', katakana: 'ホ' }, // M row ma: { hiragana: 'ま', katakana: 'マ' }, mi: { hiragana: 'み', katakana: 'ミ' }, mu: { hiragana: 'む', katakana: 'ム' }, me: { hiragana: 'め', katakana: 'メ' }, mo: { hiragana: 'も', katakana: 'モ' }, // Y row — only three real cells exist (ya/yu/yo) ya: { hiragana: 'や', katakana: 'ヤ' }, yu: { hiragana: 'ゆ', katakana: 'ユ' }, yo: { hiragana: 'よ', katakana: 'ヨ' }, // R row ra: { hiragana: 'ら', katakana: 'ラ' }, ri: { hiragana: 'り', katakana: 'リ' }, ru: { hiragana: 'る', katakana: 'ル' }, re: { hiragana: 'れ', katakana: 'レ' }, ro: { hiragana: 'ろ', katakana: 'ロ' }, // W row — only two real cells; "wo" is the object-particle character を wa: { hiragana: 'わ', katakana: 'ワ' }, wo: { hiragana: 'を', katakana: 'ヲ' }, // The standalone syllabic n — its own real, single-character sound n: { hiragana: 'ん', katakana: 'ン' }, };

Voiced and Semi-Voiced Sounds

A dakuten mark (゛) or handakuten mark (゜) shifts a base consonant into a real, distinct sound — G/Z/D/B rows, and a P row from H. There's a real, genuine ambiguity worth naming honestly rather than hiding: standard Hepburn romanizes both じ and ぢ as "ji", and both ず and づ as "zu" — a flat romaji-keyed table simply cannot tell these apart. This table maps "ji"/"zu" to their far more common Z-row readings (じ/ず); the rarer ぢ/づ, mostly confined to specific compound-word contexts, aren't reachable through plain romaji input at all.

// conversion-engine.ts — added to the same KANA_MAP object // G row ga: { hiragana: 'が', katakana: 'ガ' }, gi: { hiragana: 'ぎ', katakana: 'ギ' }, gu: { hiragana: 'ぐ', katakana: 'グ' }, ge: { hiragana: 'げ', katakana: 'ゲ' }, go: { hiragana: 'ご', katakana: 'ゴ' }, // Z row — "ji"/"zu" deliberately resolve here, per the note above za: { hiragana: 'ざ', katakana: 'ザ' }, ji: { hiragana: 'じ', katakana: 'ジ' }, zu: { hiragana: 'ず', katakana: 'ズ' }, ze: { hiragana: 'ぜ', katakana: 'ゼ' }, zo: { hiragana: 'ぞ', katakana: 'ゾ' }, // D row — only de/do have their own unambiguous romaji da: { hiragana: 'だ', katakana: 'ダ' }, de: { hiragana: 'で', katakana: 'デ' }, do: { hiragana: 'ど', katakana: 'ド' }, // B row ba: { hiragana: 'ば', katakana: 'バ' }, bi: { hiragana: 'び', katakana: 'ビ' }, bu: { hiragana: 'ぶ', katakana: 'ブ' }, be: { hiragana: 'べ', katakana: 'ベ' }, bo: { hiragana: 'ぼ', katakana: 'ボ' }, // P row (semi-voiced) pa: { hiragana: 'ぱ', katakana: 'パ' }, pi: { hiragana: 'ぴ', katakana: 'ピ' }, pu: { hiragana: 'ぷ', katakana: 'プ' }, pe: { hiragana: 'ぺ', katakana: 'ペ' }, po: { hiragana: 'ぽ', katakana: 'ポ' },

Combination Sounds (Yōon)

A consonant fused with a small ゃ/ゅ/ょ produces a real third category — genuinely three-character romaji tokens like kya, sha, cha — with one more real Hepburn simplification: the sounds spelled じゃ/じゅ/じょ (from the Z row) romanize as ja/ju/jo, just two characters, not "jya."

// conversion-engine.ts — added to the same KANA_MAP object kya: { hiragana: 'きゃ', katakana: 'キャ' }, kyu: { hiragana: 'きゅ', katakana: 'キュ' }, kyo: { hiragana: 'きょ', katakana: 'キョ' }, sha: { hiragana: 'しゃ', katakana: 'シャ' }, shu: { hiragana: 'しゅ', katakana: 'シュ' }, sho: { hiragana: 'しょ', katakana: 'ショ' }, cha: { hiragana: 'ちゃ', katakana: 'チャ' }, chu: { hiragana: 'ちゅ', katakana: 'チュ' }, cho: { hiragana: 'ちょ', katakana: 'チョ' }, nya: { hiragana: 'にゃ', katakana: 'ニャ' }, nyu: { hiragana: 'にゅ', katakana: 'ニュ' }, nyo: { hiragana: 'にょ', katakana: 'ニョ' }, hya: { hiragana: 'ひゃ', katakana: 'ヒャ' }, hyu: { hiragana: 'ひゅ', katakana: 'ヒュ' }, hyo: { hiragana: 'ひょ', katakana: 'ヒョ' }, mya: { hiragana: 'みゃ', katakana: 'ミャ' }, myu: { hiragana: 'みゅ', katakana: 'ミュ' }, myo: { hiragana: 'みょ', katakana: 'ミョ' }, rya: { hiragana: 'りゃ', katakana: 'リャ' }, ryu: { hiragana: 'りゅ', katakana: 'リュ' }, ryo: { hiragana: 'りょ', katakana: 'リョ' }, gya: { hiragana: 'ぎゃ', katakana: 'ギャ' }, gyu: { hiragana: 'ぎゅ', katakana: 'ギュ' }, gyo: { hiragana: 'ぎょ', katakana: 'ギョ' }, ja: { hiragana: 'じゃ', katakana: 'ジャ' }, ju: { hiragana: 'じゅ', katakana: 'ジュ' }, jo: { hiragana: 'じょ', katakana: 'ジョ' }, bya: { hiragana: 'びゃ', katakana: 'ビャ' }, byu: { hiragana: 'びゅ', katakana: 'ビュ' }, byo: { hiragana: 'びょ', katakana: 'ビョ' }, pya: { hiragana: 'ぴゃ', katakana: 'ピャ' }, pyu: { hiragana: 'ぴゅ', katakana: 'ピュ' }, pyo: { hiragana: 'ぴょ', katakana: 'ピョ' },
Deliberately Not Handled Yet
Long vowels (chōon — real double-length vowel sounds) and sokuon (the small っ doubling the following consonant) both need real, dedicated handling of their own — Chapter 3's job specifically. An input containing either isn't converted correctly by anything built in this chapter; that's a scope boundary, not an oversight.

A Naive Two-Character Tokenizer — And a Real Bug

Romaji tokens genuinely vary in length — one character (a, n), two (ka, shi), or three (kya). A first, reasonable-looking tokenizer only ever tries two characters, then falls back to one:

// A naive tokenizer that never tries a 3-character token at all function convertNaiveV1(romaji: string): string { let hiragana = ''; let i = 0; while (i < romaji.length) { const two = romaji.slice(i, i + 2); const one = romaji.slice(i, i + 1); if (KANA_MAP[two]) { hiragana += KANA_MAP[two].hiragana; i += 2; } else if (KANA_MAP[one]) { hiragana += KANA_MAP[one].hiragana; i += 1; } else { hiragana += romaji[i]; i += 1; } } return hiragana; }

convertNaiveV1('kyaku') — a real word, 客 (customer) — walks through position 0 trying "ky" (not a valid token) and then "k" alone (also not valid, since only vowels and n are genuine one-character tokens), falling all the way back to emitting the raw letter k unconverted. The result is kやく — a broken mix of a literal Latin letter and two real kana, because the algorithm never even attempts a 3-character match at all.

Adding a Third Length Isn't Enough on Its Own

The obvious fix adds a length-3 check into the same loop — but the order those lengths are tried in matters just as much as trying them at all:

// Tries all three lengths — but in ascending order function convertNaiveV2(romaji: string): string { let hiragana = ''; let i = 0; while (i < romaji.length) { let matched = false; for (let len = 1; len <= 3; len++) { const chunk = romaji.slice(i, i + len); if (KANA_MAP[chunk]) { hiragana += KANA_MAP[chunk].hiragana; i += len; matched = true; break; } } if (!matched) { hiragana += romaji[i]; i += 1; } } return hiragana; }

convertNaiveV2('kyaku') actually works — "k" and "ky" both genuinely fail to match anything, so the loop reaches length 3, finds "kya", and produces the correct きゃく. But convertNaiveV2('na') — the single, ordinary syllable な — fails in a real, different way: at position 0, length 1 is tried first, and "n" is a genuinely valid one-character token on its own (ん). The loop matches it immediately and stops, never trying the longer, correct "na" at all. The remaining "a" is then converted separately, producing んあ — two syllables — instead of the single, correct な.

The Real Cause: A Valid Short Token That's Also a Prefix of a Longer One
n is unusual in this table — it's one of the only single-character tokens that's also a genuine prefix of several longer, different tokens (na, ni, nu, ne, no). Checking shortest-first means the algorithm can find a real, valid match and stop — before ever discovering a longer, equally real match starting at the exact same position. Trying "kya" never hit this problem only because no 1- or 2-character prefix of it happens to be a valid token in its own right.

The Real Fix: Always Check Longest-to-Shortest

Trying lengths in descending order guarantees the longest possible match is always found first — and trying a longer chunk that turns out not to match costs nothing, since the loop simply falls through to a shorter length next:

// conversion-engine.ts — the real, correct tokenizer const MAX_TOKEN_LENGTH = 3; export function convert(romaji: string): { hiragana: string; katakana: string } { let hiragana = ''; let katakana = ''; let i = 0; while (i < romaji.length) { let matched = false; for (let len = MAX_TOKEN_LENGTH; len >= 1; len--) { const chunk = romaji.slice(i, i + len).toLowerCase(); const entry = KANA_MAP[chunk]; if (entry) { hiragana += entry.hiragana; katakana += entry.katakana; i += len; matched = true; break; } } if (!matched) { hiragana += romaji[i]; katakana += romaji[i]; i += 1; } } return { hiragana, katakana }; }

Both real bugs are resolved by the identical, single change — checking length 3, then 2, then 1, rather than the other way around. convert('kyaku').hiragana still correctly produces きゃく, and convert('na').hiragana now correctly produces the single syllable な, with んあ never entering into it.

Computing Both Scripts in One Pass
hiragana and katakana are built up together, from the exact same tokenization decisions, in a single loop — not by calling two separate conversion functions. That guarantees the two outputs can never disagree about where one syllable ends and the next begins, since they're never tokenized independently in the first place.

Trying the Engine in Isolation

A quick script — not yet connected to Angular or Express — confirms the engine on three real words:

console.log(convert('sushi')); // { hiragana: 'すし', katakana: 'スシ' } console.log(convert('kyaku')); // { hiragana: 'きゃく', katakana: 'キャク' } console.log(convert('konnichiwa')); // { hiragana: 'こんにちわ', katakana: 'コンニチワ' }

That last one is worth pausing on. The real, formal greeting is written こんにちは — using the character は (normally "ha"), pronounced "wa" specifically in this one fossilized grammatical role. Typed literally as "konnichiwa," this engine — correctly, given only what it knows so far — produces こんにちわ instead, a genuinely common casual spelling that plenty of native speakers use interchangeably, but not the traditionally "correct" one. Recognizing when a trailing "wa" really means は is exactly the kind of particle exception Chapter 3 builds real handling for.

Where This Engine Will Actually Live

Nothing built in this chapter depends on Angular or Express — KANA_MAP and convert() are plain TypeScript, testable with nothing more than node or a browser console. Chapter 4 decides how the Angular side of this project organizes its own services around calling this logic; Chapter 5 moves this exact file, unchanged, into the real Express API this course exists to build.

Hands-On Exercises

Exercise 1

Build the complete KANA_MAP (vowels through yōon) and the final, correct convert() function, then test it against "sushi", "kyaku", and "gakusei" (学生, student), confirming correct hiragana and katakana output for all three.

📄 View solution
Exercise 2

Build convertNaiveV1 (the fixed 2-then-1-character tokenizer) exactly as this chapter describes. Run it against "kyaku" and confirm the real "kやく" output, then run it against "kaki" (a word with no 3-character exception tokens at all) and confirm it works correctly there — showing the bug specifically depends on a 3-character token being present, not a universal failure. (Note: "sushi" is not a safe comparison word for this, since "shi" is itself one of the real 3-character exception tokens and would also break.)

📄 View solution
Exercise 3

Build convertNaiveV2 (checking lengths 1, 2, 3 in that order) and confirm it produces the correct result for "kyaku" but the wrong result ("んあ") for "na". Then explain in your own words why "n" specifically is the token that exposes this bug, when most other single-character or short tokens wouldn't.

📄 View solution

Chapter 2 Quick Reference

  • KANA_MAP — the real gojūon grid, plus voiced/semi-voiced rows and yōon combinations, keyed by romaji
  • Known limitation — "ji"/"zu" can't distinguish じ/ぢ or ず/づ from romaji alone; this table resolves to the more common reading
  • Deliberately deferred — long vowels and sokuon (っ) both need dedicated handling, built in Chapter 3
  • Bug 1 — a tokenizer that never tries 3 characters at all breaks on real yōon combinations like "kya"
  • Bug 2 — even trying all three lengths, checking shortest-first can match a valid short token ("n") before ever considering a longer, correct one ("na")
  • The fix — always check lengths longest-to-shortest; a failed longer match costs nothing and simply falls through
  • Next chapter: Handling Real Edge Cases — long vowels, sokuon, and particle exceptions like は→"wa"