Handling Real Edge Cases

Romaji to Kana Converter: Angular & Express

Chapter 3 · Handling Real Edge Cases: Long Vowels, Sokuon (っ) & Particle Exceptions

Chapter 2's own closing warn-box named three things deliberately left unhandled: long vowels, the small っ (sokuon), and — surfaced by that chapter's very last test — the real は/へ/を particle exceptions. This chapter builds real, tested handling for all three. Every example below was run through an actual node script before being written down here; two of them exposed genuine bugs in the first attempt, both caught and fixed the same way Chapter 2's own tokenizer bugs were — by actually running the code rather than assuming it was correct.

Long Vowels: How Chōon Is Actually Spelled

A long (double-length) vowel sound is real and phonemically distinct in Japanese — おばさん (obasan, aunt) and おばあさん (obaasan, grandmother) are genuinely different words, not the same word said carefully. But hiragana and katakana don't spell a long vowel the same way, and hiragana itself doesn't even spell it the same way for every vowel:

Vowel RowReal Hiragana ConventionReal Exceptions
あ (a)Extend with あ — おかあさん (okaasan, mother)—
い (i)Extend with い — おにいさん (oniisan, older brother)—
う (u)Extend with う — くうき (kuuki, air)—
え (e)Usually extend with い — せんせい (sensei, teacher)Sometimes え — おねえさん (oneesan, older sister)
お (o)Usually extend with う — とうきょう (toukyou)Sometimes お — とおい (tooi, far), おおきい (ookii, big)

Katakana avoids this split entirely. A long vowel in katakana is always written with one mark, ー (chōonpu), regardless of which vowel is being extended — コーヒー (kōhī, coffee), ラーメン (rāmen, ramen). This isn't a simplification this course is inventing; it's the real, standard convention for how katakana (used overwhelmingly for loanwords) spells vowel length.

A Real, Unresolvable Ambiguity
Hepburn romanization with macrons (ā/ī/ū/ē/ō) collapses the え/い and お/う distinction on purpose — a macron just says "this vowel is long," not which of the two real hiragana spellings applies. Knowing whether a macron ō came from とう or とお requires knowing the specific word; it can't be recovered from the romaji alone. This converter makes a documented default choice below, and that choice will be wrong for real, correctly-spelled exception words like とおい and おおきい.

A First Attempt at Long Vowels — And a Real Bug

The plan: expand each macron character into an ASCII digraph before tokenizing, so the existing longest-match loop from Chapter 2 can do the rest. Hiragana gets the "best guess" native-word extension (え→い, お→う); katakana gets the same vowel doubled, since that's what will need collapsing into a single ー:

// conversion-engine.ts const MACRON_TO_HIRAGANA: Record<string, string> = { ā: 'aa', ī: 'ii', ū: 'uu', ē: 'ei', // best-guess default — real words vary, see the warn-box above ō: 'ou', // best-guess default — real words vary, see the warn-box above }; const MACRON_TO_KATAKANA: Record<string, string> = { ā: 'aa', ī: 'ii', ū: 'uu', ē: 'ee', ō: 'oo', }; function expandMacrons(romaji: string, table: Record<string, string>): string { let out = ''; for (const ch of romaji) out += table[ch.toLowerCase()] ?? ch; return out; }

Then, a first pass at turning a doubled katakana vowel into ー — scan the finished string, and whenever two consecutive characters share the same trailing vowel sound, collapse the second into ー:

// A post-processing pass over the ALREADY-BUILT katakana string function collapseChoonpuNaive(katakana: string, vowelOf: Record<string, string>): string { let out = ''; let prevVowel: string | null = null; for (const ch of katakana) { const vowel = vowelOf[ch]; if (vowel && vowel === prevVowel) { out += 'ー'; } else { out += ch; } prevVowel = vowel ?? null; } return out; }

collapseChoonpuNaive correctly turns ラアメン (from tokenizing "raamen") into ラーメン. But run the exact same function against the katakana for the ordinary word "ikimasu" (行きます, I go) — a word with no long vowel in it at all — and it produces イーマス instead of the correct イキマス.

The Real Cause: Comparing Characters, Not Syllables
イ (from i) and キ (from ki) are two entirely different, unrelated syllables that simply happen to both end in an "i" sound — the same coincidence that would make "kit" and "sit" rhyme in English without either word containing a doubled letter. collapseChoonpuNaive only looks at the two most recently emitted characters, with no memory of which whole token either one came from — so two unrelated syllables that happen to share a trailing vowel get merged exactly as if a genuine vowel had been typed twice in a row. A real vowel-lengthening collapse has to compare against the vowel of the last matched token, not the last accumulated character.

The Real Fix: Track the Vowel Inside the Tokenizer Itself

Rather than patching the finished string afterward, the chōonpu decision moves inside the same loop that's already matching tokens — it always knows exactly which real syllable it just emitted, so it can never confuse "ends in the same vowel sound" with "is a genuine repeated vowel":

// conversion-engine.ts — replaces the naive post-processing pass entirely function tokenizeScript(romaji: string, useChoonpu: boolean): string { let out = ''; let i = 0; let lastVowel: string | null = null; while (i < romaji.length) { let matched = false; for (let len = MAX_TOKEN_LENGTH; len >= 1; len--) { const chunk = romaji.slice(i, i + len).toLowerCase(); const entry = KANA_MAP[chunk]; if (entry) { // chunk.length, NOT the loop variable len — see the warn-box below const isBareVowel = chunk.length === 1 && 'aiueo'.includes(chunk); if (useChoonpu && isBareVowel && chunk === lastVowel) { out += 'ー'; } else { out += useChoonpu ? entry.katakana : entry.hiragana; } lastVowel = chunk === 'n' ? null : chunk[chunk.length - 1]; i += len; matched = true; break; } } if (!matched) { out += romaji[i]; lastVowel = null; i += 1; } } return out; }
A Second Real Bug, Found Testing "kōhī"
The first version of this function checked len === 1 instead of chunk.length === 1 to decide whether a match was a bare vowel. Near the end of a string, romaji.slice(i, i + 3) can genuinely return fewer than 3 characters — there's nothing left to slice — so a length-3 attempt at the very last character of "koohii" returns the 1-character string "i", which correctly matches the vowel token, but with len still equal to 3. Checking the loop variable instead of the actual matched string's own length made the function think this was a 3-character match, never recognizing it as the bare vowel it actually was — so the real second long vowel in "kōhī" silently failed to collapse, producing コーヒイ instead of the correct コーヒー. Switching the check to chunk.length — the true length of what was actually found, regardless of which loop iteration found it — fixed it.

Both scripts now need their own full pass, expanded with their own macron table, because they diverge on exactly the case this chapter is about:

console.log(tokenizeScript(expandMacrons('kōhī', MACRON_TO_HIRAGANA), false)); // こうひい console.log(tokenizeScript(expandMacrons('kōhī', MACRON_TO_KATAKANA), true)); // コーヒー console.log(tokenizeScript(expandMacrons('tōkyō', MACRON_TO_HIRAGANA), false)); // とうきょう — the real, correct spelling console.log(tokenizeScript(expandMacrons('tōi', MACRON_TO_HIRAGANA), false)); // とうい — the real word is spelled とおい; this is the honest limitation from the warn-box above
Chapter 2's "One Shared Loop" Promise Doesn't Survive Long Vowels
Chapter 2's own tip-box promised hiragana and katakana are always built "from the exact same tokenization decisions, in a single loop." Long vowels break that promise on purpose: とうきょう correctly needs a う in hiragana but ー in katakana at the identical position, so the two scripts now genuinely need two separate macron-expanded inputs and two separate tokenizer passes. The single-pass guarantee still holds for everything without a long vowel in it — which is most input — but it's no longer universal.

Sokuon: The Small っ

A doubled consonant in romaji (other than a doubled n) signals a small っ/ッ immediately before the syllable that follows — きって (kitte, stamp), がっこう (gakkou, school). The small tsu itself carries no sound of its own; it's a brief, silent held pause, then the following consonant releases normally:

// conversion-engine.ts const SMALL_TSU = { hiragana: 'っ', katakana: 'ッ' }; // Returns how many characters the sokuon trigger itself consumes (0 if none) — // the actual syllable that follows is left for the normal tokenizer loop. function sokuonLength(romaji: string, i: number): number { const c = romaji[i]; if (!c || 'aiueo'.includes(c)) return 0; // Real Hepburn irregularity: gemination before chi/cha/chu/cho is spelled // with a "t", not a doubled "c" — まっちゃ is "matcha," never "maccha." if (c === 't' && romaji.slice(i + 1, i + 3) === 'ch') return 1; // Doubled n is NOT sokuon — see こんにちは below. if (c !== 'n' && romaji[i + 1] === c) return 1; return 0; }

Wired into tokenizeScript, this is checked before every ordinary token attempt: if a sokuon trigger is found, emit the small tsu, advance past just the trigger, and let the loop pick up the real syllable from there on its very next iteration.

console.log(tokenizeScript('kitte', false)); // きって console.log(tokenizeScript('gakkou', false)); // がっこう console.log(tokenizeScript('matcha', false)); // まっちゃ — the real "tch" exception, verified console.log(tokenizeScript('kekkon', false)); // けっこん (結婚, marriage)
Why "matcha" Isn't "maccha"
Trace it by hand: at t, the next two characters are exactly "ch", so the special case fires and consumes only the t — emitting っ and leaving the string positioned right at "cha", which the ordinary tokenizer then matches as a normal 3-character token, ちゃ. The doubled-consonant rule alone (t followed by another t) would never have fired here at all, since the second letter is c, not t — this is a genuinely separate, irregular case, not a variation on the ordinary rule.

The doubled-n exclusion matters for a real, common word — こんにちは (konnichiwa, hello) is ko-n-ni-chi-wa, five real syllables with no small tsu anywhere in it. Treating a doubled n as sokuon would insert a phantom っ that doesn't belong:

console.log(tokenizeScript('konnichiwa', false)); // こんにちわ — correct so far as sokuon goes (no phantom っ), but still ends in // わ rather than は — that's a completely different, particle-level problem, next

The Real は/へ/を Particle Exceptions

Three hiragana keep an old, "wrong-looking" spelling in one specific grammatical role — the character is written one way but pronounced another:

CharacterOrdinary ReadingAs a ParticleReal Example
はhawa (topic marker)これは本です — kore wa hon desu
へhee (direction marker)これへ行きます — kore e ikimasu
をwoo (object marker)これをください — kore o kudasai

Our converter runs romaji into kana, so the direction that actually matters is: given typed "wa", "e", or "o", when should the output be は/へ/を instead of the ordinary わ/え/お? A workable, honest heuristic: only when that romaji appears as its own complete, space-separated word — never when it's a letter sequence inside a longer word:

// conversion-engine.ts const PARTICLE_OVERRIDES: Record<string, string> = { wa: 'ha', e: 'he', o: 'wo' }; export function convertSentence(romaji: string): { hiragana: string; katakana: string } { const words = romaji.split(' '); const hiraganaParts: string[] = []; const katakanaParts: string[] = []; for (const word of words) { const lower = word.toLowerCase(); const effective = PARTICLE_OVERRIDES[lower] ?? word; hiraganaParts.push(tokenizeScript(expandMacrons(effective, MACRON_TO_HIRAGANA), false)); katakanaParts.push(tokenizeScript(expandMacrons(effective, MACRON_TO_KATAKANA), true)); } return { hiragana: hiraganaParts.join(' '), katakana: katakanaParts.join(' ') }; }

This resolves the two grammatically unambiguous cases correctly, and correctly leaves わたし (watashi, I) — which genuinely contains the letters "wa" but never as its own word — completely untouched:

console.log(convertSentence('kore wa hon desu').hiragana); // これ は ほん です — the topic particle, correctly は console.log(convertSentence('kore o kudasai').hiragana); // これ を ください — the object particle, correctly を console.log(convertSentence('watashi wa gakusei desu').hiragana); // わたし は がくせい です — the "wa" inside watashi is left alone; only the // standalone word is overridden
Where This Heuristic Genuinely Breaks: A Real, Unresolvable Ambiguity
"wa" and "o" standing alone are almost always the particle — there's little else a lone "wa" or "o" would normally be. But standalone "e" is a real, common Japanese word in its own right — 絵 (e), meaning "picture." Compare two genuinely grammatical sentences:
  • "kore e ikimasu" ("I'm going this way") — "e" really is the へ particle
  • "kono e wa kirei desu" ("This picture is pretty") — "e" here means 絵, and should stay え, not へ
A word-boundary rule alone can't tell these two sentences apart — both have a standalone "e" in the identical grammatical position. This converter always resolves standalone "e" to へ, which is right for the first sentence and genuinely wrong for the second. Correctly distinguishing them would need real part-of-speech context — something well outside a lookup-table-driven converter — so this is documented as a known, permanent limitation rather than quietly hidden.

Verified against both real sentences directly — the second one confirms the honest limitation exists, not just that it's theoretically possible:

console.log(convertSentence('kore e ikimasu').hiragana); // これ へ いきます — correct console.log(convertSentence('kono e wa kirei desu').hiragana); // この へ は きれい です — "wa" correctly resolves to は, but "e" incorrectly // becomes へ instead of the correct え (this sentence is really about a picture)

こんにちは: A Whole-Word Exception, Not a Grammatical One

こんにちは itself needs a genuinely different fix from the general particle rule above. Its own は isn't a standalone particle word at all — it's fused inside one continuous greeting, historically built from "today, as for..." with the rest of the sentence traditionally left unsaid. The word-boundary heuristic can't reach it, because there's no space anywhere near that は to detect. This needs a small, explicit dictionary of whole-word exceptions, checked before anything else runs:

// conversion-engine.ts const LEXICALIZED_EXCEPTIONS: Record<string, { hiragana: string; katakana: string }> = { konnichiwa: { hiragana: 'こんにちは', katakana: 'コンニチハ' }, konbanwa: { hiragana: 'こんばんは', katakana: 'コンバンハ' }, };
Two Genuinely Different Kinds of Exception
This chapter now has two structurally different fixes living side by side: the general PARTICLE_OVERRIDES rule, which applies whenever a real grammatical pattern holds (a standalone word between two others), and LEXICALIZED_EXCEPTIONS, a small closed list of specific, memorized whole words whose irregular spelling can't be derived from any general rule at all. こんばんは (konbanwa, good evening) is a real, second example of the exact same lexicalized pattern — worth knowing the category exists, since a third such word could show up later and need the same treatment.

Putting It All Together

The final convert() checks the lexicalized exception dictionary first, applies the particle override to a standalone word second, then runs the full macron-expansion-plus-sokuon-aware tokenizer for each script — every layer from this chapter, combined:

// conversion-engine.ts — the real Chapter 3 convert() export function convert(romaji: string): { hiragana: string; katakana: string } { const words = romaji.split(' '); const hiraganaParts: string[] = []; const katakanaParts: string[] = []; for (const word of words) { const lower = word.toLowerCase(); const lexical = LEXICALIZED_EXCEPTIONS[lower]; if (lexical) { hiraganaParts.push(lexical.hiragana); katakanaParts.push(lexical.katakana); continue; } const effective = PARTICLE_OVERRIDES[lower] ?? word; hiraganaParts.push(tokenizeScript(expandMacrons(effective, MACRON_TO_HIRAGANA), false)); katakanaParts.push(tokenizeScript(expandMacrons(effective, MACRON_TO_KATAKANA), true)); } return { hiragana: hiraganaParts.join(' '), katakana: katakanaParts.join(' ') }; }

Every feature from this chapter, exercised together in one sentence — a long vowel from a native word, a doubled consonant, and a direction particle, all at once:

console.log(convert('gakkou e ikimasu')); // { hiragana: 'がっこう へ いきます', katakana: 'ガッコウ ヘ イキマス' } console.log(convert('konnichiwa')); // { hiragana: 'こんにちは', katakana: 'コンニチハ' } — finally correct, resolving // the exact cliffhanger Chapter 2 ended on

Hands-On Exercises

Exercise 1

Build sokuonLength exactly as this chapter describes, wire it into a tokenizer, and verify it against three real words: "zutto" (ずっと, all along), "issho" (いっしょ, together), and "matcha" (まっちゃ, matcha). Then explain, in your own words, why checking for "ch" specifically after a "t" has to happen before the general doubled-consonant check runs — what would go wrong if the order were reversed?

📄 View solution
Exercise 2

Build collapseChoonpuNaive (the character-comparison version) and confirm it works for "raamen" but produces "イーマス" for "ikimasu". Then build the fixed tokenizeScript version and confirm it correctly produces "イキマス" for the same word while still correctly producing "ラーメン" for "raamen".

📄 View solution
Exercise 3

Build convertSentence with the particle overrides, then run it against both "kore e ikimasu" and "kono e wa kirei desu". Confirm the first is correct and the second is wrong (produces へ where え would be correct), and write one sentence explaining why no purely lookup-table-based fix could resolve this specific ambiguity.

📄 View solution

Chapter 3 Quick Reference

  • Long vowels — hiragana extends あ/い/う-row vowels with themselves, but usually extends え-row with い and お-row with う (real exceptions exist: とおい, おおきい); katakana always uses one mark, ー
  • Bug 1 — collapsing a long vowel by comparing accumulated katakana characters (not matched tokens) wrongly merges unrelated syllables that happen to share a trailing vowel (イキ → イー)
  • Bug 2 — checking the loop variable len instead of chunk.length misses a bare-vowel match truncated by the string's own end
  • Sokuon — a doubled consonant (except doubled n) inserts っ/ッ before the following syllable; gemination before ち/ちゃ/ちゅ/ちょ is spelled with "t" ("matcha," never "maccha")
  • Particle exceptions — standalone "wa"/"e"/"o" override to は/へ/を; standalone "e" is genuinely ambiguous with 絵 (picture) and has no lookup-table fix
  • Lexicalized exceptions — こんにちは and こんばんは need a small whole-word dictionary, since their own は isn't a standalone particle at all
  • Next chapter: Angular Services & Components — structuring the conversion tool around this now-complete engine