Handling Real Edge Cases

Romaji to Kana Converter: React & Next.js

Chapter 3 · Handling Real Edge Cases

Chapter 2 ended on a real cliffhanger: convert('konnichiwa') produces こんにちわ, the common casual spelling, not the traditionally "correct" こんにちは — because a plain lookup table has no way to know that a trailing は should sometimes be pronounced "wa" rather than "ha." That's one of three genuine edge cases this chapter closes out: long vowels, sokuon (small っ), and は/へ/を particle exceptions. Every claim below was checked by actually running the code, including two real chōon bugs found along the way — one that reproduces a failure mode both sibling courses already documented, and a second, subtler one that's specific to how this variant's own approach is structured.

Long Vowels: Two Real Ways to Get This Wrong First

In doubled-vowel romaji notation, a repeated vowel letter (or the alternate spellings "ei" and "ou") signals a long vowel — katakana always marks this with a single chōonpu (ー); hiragana conventionally extends with a second vowel kana (あ for a, い for i, う for u, and by default い for e and う for o, with real, genuine exception words that don't follow the default, covered in the warn-box below).

Attempt 1: Collapse the Output, After the Fact

A tempting first approach runs after tokenization is already finished — walk the finished kana string, and whenever two adjacent characters "sound like" the same vowel, collapse the second one into a chōonpu:

function collapseChoonpuNaive(kanaStr: string): string { let out = '', prevVowel: string | null = null; for (const ch of kanaStr) { const v = KANA_VOWEL[ch]; // the vowel sound this kana character represents if (v && v === prevVowel) { out += 'ー'; } // BUG below else { out += ch; } prevVowel = v ?? null; } return out; }

Run against real, already-correct tokenizer output, this immediately corrupts a word that never contained a long vowel in the first place:

> convert('ikimasu') // raw tokenizer output, no chōon involved { hiragana: 'いきます', katakana: 'イキマス' } > collapseChoonpuNaive('イキマス') 'イーマス' // い and き just happen to both sound like "i" — wrongly collapsed
A Genuine Failure Mode, Verified — and Fixed by Working on the Input, Not the Output
い ("i") and き ("ki") aren't a real doubled vowel in "ikimasu" at all — they're two completely unrelated tokens that happen to end in the same vowel sound. Comparing the output kana characters' own vowel sounds can never distinguish "these two syllables coincidentally share a vowel" from "the user actually typed a doubled vowel letter here." The only reliable signal is in the original romaji input, not the finished kana.

Attempt 2: Check the Input — But in the Wrong Order

Moving the check into the tokenizer itself, working directly against the romaji input rather than the finished kana, looks like the right fix. A first version checks for a normal token match first, falling back to a chōon check only when nothing else matches:

// inside the main loop, checked AFTER the normal 3/2/1-length match attempt if (lastVowel) { const c = romaji[i]; if (c === lastVowel || (lastVowel === 'e' && c === 'i') || (lastVowel === 'o' && c === 'u')) { kata += 'ー'; hira += HIRA_EXT[lastVowel]; i += 1; continue; } }

Run against a real word with a genuine long vowel, this version detects nothing at all:

> convertWithChoonAttemptA('koohii') { hiragana: 'こおひい', katakana: 'コオヒイ' } // no ー anywhere — the chōon branch never runs
A Second, Genuinely Different Bug — Not Documented by Either Sibling Course
Every single vowel letter — a, i, u, e, o — is itself a valid, standalone 1-character key in KANA_MAP. Because the normal token-match loop runs first, that second, repeated vowel letter always gets matched as a brand-new ordinary kana character before the chōon-check branch is ever reached — the fallback code is real, correct, and completely unreachable. The one fix that actually works is checking for a chōon continuation before attempting a fresh token match, not after:
// src/lib/convert.ts — the real, verified order function tokenizeWord(word: string) { let hira = '', kata = '', i = 0; let lastVowel: string | null = null; while (i < word.length) { // 1. chōon continuation — checked FIRST if (lastVowel) { const c = word[i]; const extends_ = c === lastVowel || (lastVowel === 'e' && c === 'i') || (lastVowel === 'o' && c === 'u'); if (extends_) { kata += 'ー'; hira += HIRA_EXT[lastVowel]; i += 1; continue; // lastVowel deliberately unchanged } } // 2. sokuon check — see below // 3. normal longest-to-shortest token match — see Chapter 2 } }

Re-verified against every case that broke either earlier attempt:

> convert('ikimasu') { hiragana: 'いきます', katakana: 'イキマス' } > convert('okaasan') { hiragana: 'おかあさん', katakana: 'オカーサン' } > convert('sensee') { hiragana: 'せんせい', katakana: 'センセー' } > convert('obaasan') { hiragana: 'おばあさん', katakana: 'オバーサン' }
A Real, Honest Limitation
The hiragana default rule above extends an o-row long vowel with う, matching the majority convention (とうきょう, not とおきょう). A small number of genuine exception words — とおい (far), おおきい (big), おおい (many) — really are spelled with おお instead. No lookup rule can tell "koohii" (a default-rule word) apart from a real お-doubling exception from the romaji spelling alone; this converter's default will get those specific words wrong, exactly as both sibling courses' own converters do.

Sokuon: A Doubled Consonant Signals っ/ッ

A doubled consonant (not a vowel, not "n") signals a small っ before the following syllable — only the first of the pair is consumed as the sokuon signal; the second starts the real next token normally. Japanese has one genuine, real irregularity in how this is spelled before the ち-row: a sokuon before ち/ちゃ/ちゅ/ちょ is written as a single "t" followed by "ch," not by doubling the "ch" digraph itself:

const CONSONANTS = 'kstpghzjbcmry'.split(''); // checked after the chōon check, before the normal token-match loop const c0 = word[i], c1 = word[i + 1]; const isDoubledConsonant = c0 === c1 && CONSONANTS.includes(c0) && c0 !== 'n'; const isTchIrregular = c0 === 't' && word.slice(i + 1, i + 3) === 'ch'; if (isDoubledConsonant || isTchIrregular) { hira += 'っ'; kata += 'ッ'; i += 1; // consume only the first consonant; the second starts the next real token continue; }

Verified against a real word for each case, including the "nn" exclusion that must not trigger a sokuon:

> convert('matcha') { hiragana: 'まっちゃ', katakana: 'マッチャ' } // the real tch irregularity > convert('kitte') { hiragana: 'きって', katakana: 'キッテ' } // ordinary doubled consonant > convert('gakkou') { hiragana: 'がっこう', katakana: 'ガッコー' } // sokuon AND a chōon, in one word
A Real, Verified Multi-Feature Interaction
"gakkou" exercises the sokuon check ("kk") and the chōon check ("ou") in the same nine-character word, and both fire correctly against the exact same tokenizeWord function — the priority order established above (chōon check, then sokuon check, then normal match) holds up without any special-casing needed for words that use more than one edge case at once.

A Third Bug: The Syllabic ん Has No Real Vowel to Extend

Every normal token sets lastVowel to that token's own trailing letter, so the chōon check on the next loop iteration knows what vowel to look for a repeat of. The syllabic n token (n → ん) was, without thinking it through, treated the same way — its "trailing letter" is n itself, so lastVowel ends up set to the literal character 'n'. That's fine right up until a second n genuinely follows the first — a real, common pattern, not an edge case:

> convert('annai') // 案内, "guidance" — a real word with a genuine doubled n { hiragana: 'あんundefinedあい', katakana: 'アンーアイ' }
A Real, Severe Bug — Verified, Then Fixed
Once the first "n" is correctly matched as ん, lastVowel is literally the string 'n'. The very next character is a second, genuine "n" — and the chōon check's own c === lastVowel comparison treats 'n' === 'n' exactly like a doubled vowel, firing the chōon branch. HIRA_EXT has no 'n' key at all, so HIRA_EXT['n'] evaluates to undefined — and string concatenation with undefined in JavaScript silently produces the literal four-letter word "undefined" spliced directly into the output, a genuinely worse failure than a merely wrong kana character. The real fix is one line: the syllabic n has no vowel to extend, so it must never be allowed to set lastVowel to anything but null.
// the corrected line inside tokenizeWord()'s own normal-match branch lastVowel = (chunk === 'n') ? null : chunk[chunk.length - 1];
> convert('annai') { hiragana: 'あんない', katakana: 'アンナイ' } > convert('konnichiwa') { hiragana: 'こんにちわ', katakana: 'コンニチワ' } // re-verified clean, still correct > convert('matcha') { hiragana: 'まっちゃ', katakana: 'マッチャ' } // re-verified clean, still correct

Particle Exceptions: は, へ & を

は (usually the syllable "ha"), へ (usually "he"), and を (always romanized "o," never appearing anywhere else in ordinary vocabulary) each have a real second job as a grammatical particle, pronounced "wa," "e," and "o" respectively when used that way. A plain lookup table sees only the romaji spelling, with no way to know which job a given "wa"/"e"/"o" is doing — the practical heuristic used here is that a standalone word consisting of exactly "wa," "e," or "o" is far more likely to be functioning as a particle than as the plain syllable:

const PARTICLE_OVERRIDES: Record<string, KanaPair> = { wa: ['は', 'ハ'], e: ['へ', 'ヘ'], o: ['を', 'ヲ'], }; // checked per standalone word, before tokenizeWord() runs on that word
> convert('watashi wa gakusei desu') { hiragana: 'わたし は がくせい です', katakana: 'ワタシ ハ ガクセー デス' } > convert('kore o kudasai') { hiragana: 'これ を ください', katakana: 'コレ ヲ クダサイ' } > convert('gakkou e ikimasu') { hiragana: 'がっこう へ いきます', katakana: 'ガッコー ヘ イキマス' }

Resolving Chapter 2's Own Cliffhanger: A Small Lexicalized Dictionary

The standalone-word heuristic above can't help with こんにちは, since "wa" there isn't a standalone word — it's the tail end of one single continuous typed word, "konnichiwa." A small, explicit exception dictionary, checked before ordinary tokenization, closes that specific gap:

const LEXICAL_EXCEPTIONS: Record<string, KanaPair> = { konnichiwa: ['こんにちは', 'コンニチハ'], konbanwa: ['こんばんは', 'コンバンハ'], };
> convert('konnichiwa') { hiragana: 'こんにちは', katakana: 'コンニチハ' } // Chapter 2's own cliffhanger, resolved > convert('konbanwa') { hiragana: 'こんばんは', katakana: 'コンバンハ' }
A Real, Demonstrated, Unresolvable Ambiguity
Standalone "e" is genuinely ambiguous between the へ direction particle and 絵 (e, "picture") — a real Japanese noun that happens to be spelled with the exact same single romaji letter. Two real, contrasting sentences make this concrete:
> convert('gakkou e ikimasu') // "I'm going to school" — e IS the particle { hiragana: 'がっこう へ いきます', katakana: 'ガッコー ヘ イキマス' } // correct > convert('kore wa e desu') // "This is a picture" — e means 絵, NOT the particle { hiragana: 'これ は へ です', katakana: 'コレ ハ ヘ デス' } // wrong — へ instead of え
Both sentences produce へ for the standalone "e," because the heuristic can only see that "e" is a standalone word — it has no real understanding of what the sentence actually means. This is a genuine limit of a lookup-table approach, not a bug to fix; no rule based on spelling alone can recover which sense was intended.

Hands-On Exercises

Exercise 1

Run convert('kuuki') ("air," くうき) and convert('oneesan') ("older sister," おねえさん) with real code. Predict each output first, then verify, and explain any surprise.

📄 View solution
Exercise 2

Verify with real code that convert('kekkon') ("marriage," けっこん — a genuine doubled-k sokuon) and convert('sannin') ("three people," さんにん — a genuine doubled "n" that must resolve cleanly through the syllabic-n fix, not the sokuon check) both produce the correct result, and explain why each one exercises a genuinely different part of this chapter's own logic.

📄 View solution
Exercise 3

Explain, in your own words, why the check-order fix for the chōon bug (checking for a continuation before attempting a normal token match) is necessary specifically because every vowel is also a valid standalone KANA_MAP key — and why the analogous problem never came up for the sokuon check in this same chapter.

📄 View solution

Chapter 3 Quick Reference

  • Chōon — three real bugs found: (1) naive output-character collapse, corrupting coincidental same-vowel neighbors; (2) checking a normal match before the chōon continuation, which is silently unreachable since every vowel is itself a valid token; (3) a real "undefined" leak when a genuine doubled "n" (e.g. "annai") was wrongly treated as a repeated vowel
  • Fix — check the chōon continuation first against the real romaji input, and never let the syllabic n set lastVowel to anything but null
  • Sokuon — a doubled consonant (excluding "n"), or the real Hepburn "t"+"ch" irregularity before the ち-row, verified via "matcha"
  • Particle overrides — standalone "wa"/"e"/"o" map to は/へ/を instead of わ/え/お
  • Lexicalized exceptions — こんにちは/こんばんは resolved directly, closing Chapter 2's own cliffhanger
  • Honest limit — standalone "e" as へ vs. 絵 is genuinely unresolvable from spelling alone