Handling Real Edge Cases
Romaji to Kana Converter: React & Next.js
Chapter 3 · Handling Real Edge Cases
Chapter 2 ended on a real cliffhanger: convert('konnichiwa') produces こんにちわ,
the common casual spelling, not the traditionally "correct" こんにちは — because a plain
lookup table has no way to know that a trailing は should sometimes be pronounced "wa" rather
than "ha." That's one of three genuine edge cases this chapter closes out: long vowels, sokuon
(small っ), and は/へ/を particle exceptions. Every claim below was checked by actually running
the code, including two real chōon bugs found along the way — one that reproduces a failure
mode both sibling courses already documented, and a second, subtler one that's specific to how
this variant's own approach is structured.
Long Vowels: Two Real Ways to Get This Wrong First
In doubled-vowel romaji notation, a repeated vowel letter (or the alternate spellings "ei" and "ou") signals a long vowel — katakana always marks this with a single chōonpu (ー); hiragana conventionally extends with a second vowel kana (あ for a, い for i, う for u, and by default い for e and う for o, with real, genuine exception words that don't follow the default, covered in the warn-box below).
Attempt 1: Collapse the Output, After the Fact
A tempting first approach runs after tokenization is already finished — walk the finished kana string, and whenever two adjacent characters "sound like" the same vowel, collapse the second one into a chōonpu:
Run against real, already-correct tokenizer output, this immediately corrupts a word that never contained a long vowel in the first place:
Attempt 2: Check the Input — But in the Wrong Order
Moving the check into the tokenizer itself, working directly against the romaji input rather than the finished kana, looks like the right fix. A first version checks for a normal token match first, falling back to a chōon check only when nothing else matches:
Run against a real word with a genuine long vowel, this version detects nothing at all:
a, i, u, e,
o — is itself a valid, standalone 1-character key in KANA_MAP. Because
the normal token-match loop runs first, that second, repeated vowel letter always gets matched
as a brand-new ordinary kana character before the chōon-check branch is ever reached — the
fallback code is real, correct, and completely unreachable. The one fix that actually works is
checking for a chōon continuation before attempting a fresh token match, not after:
Re-verified against every case that broke either earlier attempt:
とおい (far), おおきい (big), おおい (many) — really are
spelled with おお instead. No lookup rule can tell "koohii" (a default-rule word) apart from a
real お-doubling exception from the romaji spelling alone; this converter's default will get
those specific words wrong, exactly as both sibling courses' own converters do.
Sokuon: A Doubled Consonant Signals っ/ッ
A doubled consonant (not a vowel, not "n") signals a small っ before the following syllable — only the first of the pair is consumed as the sokuon signal; the second starts the real next token normally. Japanese has one genuine, real irregularity in how this is spelled before the ち-row: a sokuon before ち/ちゃ/ちゅ/ちょ is written as a single "t" followed by "ch," not by doubling the "ch" digraph itself:
Verified against a real word for each case, including the "nn" exclusion that must not trigger a sokuon:
tokenizeWord function — the
priority order established above (chōon check, then sokuon check, then normal match) holds up
without any special-casing needed for words that use more than one edge case at once.
A Third Bug: The Syllabic ん Has No Real Vowel to Extend
Every normal token sets lastVowel to that token's own trailing letter, so the
chōon check on the next loop iteration knows what vowel to look for a repeat of. The syllabic
n token (n → ん) was, without thinking it through, treated the same way — its
"trailing letter" is n itself, so lastVowel ends up set to the
literal character 'n'. That's fine right up until a second n
genuinely follows the first — a real, common pattern, not an edge case:
lastVowel is literally the string
'n'. The very next character is a second, genuine "n" — and the chōon check's own
c === lastVowel comparison treats 'n' === 'n' exactly like a doubled
vowel, firing the chōon branch. HIRA_EXT has no 'n' key at all, so
HIRA_EXT['n'] evaluates to undefined — and string concatenation with
undefined in JavaScript silently produces the literal four-letter word
"undefined" spliced directly into the output, a genuinely worse failure than a merely wrong
kana character. The real fix is one line: the syllabic n has no vowel to extend, so it must
never be allowed to set lastVowel to anything but null.
Particle Exceptions: は, へ & を
は (usually the syllable "ha"), へ (usually "he"), and を (always romanized "o," never appearing anywhere else in ordinary vocabulary) each have a real second job as a grammatical particle, pronounced "wa," "e," and "o" respectively when used that way. A plain lookup table sees only the romaji spelling, with no way to know which job a given "wa"/"e"/"o" is doing — the practical heuristic used here is that a standalone word consisting of exactly "wa," "e," or "o" is far more likely to be functioning as a particle than as the plain syllable:
Resolving Chapter 2's Own Cliffhanger: A Small Lexicalized Dictionary
The standalone-word heuristic above can't help with こんにちは, since "wa" there isn't a standalone word — it's the tail end of one single continuous typed word, "konnichiwa." A small, explicit exception dictionary, checked before ordinary tokenization, closes that specific gap:
Hands-On Exercises
Run convert('kuuki') ("air," くうき) and convert('oneesan') ("older sister," おねえさん) with real code. Predict each output first, then verify, and explain any surprise.
Verify with real code that convert('kekkon') ("marriage," けっこん — a genuine doubled-k sokuon) and convert('sannin') ("three people," さんにん — a genuine doubled "n" that must resolve cleanly through the syllabic-n fix, not the sokuon check) both produce the correct result, and explain why each one exercises a genuinely different part of this chapter's own logic.
Explain, in your own words, why the check-order fix for the chōon bug (checking for a continuation before attempting a normal token match) is necessary specifically because every vowel is also a valid standalone KANA_MAP key — and why the analogous problem never came up for the sokuon check in this same chapter.
Chapter 3 Quick Reference
- Chōon — three real bugs found: (1) naive output-character collapse, corrupting coincidental same-vowel neighbors; (2) checking a normal match before the chōon continuation, which is silently unreachable since every vowel is itself a valid token; (3) a real "undefined" leak when a genuine doubled "n" (e.g. "annai") was wrongly treated as a repeated vowel
- Fix — check the chōon continuation first against the real romaji input, and never let the syllabic n set
lastVowelto anything butnull - Sokuon — a doubled consonant (excluding "n"), or the real Hepburn "t"+"ch" irregularity before the ち-row, verified via "matcha"
- Particle overrides — standalone "wa"/"e"/"o" map to は/へ/を instead of わ/え/お
- Lexicalized exceptions — こんにちは/こんばんは resolved directly, closing Chapter 2's own cliffhanger
- Honest limit — standalone "e" as へ vs. 絵 is genuinely unresolvable from spelling alone