Handling Real Edge Cases
Romaji to Kana Converter: Angular & Express
Chapter 3 · Handling Real Edge Cases: Long Vowels, Sokuon (っ) & Particle Exceptions
Chapter 2's own closing warn-box named three things deliberately left unhandled: long vowels, the small っ
(sokuon), and — surfaced by that chapter's very last test — the real は/へ/を particle exceptions. This
chapter builds real, tested handling for all three. Every example below was run through an actual
node script before being written down here; two of them exposed genuine bugs in the first
attempt, both caught and fixed the same way Chapter 2's own tokenizer bugs were — by actually running the
code rather than assuming it was correct.
Long Vowels: How Chōon Is Actually Spelled
A long (double-length) vowel sound is real and phonemically distinct in Japanese — おばさん
(obasan, aunt) and おばあさん (obaasan, grandmother) are genuinely different words, not the
same word said carefully. But hiragana and katakana don't spell a long vowel the same way, and hiragana
itself doesn't even spell it the same way for every vowel:
| Vowel Row | Real Hiragana Convention | Real Exceptions |
|---|---|---|
| あ (a) | Extend with あ — おかあさん (okaasan, mother) | — |
| い (i) | Extend with い — おにいさん (oniisan, older brother) | — |
| う (u) | Extend with う — くうき (kuuki, air) | — |
| え (e) | Usually extend with い — せんせい (sensei, teacher) | Sometimes え — おねえさん (oneesan, older sister) |
| お (o) | Usually extend with う — とうきょう (toukyou) | Sometimes お — とおい (tooi, far), おおきい (ookii, big) |
Katakana avoids this split entirely. A long vowel in katakana is always written with one mark, ー (chōonpu), regardless of which vowel is being extended — コーヒー (kōhī, coffee), ラーメン (rāmen, ramen). This isn't a simplification this course is inventing; it's the real, standard convention for how katakana (used overwhelmingly for loanwords) spells vowel length.
A First Attempt at Long Vowels — And a Real Bug
The plan: expand each macron character into an ASCII digraph before tokenizing, so the existing longest-match loop from Chapter 2 can do the rest. Hiragana gets the "best guess" native-word extension (え→い, お→う); katakana gets the same vowel doubled, since that's what will need collapsing into a single ー:
Then, a first pass at turning a doubled katakana vowel into ー — scan the finished string, and whenever two consecutive characters share the same trailing vowel sound, collapse the second into ー:
collapseChoonpuNaive correctly turns ラアメン (from tokenizing "raamen")
into ラーメン. But run the exact same function against the katakana for the ordinary word
"ikimasu" (行きます, I go) — a word with no long vowel in it at all — and it produces
イーマス instead of the correct イキマス.
i) and キ (from ki) are two entirely different, unrelated syllables that
simply happen to both end in an "i" sound — the same coincidence that would make "kit" and "sit" rhyme in
English without either word containing a doubled letter. collapseChoonpuNaive only looks at the
two most recently emitted characters, with no memory of which whole token either one came
from — so two unrelated syllables that happen to share a trailing vowel get merged exactly as if a genuine
vowel had been typed twice in a row. A real vowel-lengthening collapse has to compare against the vowel of
the last matched token, not the last accumulated character.
The Real Fix: Track the Vowel Inside the Tokenizer Itself
Rather than patching the finished string afterward, the chōonpu decision moves inside the same loop that's already matching tokens — it always knows exactly which real syllable it just emitted, so it can never confuse "ends in the same vowel sound" with "is a genuine repeated vowel":
len === 1 instead of chunk.length === 1
to decide whether a match was a bare vowel. Near the end of a string, romaji.slice(i, i + 3)
can genuinely return fewer than 3 characters — there's nothing left to slice — so a length-3 attempt at the
very last character of "koohii" returns the 1-character string "i", which
correctly matches the vowel token, but with len still equal to 3. Checking the loop variable
instead of the actual matched string's own length made the function think this was a 3-character match,
never recognizing it as the bare vowel it actually was — so the real second long vowel in "kōhī"
silently failed to collapse, producing コーヒイ instead of the correct コーヒー.
Switching the check to chunk.length — the true length of what was actually found, regardless of
which loop iteration found it — fixed it.
Both scripts now need their own full pass, expanded with their own macron table, because they diverge on exactly the case this chapter is about:
Sokuon: The Small っ
A doubled consonant in romaji (other than a doubled n) signals a small っ/ッ immediately before
the syllable that follows — きって (kitte, stamp), がっこう (gakkou, school). The small tsu itself carries no
sound of its own; it's a brief, silent held pause, then the following consonant releases normally:
Wired into tokenizeScript, this is checked before every ordinary token attempt: if a sokuon
trigger is found, emit the small tsu, advance past just the trigger, and let the loop pick up the real
syllable from there on its very next iteration.
t, the next two characters are exactly "ch", so the special
case fires and consumes only the t — emitting っ and leaving the string positioned right at
"cha", which the ordinary tokenizer then matches as a normal 3-character token, ちゃ. The
doubled-consonant rule alone (t followed by another t) would never have fired
here at all, since the second letter is c, not t — this is a genuinely separate,
irregular case, not a variation on the ordinary rule.
The doubled-n exclusion matters for a real, common word — こんにちは (konnichiwa, hello) is
ko-n-ni-chi-wa, five real syllables with no small tsu anywhere in it. Treating a doubled n as
sokuon would insert a phantom っ that doesn't belong:
The Real は/へ/を Particle Exceptions
Three hiragana keep an old, "wrong-looking" spelling in one specific grammatical role — the character is written one way but pronounced another:
| Character | Ordinary Reading | As a Particle | Real Example |
|---|---|---|---|
| は | ha | wa (topic marker) | これは本です — kore wa hon desu |
| へ | he | e (direction marker) | これへ行きます — kore e ikimasu |
| を | wo | o (object marker) | これをください — kore o kudasai |
Our converter runs romaji into kana, so the direction that actually matters is: given typed
"wa", "e", or "o", when should the output be は/へ/を instead of the
ordinary わ/え/お? A workable, honest heuristic: only when that romaji appears as its own complete,
space-separated word — never when it's a letter sequence inside a longer word:
This resolves the two grammatically unambiguous cases correctly, and correctly leaves わたし (watashi, I) —
which genuinely contains the letters "wa" but never as its own word — completely untouched:
"wa" and "o" standing alone are almost always the particle — there's little else a
lone "wa" or "o" would normally be. But standalone "e" is a real,
common Japanese word in its own right — 絵 (e), meaning "picture." Compare two genuinely grammatical
sentences:
"kore e ikimasu"("I'm going this way") —"e"really is the へ particle"kono e wa kirei desu"("This picture is pretty") —"e"here means 絵, and should stay え, not へ
"e" in
the identical grammatical position. This converter always resolves standalone "e" to へ, which
is right for the first sentence and genuinely wrong for the second. Correctly distinguishing them would
need real part-of-speech context — something well outside a lookup-table-driven converter — so this is
documented as a known, permanent limitation rather than quietly hidden.
Verified against both real sentences directly — the second one confirms the honest limitation exists, not just that it's theoretically possible:
こんにちは: A Whole-Word Exception, Not a Grammatical One
こんにちは itself needs a genuinely different fix from the general particle rule above. Its own は isn't a standalone particle word at all — it's fused inside one continuous greeting, historically built from "today, as for..." with the rest of the sentence traditionally left unsaid. The word-boundary heuristic can't reach it, because there's no space anywhere near that は to detect. This needs a small, explicit dictionary of whole-word exceptions, checked before anything else runs:
PARTICLE_OVERRIDES rule, which applies whenever a real grammatical pattern holds (a standalone
word between two others), and LEXICALIZED_EXCEPTIONS, a small closed list of specific,
memorized whole words whose irregular spelling can't be derived from any general rule at all. こんばんは
(konbanwa, good evening) is a real, second example of the exact same lexicalized pattern — worth knowing the
category exists, since a third such word could show up later and need the same treatment.
Putting It All Together
The final convert() checks the lexicalized exception dictionary first, applies the particle
override to a standalone word second, then runs the full macron-expansion-plus-sokuon-aware tokenizer for
each script — every layer from this chapter, combined:
Every feature from this chapter, exercised together in one sentence — a long vowel from a native word, a doubled consonant, and a direction particle, all at once:
Hands-On Exercises
Build sokuonLength exactly as this chapter describes, wire it into a tokenizer, and verify it against three real words: "zutto" (ずっと, all along), "issho" (いっしょ, together), and "matcha" (まっちゃ, matcha). Then explain, in your own words, why checking for "ch" specifically after a "t" has to happen before the general doubled-consonant check runs — what would go wrong if the order were reversed?
📄 View solutionBuild collapseChoonpuNaive (the character-comparison version) and confirm it works for "raamen" but produces "イーマス" for "ikimasu". Then build the fixed tokenizeScript version and confirm it correctly produces "イキマス" for the same word while still correctly producing "ラーメン" for "raamen".
📄 View solutionBuild convertSentence with the particle overrides, then run it against both "kore e ikimasu" and "kono e wa kirei desu". Confirm the first is correct and the second is wrong (produces へ where え would be correct), and write one sentence explaining why no purely lookup-table-based fix could resolve this specific ambiguity.
📄 View solutionChapter 3 Quick Reference
- Long vowels — hiragana extends あ/い/う-row vowels with themselves, but usually extends え-row with い and お-row with う (real exceptions exist: とおい, おおきい); katakana always uses one mark, ー
- Bug 1 — collapsing a long vowel by comparing accumulated katakana characters (not matched tokens) wrongly merges unrelated syllables that happen to share a trailing vowel (イキ → イー)
- Bug 2 — checking the loop variable
leninstead ofchunk.lengthmisses a bare-vowel match truncated by the string's own end - Sokuon — a doubled consonant (except doubled
n) inserts っ/ッ before the following syllable; gemination before ち/ちゃ/ちゅ/ちょ is spelled with "t" ("matcha," never "maccha") - Particle exceptions — standalone "wa"/"e"/"o" override to は/へ/を; standalone "e" is genuinely ambiguous with 絵 (picture) and has no lookup-table fix
- Lexicalized exceptions — こんにちは and こんばんは need a small whole-word dictionary, since their own は isn't a standalone particle at all
- Next chapter: Angular Services & Components — structuring the conversion tool around this now-complete engine