Handling Real Edge Cases: Long Vowels, Sokuon (っ) & Particle Exceptions
Romaji to Kana Converter: Astro
Chapter 3 · Handling Real Edge Cases: Long Vowels, Sokuon (っ) & Particle Exceptions
Chapter 2's own closing warn-box named two things deliberately left unhandled — long vowels and the small
っ — and its very last test surfaced a third: convert('konnichiwa') produced こんにちわ
instead of the traditionally correct こんにちは. This chapter builds real, tested handling
for all three, adding directly onto the same src/lib/convert.ts module Chapter 2 started. Every
example below was independently re-run through a real Node script before being written down — two of them
exposed genuine bugs on the first attempt, caught the exact same way Chapter 2's own tokenizer bugs were:
by actually running the code rather than trusting it looked right.
Long Vowels: How Chōon Is Actually Spelled
A long (double-length) vowel sound is real and phonemically distinct in Japanese — おばさん
(obasan, aunt) and おばあさん (obaasan, grandmother) are genuinely different words, not the
same word said carefully. But hiragana and katakana don't spell a long vowel the same way, and hiragana
itself doesn't even spell it the same way for every vowel:
| Vowel Row | Real Hiragana Convention | Real Exceptions |
|---|---|---|
| あ (a) | Extend with あ — おかあさん (okaasan, mother) | — |
| い (i) | Extend with い — おにいさん (oniisan, older brother) | — |
| う (u) | Extend with う — くうき (kuuki, air) | — |
| え (e) | Usually extend with い — せんせい (sensei, teacher) | Sometimes え — おねえさん (oneesan, older sister) |
| お (o) | Usually extend with う — とうきょう (toukyou) | Sometimes お — とおい (tooi, far), おおきい (ookii, big) |
Katakana avoids this split entirely — a long vowel in katakana is always written with a single mark, ー (chōonpu), regardless of which vowel is being extended: コーヒー (kōhī, coffee), ラーメン (rāmen, ramen). That's the real, standard convention katakana (used overwhelmingly for loanwords) actually follows, not a simplification invented for this course.
A First Attempt — And a Real Bug
The plan: expand each macron character into an ASCII digraph before tokenizing, so the longest-match loop already built in Chapter 2 can do the rest. Hiragana gets the "best guess" native-word extension (え→い, お→う); katakana gets the same vowel doubled, since that's what will need collapsing into a single ー:
Then a first, plausible pass at turning a doubled katakana vowel into ー — scan the already-tokenized string, and whenever two consecutive characters happen to share the same trailing vowel sound, collapse the second into ー:
Run for real against the katakana for "raamen", this correctly collapses ラアメン
into ラーメン. Run the identical function against the katakana for the ordinary word
"ikimasu" (行きます, I go) — a word with no long vowel in it at all — and it produces
イーマス instead of the correct イキマス, verified directly
against real Node output rather than assumed from the algorithm's own description.
i) and キ (from ki) are two entirely different, unrelated syllables that
simply happen to both end in an "i" sound — the same coincidence that makes "kit" and "sit" rhyme in
English without either word containing a doubled letter. collapseChoonpuNaive only looks at
the two most recently emitted characters, with no memory of which whole token either one
came from — so two unrelated syllables that happen to share a trailing vowel get merged exactly as if a
genuine vowel had been typed twice in a row. A real vowel-lengthening collapse has to compare against the
vowel of the last matched token, not the last accumulated character.
The Real Fix: Track the Vowel Inside the Tokenizer Itself
Rather than patching the finished string afterward, the chōonpu decision moves inside the same loop that's already matching tokens — it always knows exactly which real syllable it just emitted, so it can never confuse "ends in the same vowel sound" with "is a genuine repeated vowel":
len === 1 instead of chunk.length === 1
to decide whether a match was a bare vowel. Near the end of a string, romaji.slice(i, i + 3)
can genuinely return fewer than 3 characters — there's nothing left to slice — so a length-3 attempt at the
very last character of "koohii" returns the 1-character string "i", which
correctly matches the vowel token, but with len still equal to 3. Checking the loop variable
instead of the matched string's own real length made the function think this was a 3-character match,
never recognizing it as the bare vowel it actually was — re-run for real, that buggy version produces
コーヒイ instead of the correct コーヒー. Switching the check to
chunk.length — the true length of what was actually found, regardless of which loop iteration
found it — fixed it, re-verified producing the correct コーヒー.
Both scripts now need their own full pass, expanded with their own macron table, because they genuinely diverge on exactly the case this chapter is about:
Sokuon: The Small っ
A doubled consonant in romaji (other than a doubled n) signals a small っ/ッ immediately before
the syllable that follows — きって (kitte, stamp), がっこう (gakkou, school). The small tsu itself carries
no sound of its own; it's a brief, silent held pause, then the following consonant releases normally:
Wired into tokenizeScript, this is checked before every ordinary token attempt: if a sokuon
trigger is found, emit the small tsu, advance past just the trigger, and let the loop pick up the real
syllable from there on its very next iteration.
t, the next two characters are exactly "ch", so the special
case fires and consumes only the t — emitting っ and leaving the string positioned right at
"cha", which the ordinary tokenizer then matches as a normal 3-character token, ちゃ. The
doubled-consonant rule alone (t followed by another t) would never have fired
here at all, since the second letter is c, not t — this is a genuinely separate,
irregular case, not a variation on the ordinary rule.
The doubled-n exclusion matters for a real, common word — こんにちは (konnichiwa, hello) is
ko-n-ni-chi-wa, five real syllables with no small tsu anywhere in it. Treating a doubled n as
sokuon would insert a phantom っ that doesn't belong:
The Real は/へ/を Particle Exceptions
Three hiragana keep an old, "wrong-looking" spelling in one specific grammatical role — the character is written one way but pronounced another:
| Character | Ordinary Reading | As a Particle | Real Example |
|---|---|---|---|
| は | ha | wa (topic marker) | これは本です — kore wa hon desu |
| へ | he | e (direction marker) | これへ行きます — kore e ikimasu |
| を | wo | o (object marker) | これをください — kore o kudasai |
This converter runs romaji into kana, so the direction that actually matters is: given typed
"wa", "e", or "o", when should the output be は/へ/を instead of the
ordinary わ/え/お? A workable, honest heuristic: only when that romaji appears as its own complete,
space-separated word — never when it's a letter sequence inside a longer word:
This resolves both grammatically unambiguous cases correctly, and correctly leaves わたし (watashi, I) —
which genuinely contains the letters "wa" but never as its own word — completely untouched:
"wa" and "o" standing alone are almost always the particle — there's little else
a lone "wa" or "o" would normally be. But standalone "e" is a real,
common Japanese word in its own right — 絵 (e), meaning "picture." Compare two genuinely grammatical
sentences:
"kore e ikimasu"("I'm going this way") —"e"really is the へ particle"kono e wa kirei desu"("This picture is pretty") —"e"here means 絵, and should stay え, not へ
"e"
in the identical grammatical position. This converter always resolves standalone "e" to へ,
which is right for the first sentence and genuinely wrong for the second. Correctly distinguishing them
would need real part-of-speech context — something well outside a lookup-table-driven converter — so this
is documented as a known, permanent limitation rather than quietly hidden.
Verified against both real sentences directly — the second one confirms the honest limitation actually exists, not just that it's theoretically possible:
こんにちは: A Whole-Word Exception, Not a Grammatical One
こんにちは itself needs a genuinely different fix from the general particle rule above. Its own は isn't a standalone particle word at all — it's fused inside one continuous greeting, historically built from "today, as for..." with the rest of the sentence traditionally left unsaid. The word-boundary heuristic can't reach it, because there's no space anywhere near that は to detect. This needs a small, explicit dictionary of whole-word exceptions, checked before anything else runs:
PARTICLE_OVERRIDES rule, which applies whenever a real grammatical pattern holds (a standalone
word between two others), and LEXICALIZED_EXCEPTIONS, a small closed list of specific,
memorized whole words whose irregular spelling can't be derived from any general rule at all. こんばんは
(konbanwa, good evening) is a real, second example of the exact same lexicalized pattern — worth knowing
the category exists, since a third such word could show up later and need the same treatment.
Putting It All Together
The final convert() checks the lexicalized exception dictionary first, applies the particle
override to a standalone word second, then runs the full macron-expansion-plus-sokuon-aware tokenizer for
each script — every layer from this chapter, combined:
Every feature from this chapter, exercised together in one sentence — a doubled consonant and a direction particle at once — plus the exact greeting Chapter 2 ended on:
Nothing built across this chapter or the last one depends on Astro yet — src/lib/convert.ts
is still plain TypeScript. Chapter 4 imports this exact, unmodified module into a real input/output UI;
Chapter 5 is where it finally gets a server-side alternative to be measured against.
Hands-On Exercises
Build sokuonLength exactly as this chapter describes, wire it into a tokenizer, and verify it against three real words: "zutto" (ずっと, all along), "issho" (いっしょ, together), and "matcha" (まっちゃ, matcha). Then explain, in your own words, why checking for "ch" specifically after a "t" has to happen before the general doubled-consonant check runs — what would go wrong if the order were reversed?
📄 View solutionBuild collapseChoonpuNaive (the character-comparison version) and confirm it works for "raamen" but produces "イーマス" for "ikimasu". Then build the fixed tokenizeScript version and confirm it correctly produces "イキマス" for the same word while still correctly producing "ラーメン" for "raamen".
📄 View solutionBuild convertSentence with the particle overrides, then run it against both "kore e ikimasu" and "kono e wa kirei desu". Confirm the first is correct and the second is wrong (produces へ where え would be correct), and write one sentence explaining why no purely lookup-table-based fix could resolve this specific ambiguity.
📄 View solutionChapter 3 Quick Reference
- Long vowels — hiragana extends あ/い/う-row vowels with themselves, but usually extends え-row with い and お-row with う (real exceptions exist: とおい, おおきい); katakana always uses one mark, ー
- Bug 1 — collapsing a long vowel by comparing accumulated katakana characters (not matched tokens) wrongly merges unrelated syllables that happen to share a trailing vowel (イキ → イー)
- Bug 2 — checking the loop variable
leninstead ofchunk.lengthmisses a bare-vowel match truncated by the string's own end - Sokuon — a doubled consonant (except doubled
n) inserts っ/ッ before the following syllable; gemination before ち/ちゃ/ちゅ/ちょ is spelled with "t" ("matcha," never "maccha") - Particle exceptions — standalone "wa"/"e"/"o" override to は/へ/を; standalone "e" is genuinely ambiguous with 絵 (picture) and has no lookup-table fix
- Lexicalized exceptions — こんにちは and こんばんは need a small whole-word dictionary, since their own は isn't a standalone particle at all
- Still framework-agnostic — src/lib/convert.ts is complete and correct-as-documented, with no Astro dependency anywhere yet
- Next chapter: Building the Input/Output UI in Astro — this module gets its first real caller