Romaji to Kana Converter: Angular & Express — Chapter 3, Exercise 1 ===================================================================== TASK Build sokuonLength exactly as this chapter describes, wire it into a tokenizer, and verify it against three real words: "zutto" (ずっと, all along), "issho" (いっしょ, together), and "matcha" (まっちゃ, matcha). Then explain, in your own words, why checking for "ch" specifically after a "t" has to happen before the general doubled-consonant check runs -- what would go wrong if the order were reversed? SOLUTION const SMALL_TSU = { hiragana: 'っ', katakana: 'ッ' }; function sokuonLength(romaji: string, i: number): number { const c = romaji[i]; if (!c || 'aiueo'.includes(c)) return 0; // Real Hepburn irregularity: gemination before chi/cha/chu/cho is // spelled with "t", not a doubled "c" -- matcha, never maccha. if (c === 't' && romaji.slice(i + 1, i + 3) === 'ch') return 1; // Doubled n is NOT sokuon. if (c !== 'n' && romaji[i + 1] === c) return 1; return 0; } function tokenizeWithSokuon(romaji: string): { hiragana: string; katakana: string } { let hiragana = ''; let katakana = ''; let i = 0; while (i < romaji.length) { const consumed = sokuonLength(romaji, i); if (consumed > 0) { hiragana += SMALL_TSU.hiragana; katakana += SMALL_TSU.katakana; i += consumed; continue; } let matched = false; for (let len = MAX_TOKEN_LENGTH; len >= 1; len--) { const chunk = romaji.slice(i, i + len).toLowerCase(); const entry = KANA_MAP[chunk]; if (entry) { hiragana += entry.hiragana; katakana += entry.katakana; i += len; matched = true; break; } } if (!matched) { hiragana += romaji[i]; katakana += romaji[i]; i += 1; } } return { hiragana, katakana }; } Testing: console.log(tokenizeWithSokuon('zutto')); console.log(tokenizeWithSokuon('issho')); console.log(tokenizeWithSokuon('matcha')); Run with: node (or ts-node) Output (verified by hand-running the real function): { hiragana: 'ずっと', katakana: 'ズット' } { hiragana: 'いっしょ', katakana: 'イッショ' } { hiragana: 'まっちゃ', katakana: 'マッチャ' } zutto: at position 2, "t" is followed by another "t" (not "n," and not the "t"+"ch" case), so sokuonLength returns 1 -- っ is emitted, the first "t" is consumed, and the remaining "to" is tokenized normally as と. issho: at position 1, "s" is followed by another "s" -- っ is emitted, one "s" is consumed, and the remaining "sho" matches the real 3-character yōon token しょ directly. matcha: at position 2, "t" is checked against the "ch"-lookahead case FIRST and matches (the next two characters really are "c" then "h"), consuming just the "t" and leaving "cha" for the ordinary tokenizer to match as ちゃ. WHY THE "tch" CHECK MUST COME BEFORE THE GENERAL DOUBLED-CONSONANT CHECK -------------------------------------------------------------------------- The general doubled-consonant rule only fires when the CURRENT character is literally repeated as the NEXT character (romaji[i+1] === c). In "matcha," the character at the "t" position is "t," and the very next character is "c" -- not "t" again -- so the general rule would never match this case at all, doubled or not. If the "tch" check were removed or placed after a check that assumed "the general rule already covers every geminated consonant," the "t" in "matcha" would simply fail both checks and be treated as an ordinary, un-doubled letter: sokuonLength would return 0, no っ would be emitted, and the tokenizer would instead try to match "t" alone (not a valid one-character token) and fall through to emitting a raw, unconverted "t" character before separately tokenizing "cha" as ちゃ -- producing the broken "tちゃ" instead of the correct まっちゃ. Order doesn't actually matter here in the sense of one check "stealing" cases from the other (the two conditions can never both be true for the same position, since one requires the next character to equal "t" and the other requires it to equal "c"), but the "tch" case would be silently missed entirely if it weren't checked at all -- it's a genuinely separate rule, not a variant of the general one. WHY THIS WORKS AS AN ANSWER ---------------------------- It reproduces sokuonLength exactly as specified, verifies three real words covering three different real triggers (a plain doubled consonant, a doubled consonant before a yōon combination, and the irregular "tch" spelling), and correctly explains that the "tch" case is a structurally separate rule from the general doubled-consonant rule -- not a matter of which check "wins" when both could apply, since they never both apply to the same input at once.