Romaji to Kana Converter: Angular & Express — Chapter 3, Exercise 2 ===================================================================== TASK Build collapseChoonpuNaive (the character-comparison version) and confirm it works for "raamen" but produces "イーマス" for "ikimasu". Then build the fixed tokenizeScript version and confirm it correctly produces "イキマス" for the same word while still correctly producing "ラーメン" for "raamen". SOLUTION Naive, broken version (compares individual accumulated characters): function collapseChoonpuNaive(katakana: string, vowelOf: Record): string { let out = ''; let prevVowel: string | null = null; for (const ch of katakana) { const vowel = vowelOf[ch]; if (vowel && vowel === prevVowel) { out += 'ー'; } else { out += ch; } prevVowel = vowel ?? null; } return out; } // vowelOf keyed by the trailing katakana character of every KANA_MAP entry const vowelOf: Record = {}; for (const [key, entry] of Object.entries(KANA_MAP)) { if (key === 'n') continue; vowelOf[entry.katakana.slice(-1)] = key[key.length - 1]; } Testing the naive version: console.log(collapseChoonpuNaive(rawTokenize('raamen').katakana, vowelOf)); console.log(collapseChoonpuNaive(rawTokenize('ikimasu').katakana, vowelOf)); Run with: node (or ts-node) Output (verified by hand-running the real function): ラーメン イーマス "raamen" correctly collapses (ラ then ア genuinely repeats the same "a" vowel, a real lengthening). "ikimasu" incorrectly collapses -- イ ("i") and キ (from "ki," also ending in "i") are two unrelated syllables that just happen to share a trailing vowel sound, and the naive function has no way to tell that apart from a genuine repeated vowel, since it only ever looks at the two most recently emitted CHARACTERS with no memory of which whole token either one came from. Fixed version (tracks the vowel of the last MATCHED TOKEN, inside the tokenizer loop itself, exactly as this chapter builds it): function tokenizeScript(romaji: string, useChoonpu: boolean): string { let out = ''; let i = 0; let lastVowel: string | null = null; while (i < romaji.length) { let matched = false; for (let len = MAX_TOKEN_LENGTH; len >= 1; len--) { const chunk = romaji.slice(i, i + len).toLowerCase(); const entry = KANA_MAP[chunk]; if (entry) { const isBareVowel = chunk.length === 1 && 'aiueo'.includes(chunk); if (useChoonpu && isBareVowel && chunk === lastVowel) { out += 'ー'; } else { out += useChoonpu ? entry.katakana : entry.hiragana; } lastVowel = chunk === 'n' ? null : chunk[chunk.length - 1]; i += len; matched = true; break; } } if (!matched) { out += romaji[i]; lastVowel = null; i += 1; } } return out; } Testing the fixed version: console.log(tokenizeScript('raamen', true)); console.log(tokenizeScript('ikimasu', true)); Output (verified by hand-running the real function): ラーメン イキマス Both are now correct. The fix works because "lastVowel" only ever gets set to the vowel of the token that was JUST matched -- when "ki" is matched, lastVowel becomes "i" because "ki" itself ends in "i," which is real and correct, but the NEXT position (whatever comes after "ki") only collapses to ー if IT is also a genuine bare vowel match equal to that same "i" -- and in "ikimasu," the character after "ki" is "m," starting an entirely different token ("ma"), so no false collapse is ever triggered. The naive version's mistake was comparing raw output characters after the fact; the fix compares the actual romaji tokens as they're matched, which is the only place the real distinction between "a repeated vowel" and "two syllables that happen to end the same way" can genuinely be seen. WHY THIS WORKS AS AN ANSWER ---------------------------- It reproduces both the real "イーマス" bug from the naive character- level approach and confirms the token-level fix resolves it while preserving the correct "raamen" → "ラーメン" behavior -- demonstrating the actual root cause (comparing characters instead of matched tokens) rather than just patching the specific failing test case.