Romaji to Kana Converter: Astro — Chapter 3, Exercise 2 ===================================================================== TASK Build collapseChoonpuNaive (the character-comparison version) and confirm it works for "raamen" but produces "イーマス" for "ikimasu". Then build the fixed tokenizeScript version and confirm it correctly produces "イキマス" for the same word while still correctly producing "ラーメン" for "raamen". SOLUTION Tokenizing "raamen" into katakana first (before any collapsing) gives the raw string ラアメン. Running collapseChoonpuNaive against that raw string: console.log(collapseChoonpuNaive(tokenizeSimple('raamen', 'katakana'))); // 'ラーメン' -- correct Tokenizing "ikimasu" into katakana first gives the raw string イキマス -- already fully correct on its own, with no long vowel anywhere in it. Running the exact same collapse function against it anyway: console.log(collapseChoonpuNaive(tokenizeSimple('ikimasu', 'katakana'))); // 'イーマス' -- WRONG, should stay 'イキマス' The bug fires here because イ (from the token "i") and キ (from the token "ki") both happen to end in an "i" sound as far as the character-level vowel lookup is concerned, even though they're two completely unrelated syllables from two completely different tokens. The naive function has no idea "i" and "ki" came from separate, independent matches; it just sees two characters in a row that both map back to the vowel "i" and collapses the second one into ー. Now running the fixed tokenizeScript version, which tracks the actual matched token's own vowel inside the same loop that does the tokenizing, rather than comparing already-emitted characters afterward: console.log(tokenizeScript('ikimasu', true)); // 'イキマス' -- correct console.log(tokenizeScript('raamen', true)); // 'ラーメン' -- still correct Both fixed. "raamen" still collapses correctly because the tokenizer genuinely does match two separate "a" tokens back to back at that position -- a real, repeated vowel this time, not a coincidence of two different syllables happening to share a trailing sound. "ikimasu" no longer breaks, because the fix checks whether the CURRENT matched token is itself a bare vowel equal to the PREVIOUS matched token -- "ki" is never a bare vowel at all, so it can never be mistaken for a repeated "i" regardless of what character it happens to display. WHY THIS WORKS AS AN ANSWER ---------------------------- It reproduces both the real failure and the real fix against actual function output, and correctly identifies the structural difference between the two versions -- comparing emitted characters (naive, and therefore blind to which token produced them) versus comparing matched tokens directly (fixed, and therefore only ever collapsing a genuine repeated vowel).