What the preprint reports
An arXiv paper (authors not named in the supplied abstract) introduces UGTPhon, described as the first grapheme-to-phoneme (G2P) benchmark for user-generated text in English, Vietnamese, and Korean, plus an inference-grounded taxonomy for fine-grained diagnosis. The authors say existing G2P models and frontier LLMs show a systematic canonical-to-non-canonical performance gap reaching up to 66.8 PER points. As a baseline, they propose a compositional G2P approach that uses canonical-form evidence through exact-match lookup and staged decoding. Across matched ByT5 and Qwen2.5-0.5B backbones, the paper reports explicit canonical-form modeling consistently reduces non-canonical G2P errors, and that the 0.5B variant is competitive with much larger few-shot frontier LLMs. All figures and comparisons here are the authors' own claims; the abstract provides no independent replication. [1]
Sources
- Phonemizing User-Generated Text: A Benchmark, Taxonomy, and Compositional Approach
ArXiv CS.CL (Computation and Language) · Reporting ·