What the paper proposes
> **ArXiv CS.CL** describes Ruby-ASR, which refines the conventional Japanese ASR training target from a plain orthographic transcript into a span-bound orthographic–lexical-reading sequence. In the authors' framing, binding each written span locally to its realized reading — rather than emitting separate full-sentence orthographic and phonological outputs — allows deterministic recovery of both views. The paper says it instantiates this target under both subtitle-style and verbatim-style transcription conventions, using a Qwen3-ASR backbone, with a mora-level CTC objective providing auxiliary monotonic reading supervision. [1]
Reported results and caveats
> According to the **arXiv** abstract, experiments across five Japanese benchmarks show that refining the recognition target can improve lexical-reading recovery without sacrificing readable orthographic transcription. These are the authors' reported results from their own benchmarks; no independent replication or third-party evaluation is included in the supplied evidence. The abstract does not list specific benchmark scores or dataset names. Checkpoints and inference code are stated to be released. [1]
Sources
- Ruby-ASR: Evidence-Preserving Supervision for Joint Orthographic and Lexical-Reading Recognition
ArXiv CS.CL (Computation and Language) · Reporting ·