Ki Woong Moon, Daniel Brenner · arXiv (Cornell University) 2026 · 2026
DOI: 10.48550/arxiv.2609.36754
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Explicit prosodic cues may help automatic speech recognition (ASR) of spontaneous speech, but auxiliary representations typically require additional trainable components, making it unclear whether gains come from the auxiliary information or the fusion mechanism. We address this using a frozen HuBERT backbone and a 64-dimensional representation trained to predict log F0, voicing, Delta log F0, log energy, and spectral tilt. We compare a frozen-backbone recognizer (Baseline), trainable fusion with zero auxiliary input (Null), and the same fusion supplied with the learned representation (Learned). Across Buckeye, Switchboard, and AMI IHM, Null reduces WER by 0.71-1.45 points over Baseline, whereas Learned differs from Null by +0.07, -0.09, and +0.00 points, with no significant differences. However, removing or mismatching the representation at inference increases Learned WER. Thus, Learned depends on the representation yet shows no measurable incremental WER benefit over the parameter-matched control.
No comments yet — start the discussion below.