Edem Ahadzi, Ruchi Pandey, Tomi H. Kinnunen · arXiv (Cornell University) 2026 · 2026
DOI: 10.48550/arxiv.2609.26937
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Temporal envelopes carry cues critical to speech intelligibility, yet ASR frontends based on log-mel spectrograms do not explicitly model continuous sub-band envelope structure. This limitation is particularly acute for children's speech, where high acoustic variability demands robust feature representations. We propose a modular time-domain frontend that decomposes speech into sub-band envelopes using mel-spaced windowed-sinc filters and the Hilbert transform, with learnable per-channel energy normalization (PCEN) jointly optimized with the Whisper model. On the MyST children's speech corpus, systematic ablations identify full-band windowed-sinc filters, Hilbert envelopes, a 25 Hz smoothing cutoff, and learnable PCEN as the best configuration. Under the same Whisper-small fine-tuning setup, the frontend reduces WER from 13.16% to 11.08%, a 15.8% relative reduction over the log-mel baseline, and outperforms the evaluated Kid-Whisper checkpoint on the same cleaned test split. These results show that temporal-envelope representations and learnable frontend normalization are effective complements to backend adaptation for children's ASR.
No comments yet — start the discussion below.