Steven Coats · Publication Server of the Institute for German Language (Institute for German Language) 2026 · 2026
DOI: 10.14618/ids-pub-14097
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Automatic speech recognition (ASR) is increasingly used for large-scale corpus construction, yet the impact of decoding parameters on resulting datasets remains underexplored. We compare two operational decoding configurations selected by fixed temperature settings (0.0 vs. 0.1) in a Singapore-English-adapted Whisper pipeline. In a paired rerun of 1,322 podcast recordings (774 hours), sampling at temperature 0.1 produced 99,011 more aligned words than beam search at temperature 0.0 (+1.02%; 95% paired-bootstrap interval: +0.89–1.16%). Hand-checked passages show that these differences include both recovered audible material and acoustically unsupported insertions. An additional sampling-only control at temperature 0.01 produced more output than at 0.1, showing that output quantity was not monotonic with temperature. We argue that decoding configuration should be treated as a methodological decision in corpus construction rather than a purely technical setting.
No comments yet — start the discussion below.