Shaotong Huang, Shihao Fang · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23092852
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Forced alignment is a key preprocessing step in speech processing, and the Montreal Forced Aligner (MFA) is one of the most widely used open-source alignment tools. However, the stability and robustness of MFA alignment under audio perturbations, and the effect of denoising preprocessing, have not been systematically quantified. This paper presents a two-part error testing and analysis of MFA 3.4.2 on the LibriSpeech train-clean-360 dataset (921 speakers, 104,014 utterances, ~360 hours, 16 kHz mono FLAC). Experiment 1 applies 18 types of audio perturbations (gain, white/strong white noise, resampling, downsampling, time shift, MP3/OPUS codec, reverberation, telephone-bandwidth filtering, low-pass filtering, clipping, speed change, pitch shift, frequency-band masking, combined perturbations, and pink/Brown/band-pass noise) and measures phoneme-onset boundary deviations against a clean baseline. Each perturbation was aligned three times to estimate MFA's intrinsic jitter. Experiment 2 evaluates two denoising methods—spectral subtraction (noisereduce) and deep learning (audio-denoiser)—on white noise and realistic-like noise (pink, Brown, band-pass). Results: MFA is highly robust to perturbations that preserve spectral structure (σ < 5 ms, outlier rate < 0.6%), but sensitive to perturbations that alter spectral or temporal structure (σ = 10–25 ms, outlier rate 5.5–8.9%). Speed change alters the phoneme sequence itself and is analyzed separately. Denoising is ineffective for white noise (deep-learning denoising is harmful) and only marginally, non-significantly improves realistic-like noise (p = 0.127). MFA also exhibits heavy-tailed repeated-alignment jitter (median 0 ms, P99 10 ms, P99.9 70 ms, max 330 ms), which should serve as the baseline in any perturbation experiment. At the phoneme level, SH is the most sensitive (16.55×), and pitch shift is the worst perturbation for all phonemes. The manuscript was originally written in Chinese and translated into English with the assistance of an AI-based translation tool. The author(s) have reviewed and edited the translation and take full responsibility for the content.
No comments yet — start the discussion below.