Amirhossein Ghanipour Amirhandeh · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22900286
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
This work introduces a learnable, fully differentiable raw-waveform front-end designed to overcome the rigid time-frequency trade-offs of fixed Short-Time Fourier Transforms (STFT) as well as the sample inefficiency, parameter bloat, and phase vulnerabilities of unconstrained 1D temporal convolutions. The proposed architecture parameterizes an analytic Gabor filterbank governed by three physical inductive biases: (1) a cumulative softmax reparameterization that guarantees strictly monotonic center frequency allocation, entirely preventing filter crossover, harmonic clustering, and spectral dead zones; (2) an analytic Constant-Q temporal scaling constraint coupled with Hilbert quadrature carrier decomposition to extract instantaneous amplitude envelopes without zero-crossing phase cancellation; and (3) a lightweight temporal context module that dynamically modulates channel-wise gains to adapt to non-stationary acoustic bursts and vocal-tract dynamics. When integrated into an ECAPA-TDNN backbone and optimized with Additive Angular Margin Softmax (ArcFace), the proposed front-end reduces verification error from 5.40% to 3.90% Equal Error Rate (EER)—a 27.8% relative error reduction—over a matched unconstrained 1D convolutional baseline on LibriSpeech verification trials, while reducing the core filter receptive field parameters by over 99.5% (128 versus 25,664 weights).
No comments yet — start the discussion below.