Amirhossein Ghanipour Amirhandeh · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22917885
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
The rapid adoption of discrete neural audio codecs (e.g., EnCodec) in large language models necessitates codec-invariant speech watermarking to track synthetic data provenance. However, current continuous-domain paradigms inject payloads as surface-level waveform perturbations, which are systematically destroyed by non-linear quantization bottlenecks. This extended abstract presents an empirical structural autopsy of joint codec-watermark optimization. We demonstrate that traditional approaches result in a catastrophic trade-off: either adversarial representation collapse or complete acoustic destruction (PESQ < 1.6). We mathematically prove that the continuous base latent space of these codecs is too densely packed with semantic information to harbor hidden data, even when constrained by orthogonal psychoacoustic projections. To resolve this, we introduce Hierarchical Residual Injection. By mathematically bypassing the primary semantic codebooks and explicitly targeting the mid-level residual error (Codebook 4), we decouple the payload from the speech formants. Evaluations on EnCodec-24kHz demonstrate that this targeted injection achieves a 0.0% Bit Error Rate (BER) under complete re-synthesis attacks while restoring speech intelligibility and perceptual quality.
No comments yet — start the discussion below.