
Yupei Li, Yixiong Fang, Jiahao Xue, Manuel Milling, Björn W. Schuller · Frontiers in Artificial Intelligence 2026 · 2026
DOI: 10.3389/frai.2026.1928592
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Deepfakes have raised widespread concern owing to their threats to privacy, security, and societal trust, driving growing research interest in effective detection methods. Audio deepfake detection has moved from handcrafted-feature classifiers to end-to-end deep learning architectures and, more recently, to self-supervised speech foundation models and general-purpose speech recognizers such as Whisper. The use of instruction-following speech large language models (LLMs), of which Speech Audio Language Music Open Neural Network (SALMONN) is a representative example, remains comparatively less explored, and forms the setting of this study. We argue that large language models (LLMs), with their powerful representational and information retrieval capabilities, hold significant potential for this task—provided that relevant audio features are carefully selected and prioritized. To this end, we employ a speech LLM, specifically SALMONN, for deepfake audio detection, incorporating a two-stage training strategy in which Stage 1 injects acoustic-feature awareness into the Low-Rank Adaptation (LoRA) adapters via an auxiliary feature-prediction pretext task (referred to as feature dropin), and Stage 2 performs windowed random retention on the encoder token sequence (referred to as feature dropout) before the final real/fake decision. Our approach achieves an absolute accuracy improvement of 0.188 on the FakeOrReal benchmark and attains state-of-the-art performance on the Automatic Speaker Verification Spoofing and Countermeasures (ASVspoof) 2019 LA subset with an Equal Error Rate (EER) of 0.00496, representing a 33.0% relative improvement over prior methods. Furthermore, the model transfers non-trivially to an out-of-domain artificial intelligence (AI)-generated music dataset (M6), reaching accuracy comparable to a ResNet18 baseline reported in prior study, which we interpret as evidence of transferable audio representations rather than broad cross-domain robustness.
No comments yet — start the discussion below.