Advancements of Audio Unimodal Deep Faking Detection Technology

Authors

  • Minglei Zhu School of Computer Science and Engineering, Wuhan Institute of Technology, 430200 Wuhan, China

DOI:

https://doi.org/10.54097/9j1ktg80

Keywords:

Audio Deepfake Detection, Deep Learning Models, Generalization & Robustness.

Abstract

In recent years, the generative artificial intelligence technology has undergone rapid iterations, leading to the widespread abuse of audio deep forgery methods such as speech synthesis and speech conversion, which seriously threaten information security and social credibility. Among various countermeasures, audio single-modal deep forgery detection has advantages such as lightweight and strong real-time performance, and is a key technology for security protection in pure audio scenarios, with significant research and application value. This paper systematically reviews the types of audio forgery, detection principles, authoritative datasets and evaluation indicators, compares and analyzes traditional detection methods with deep learning detection techniques, focuses on elaborating the characteristics and applicable scenarios of four mainstream detection models, and points out the core challenges in generalization, robustness, etc. of current methods. Finally, it looks forward to the development trends of this field towards generalization, high robustness, lightweight and forgery traceability, which can provide references for related research.

References

[1] A. Khan, K. M. Malik, J. Ryan, M. Saravanan, Battling Voice Spoofing: A Review, Comparative Analysis, and Generalizability Evaluation of State-of-the-Art Voice Spoofing Countermeasures. IEEE TASLP 32, 1234-1256 (2024)

[2] H. M. Tran, D. L. Olive, D. Guennec, et al., Leveraging SSL Speech Features and Mamba for Enhanced DeepFake Detection, in Proceedings of Interspeech 2025 (ISCA Press, Doha, 2025), 4572-4576

[3] X. Li, Y. Wang, H. Zhang, Robust DeepFake Audio Detection via an Improved NeXt-TDNN with Multi-Fused Self-Supervised Learning Features. Appl. Sci. 15, 9685 (2025)

[4] J. Xue, C. Fan, J. Yi, J. Zhou, Z. Lv, Dynamic Ensemble Teacher-Student Distillation Framework for Light-Weight Fake Audio Detection. IEEE Access 12, 4567-4589 (2024)

[5] Z. Zhang, W. Hao, A. Sankoh, W. Lin, E. Mendiola-Ortiz, J. Yang, C. Mao, I Can Hear You: Selective Robust Training for Deepfake Audio Detection, in Proceedings of NeurIPS 2025 (MIT Press, Montreal, 2025), 7890-7901

[6] Z. Wu, M. Li, J. Chen, CodecFake: Enhancing Anti-Spoofing Models Against Deepfake Audios from Codec-Based Speech Synthesis Systems, in Proceedings of Interspeech 2024 (ISCA Press, Kos, 2024), 1770-1774

[7] S. Chen, Y. Liu, T. Zhao, Cross-Domain Audio Deepfake Detection: Dataset and Analysis, in Proceedings of EMNLP 2024 (ACL, Stroudsburg, 2024), 4567-4579

[8] R. Liu, J. H. Zhang, G. L. Gao, et al., Betray Oneself: A Novel Audio DeepFake Detection Model via Mono-to-Stereo Conversion, in Proceedings of Interspeech 2023 (ISCA Press, Dublin, 2023), 1876-1880

[9] Y. Gao, T. Vuong, M. Elyasi, et al., Generalized Spoofing Detection Inspired from Audio Generation Artifacts, in Proceedings of Interspeech 2021 (ISCA Press, Brno, 2021), 3689-3693

[10] Y. Xie, Z. Li, H. Wang, FakeSound: Deepfake General Audio Detection, in Proceedings of Interspeech 2024 (ISCA Press, Kos, 2024), 112-116

[11] Y. Chen, J. Yi, J. Xue, et al., RawBMamba: End-to-End Bidirectional State Space Model for Audio Deepfake Detection, in Proceedings of Interspeech 2024 (ISCA Press, Kos, 2024), 2890-2894

[12] X. Li, H. Wang, S. Zhang, Where are We in Audio Deepfake Detection? A Systematic Analysis over Generative and Detection Models. ACM Trans. Internet Technol. 25, 1-28 (2025)

[13] B. Zhang, H. Cui, V. Nguyen, et al., Audio Deepfake Detection: What Has Been Achieved and What Lies Ahead. Sensors 25, 1989 (2025)

[14] X. Xiang, P.-Y. Chen, W. Wei, Measuring the Robustness of Audio Deepfake Detectors. arXiv:2503.17577 (2025)

[15] N. Klein, T. Chen, H. Tak, et al., Source Tracing of Audio Deepfake Systems, in Proceedings of Interspeech 2024 (ISCA Press, Kos, 2024), 3678-3682

[16] S. Barrington, R. Barua, G. Koorma, et al., Single and Multi-Speaker Cloned Voice Detection: From Perceptual to Learned Features, in Proceedings of IEEE WIFS 2023 (IEEE Press, Nürnberg, 2023), 1-6

[17] Z. Wu, N. Evans, T. Kinnunen, et al., Spoofing and Countermeasures for Speaker Verification: A Survey. Speech Commun. 66, 130-153 (2015)

[18] O. Pascu, A. Stan, D. Oneata, et al., Towards Generalisable and Calibrated Audio Deepfake Detection with Self-Supervised Representations, in Proceedings of Interspeech 2024 (ISCA Press, Kos, 2024), 4123-4127

[19] Z. Wu, M. Li, J. Chen, Deepfake Audio Detection in Voice Authentication: A Spectral and CNN-Based Comprehensive Review. Eng. Technol. Appl. Sci. Res. 15, 13400-13412 (2025)

[20] D.-T. Truong, R. Tao, T. Nguyen, et al., Temporal-Channel Modeling in Multi-Head Self-Attention for Synthetic Speech Detection. arXiv:2406.17376 (2024)

[21] Z. Zhang, T. Pan, H. Liu, et al., Mixture of Low-Rank Adapter Experts in Generalizable Audio Deepfake Detection. arXiv:2509.13878 (2025)

[22] J. M. Martin-Donas, A. Alvarez, E. Rosello, et al., Exploring Self-Supervised Embeddings and Synthetic Data Augmentation for Robust Audio Deepfake Detection, in Proceedings of Interspeech 2024 (ISCA Press, Kos, 2024), 4567-4571

[23] Z. Wu, M. Li, J. Chen, End-to-End Audio Deepfake Detection from RAW Waveforms: a RawNet-Based Approach with Cross-Dataset Evaluation. arXiv:2504.20923 (2025)

Downloads

Published

13-08-2026

Issue

Section

Articles

How to Cite

Zhu, M. (2026). Advancements of Audio Unimodal Deep Faking Detection Technology. Mathematical Modeling and Algorithm Application, 9(3), 63-69. https://doi.org/10.54097/9j1ktg80