Deepfake Audio Detection Using MFCC-Based Acoustic Feature Extraction and Machine Learning Classification: A Comprehensive Framework
DOI:
https://doi.org/10.63503/acset.90Keywords:
Deepfake audio detection, MFCC, SVM, audio forensics, speech processing, machine learning, cybersecurity, synthetic speechAbstract
Voice has always been one of the more trusted ways to identify people; we can identify our colleagues, family, and public figures by their voices. Meanwhile, that trust is being systematically exploited. The state of neural speech synthesis today means that AI-generated speech can sound convincing in casual conversations and, as in a handful of cited cases, has been used for financial fraud and targeted deception. In this paper, a detection system using Mel-Frequency Cepstral Coefficients (MFCCs) and a Support Vector Machine (SVM) classifier is presented to classify audio clips as real or generated. The system does not only use the direct MFCC values, but it also computes five other complementary spectral features: spectral centroid, spectral roll-off, zero crossing rate, RMS energy and spectral bandwidth, from which it forms a 45-dimensional feature vector. Noise is filtered, loudness normalised, duration trimmed, and silence standardised via audio pre-processing before any features are computed. The trained classifier is encapsulated within a Flask web app that accepts uploaded audio files, performs inference, and outputs predictions in real time. Results on a balanced benchmark test set of 4,000 clips yielded 95.2% accuracy, 94.1% precision, 95.8% recall, and 94.9% F1-score. These numbers outperform several CNN and TCN baselines in the recent literature at a significantly reduced computational cost. The paper suggests that when executed conscientiously, principled feature engineering can be a viable alternative to large-scale deep learning for this class of problems.
References
[1] A. Hamza, A. R. R. Javed, F. Iqbal, N. Kryvinska, A. S. Almadhor, Z. Jalil, and R. Borghol, "Deepfake audio detection via MFCC features using machine learning," IEEE Access, vol. 10, pp. 134018–134030, 2022.
[2] K. Verma, P. Singh, R. Ali, and A. Kumar, "Deepfake audio detection: A comparative study of advanced deep learning models," IEEE Access, vol. 13, pp. 45012–45028, 2025.
[3] A. R. Javed, W. Ahmed, M. Alazab, Z. Jalil, K. Kifayat, and T. R. Gadekallu, "A comprehensive survey on computer forensics: State-of-the-art, tools, techniques, challenges, and future directions," IEEE Access, vol. 10, pp. 11065–11089, 2022.
[4] A. Ahmed, A. R. Javed, Z. Jalil, G. Srivastava, and T. R. Gadekallu, "Privacy of web browsers: A challenge in digital forensics," in Proc. Int. Conf. Genetic Evol. Comput., Springer, 2021, pp. 493–504.
[5] H. Malik, S. Parikh, and A. Javed, "Audio forensics and deepfake detection techniques: A survey," J. Cybersecurity Res., vol. 4, no. 2, pp. 88–104, 2024.
[6] B. Paris and J. Donovan, "Deepfakes and cheap fakes," Data Soc., New York, NY, USA, 2019.
[7] R. Wijetunga, D. Matheesha, A. Al Noman, K. De Silva, M. Tissera, and L. Rupasinghe, "Deepfake audio detection: A deep learning based solution for group conversations," in Proc. ICAC, 2020, pp. 192–197.
[8] C. Stupp, "Fraudsters used AI to mimic CEO's voice in unusual cybercrime case," Wall Street J., Aug. 2019.
[9] T. T. Nguyen et al., "Deep learning for deepfakes creation and detection: A survey," arXiv:1909.11573, 2019.
[10] Z. Khanjani, G. Watson, and V. P. Janeja, "How deep are the fakes? Focusing on audio deepfake: A survey," arXiv:2111.14203, 2021.
[11] Z. Wu et al., "ASVspoof 2015: The first automatic speaker verification spoofing and countermeasures challenge," in Proc. Interspeech, 2015.
[12] T. Kinnunen et al., "The ASVspoof 2017 challenge: Assessing the limits of replay spoofing attack detection," in Proc. Interspeech, 2017, pp. 2–6.
[13] J. Yamagishi et al., "ASVspoof 2019: Automatic speaker verification spoofing and countermeasures challenge evaluation plan," 2019.
[14] Z. Almutairi and H. Elgibreen, "A review of modern audio deepfake detection methods: Challenges and future directions," Algorithms, vol. 15, no. 5, p. 155, 2022.
[15] Y. Chen et al., "Probabilistic forecasting with temporal convolutional neural network," Neurocomputing, vol. 399, pp. 491–501, 2020.
[16] Y. Kawaguchi, "Anomaly detection based on feature reconstruction from subsampled audio signals," in Proc. EUSIPCO, 2018, pp. 2524–2528.
[17] Y. Kawaguchi and T. Endo, "How can we detect anomalies from subsampled audio signals?" in Proc. IEEE MLSP, 2017, pp. 1–6.
[18] H. Yu et al., "Spoofing detection in automatic speaker verification systems using DNN classifiers and dynamic acoustic features," IEEE Trans. Neural Netw. Learn. Syst., vol. 29, no. 10, pp. 4633–4644, 2018.
[19] S. Pradhan et al., "Combating replay attacks against voice assistants," Proc. ACM IMWUT, vol. 3, no. 3, pp. 1–26, 2019.
[20] F. Tom, M. Jain, and P. Dey, "End-to-end audio replay attack detection using deep convolutional networks with attention," in Proc. Interspeech, 2018, pp. 681–685.
[21] K. Kuligowska, P. Kisielewicz, and A. Włodarz, "Speech synthesis systems: Disadvantages and limitations," Int. J. Res. Eng. Technol., vol. 7, no. 83, pp. 234–239, 2018.
[22] A. van den Oord et al., "WaveNet: A generative model for raw audio," arXiv:1609.03499, 2016.
[23] J. Shen et al., "Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions," in Proc. ICASSP, 2018, pp. 4779–4783.
[24] J. Frank and L. Schönherr, "WaveFake: A data set to facilitate audio deepfake detection," arXiv:2111.02813, 2021.
[25] R. Reimao and V. Tzerpos, "FoR: A dataset for synthetic speech detection," in Proc. SpeD, 2019, pp. 1–10.
[26] W. Ping et al., "Deep Voice 3: Scaling text-to-speech with convolutional sequence learning," arXiv:1710.07654, 2017.
[27] F. M. Rammo and M. N. Al-Hamdani, "Detecting the speaker language using CNN deep learning algorithm," Iraqi J. Comput. Sci. Math., vol. 3, no. 1, pp. 43–52, 2022.
[28] S. Ahmed et al., "Speaker identification model based on deep neural networks," Iraqi J. Comput. Sci. Math., vol. 3, no. 1, pp. 108–114, 2022.
[29] A. Winursito, R. Hidayat, and A. Bejo, "Improvement of MFCC feature extraction accuracy using PCA in Indonesian speech recognition," in Proc. ICOIACT, 2018, pp. 379–383.
[30] J. Khochare, C. Joshi, B. Yenarkar, S. Suratkar, and F. Kazi, "A deep learning framework for audio deepfake detection," Arabian J. Sci. Eng., vol. 47, pp. 1–12, 2021.
[31] P. Yu, Z. Xia, J. Fei, and Y. Lu, "A survey on deepfake video detection," IET Biometrics, vol. 10, no. 6, pp. 607–624, 2021.
[32] N. Akhtar, S. Verma, and M. Hossain, "A review of challenges and trends in deepfake audio detection," Digit. Signal Process., vol. 141, p. 103798, 2024.
[33] Z. Jiang et al., "Deepfake audio detection using dual-branch CNN with attention mechanism," IEEE Trans. Inf. Forensics Security, vol. 19, pp. 512–525, 2024.
[34] R. Mohamed Abdulhamied et al., "Deepfake audio detection using feature-based and deep learning approaches," Int. J. Adv. Comput. Sci. Appl., vol. 16, no. 6, 2025.
[35] N. Eldien, R. Ali, and F. Moussa, "Real and fake face detection: A comprehensive evaluation of machine learning and deep learning techniques," in Proc. Int. Conf. Comput. Sci., 2023, pp. 315–320.
[36] S. Ö. Arık, H. Jun, and G. Diamos, "Fast spectrogram inversion using multi-head convolutional neural networks," IEEE Signal Process. Lett., vol. 26, no. 1, pp. 94–98, Jan. 2019.
[37] Y. Zhang and P. Singh, "Machine learning approaches for speech spoofing detection," in Proc. Int. Conf. AI Security, 2023.
[38] J. Kietzmann et al., "Deepfakes: Trick or treat?" Bus. Horizons, vol. 63, no. 2, pp. 135–146, 2020.
[39] S. Basak et al., "Challenges and limitations in speech recognition technology," CMES, vol. 135, no. 2, 2023.
[40] M. Shanthamallappa et al., "Robust automatic speech recognition using wavelet based adaptive thresholding: A review," SN Comput. Sci., vol. 5, no. 2, p. 248, 2024.
Downloads
Published
Conference Proceedings Volume
Section
License
Copyright (c) 2026 Adroid Conference Series: Engineering and Technology

This work is licensed under a Creative Commons Attribution 4.0 International License.