A CLOSE-SET MULTI-SPEAKER IDENTIFICATION via LIGHTWEIGHT CNN ARCHITECTURES LEVERAGING MFCC-DERIVED SPECTRAL FEATURES

Authors:

N. K. KAPHUNGKUI,

DOI NO:

https://doi.org/10.26782/jmcms.2026.08.00012

Keywords:

MFCC,CNN,Classification,Confusion matrix,Speaker Identification,

Abstract

Speaker identification is an important task in speech processing, with applications in security, authentication, forensics, and personalised human–computer interaction. This study presents a lightweight convolutional neural network (CNN)-based system for close-set identification of 40 speakers using Mel-Frequency Cepstral Coefficients (MFCCs). A dataset comprising 1,000 speech samples was collected from 40 speakers, with 850 samples used for training and 150 for validation. MFCC features were extracted from the speech signals and used as input to a compact CNN consisting of four convolutional blocks, three intermediate pooling operations, adaptive average pooling, and a fully connected classification layer. The proposed architecture contains 102,792 trainable parameters. Under the evaluated controlled conditions, the model achieved 100% validation accuracy, with all 150 validation samples correctly classified, and achieved an ROC-AUC of 1.00 for all 40 speaker classes. Pooling ablation experiments showed that removing any one of the three intermediate pooling operations reduced the best validation accuracy from 100.00% to 99.33%, supporting the use of the complete pooling configuration. An augmentation sensitivity analysis further showed that the unaugmented configuration achieved the highest validation accuracy of 100.00%, while pitch-shift augmentation resulted in 99.00% and Gaussian-noise augmentation reduced accuracy to 2.67%, both independently and when combined with pitch shifting. These findings indicate that the unaugmented configuration was the most effective setting for the proposed model under the evaluated conditions. Overall, the proposed lightweight CNN demonstrates effective closed-set speaker identification under controlled recording conditions, while further evaluation is required for open-set recognition, speaker verification, and greater recording variability.

Refference:

I. Brümmer, Niko, and Johan de Villiers. “The BOSARIS Toolkit and Calibration Methods for Speaker Recognition Scoring.” Proceedings of the 2011 NIST Speaker and Language Recognition Workshop, 2011.
II. B. Tan, M. H. A. Hijazi, N. Khamis, P. N. E. binti Nohuddin, Z. Zainol, F. Coenen, and A. Gani, “A survey on presentation attack detection for automatic speaker verification systems: State-of-the-art, taxonomy, issues and future direction,” Multimedia Tools and Applications, vol. 80, nos. 21–23, pp. 32725–32762, 2021. doi: 10.1007/s11042-021-11235-x
III. Chakraborty, Koustav, Asmita Talele, and Savitha Upadhya. “Voice Recognition Using MFCC Algorithm.” International Journal of Innovative Research in Advanced Engineering, vol. 1, no. 10, 2014, pp. 158–161.
IV. Chung, Joon Son, Arsha Nagrani, and Andrew Zisserman. “VoxCeleb2: Deep Speaker Recognition.” Interspeech 2018, 2018, pp. 1086–1090. 10.21437/Interspeech.2018-1929.
V. Dehak, Najim, Patrick Kenny, Réda Dehak, Pierre Dumouchel, and Pierre Ouellet. “Front-End Factor Analysis for Speaker Verification.” IEEE Transactions on Audio, Speech, and Language Processing, vol. 19, no. 4, 2011, pp. 788–798. 10.1109/TASL.2010.2064307.
VI. Desplanques, Brecht, Jenthe Thienpondt, and Kris Demuynck. “ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification.” Interspeech 2020, 2020, pp. 3830–3834. 10.21437/Interspeech.2020-2650.
VII. Desplanques, Brecht, Jenthe Thienpondt, and Kris Demuynck. “Large Margin Fine-Tuning for Speaker Recognition.” 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020. 10.1109/ICASSP40776.2020.9054465.
VIII. Ferrer, Luciana, Matthew McLaren, and Niko Brümmer. “Robust Back-End Calibration and Adaptation Techniques for Speaker Recognition.” Computer Speech & Language.
IX. Garcia-Romero, Daniel, and Carol Y. Espy-Wilson. “Analysis of i-Vector Length Normalization in Speaker Recognition Systems.” Interspeech 2011, 2011, pp. 249–252. 10.21437/Interspeech.2011-53.
X. Garcia-Romero, Daniel, and Aaron McCree. “Multicondition Training of Gaussian PLDA Models in i-Vector Space for Noise and Reverberation Robust Speaker Recognition.” 2012 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2012, pp. 4257–4260. 10.1109/ICASSP.2012.6288859.
XI. Garcia-Romero, Daniel, et al. “The UMD-JHU 2011 Speaker Recognition System.” 2012 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2012, pp. 4225–4228. 10.1109/ICASSP.2012.6288852.
XII. He, Y., et al. “Exploration of Data Augmentation and Augmentation Strategies for Speaker Recognition.” Conference paper.
XIII. Heigold, Georg, Ignacio Moreno, Samy Bengio, and Noam Shazeer. “End-to-End Text-Dependent Speaker Verification.” 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2016, pp. 5115–5119. 10.1109/ICASSP.2016.7472652.
XIV. Ioffe, Sergey. “Probabilistic Linear Discriminant Analysis.” Computer Vision—ECCV 2006 Workshops, Springer, 2006, pp. 531–542. 10.1007/11744085_41.
XV. Ji, R., et al. “An End-to-End Text-Independent Speaker Identification Framework.” Interspeech 2018, 2018, pp. 3267–3271.
XVI. Kenny, Patrick. “Joint Factor Analysis of Speaker and Session Variability: Theory and Algorithms.” Technical Report CRIM-06/08-13, 2005.
XVII. Khan, Suhail Ahmad, Anil S. Thosar, Jagannath H. Nirmal, and Vinay S. Pande. “A Unique Approach in Text Independent Speaker Recognition Using MFCC Feature Sets and Probabilistic Neural Network.” 2015 Eighth International Conference on Advances in Pattern Recognition (ICAPR), IEEE, 2015, pp. 1–6.
XVIII. Kinnunen, Tomi, and Haizhou Li. “An Overview of Text-Independent Speaker Recognition: From Features to Supervectors.” Speech Communication, vol. 52, no. 1, 2010, pp. 12–40. 10.1016/j.specom.2009.08.009.
XIX. Lee, Kong Aik, et al. “Data-Efficient Speaker Recognition and Few-Shot Adaptation Techniques.” Conference paper.
XX. Lei, Yun, et al. “A Novel Scheme for Speaker Recognition Using a PLDA Variant.” 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2014, pp. 1695–1699.
XXI. Nagrani, Arsha, Joon Son Chung, and Andrew Zisserman. “VoxCeleb: A Large-Scale Speaker Identification Dataset.” Interspeech 2017, 2017, pp. 2616–2620. 10.21437/Interspeech.2017-950.
XXII. Okabe, Koji, Takafumi Koshinaka, and Koichi Shinoda. “Attentive Statistics Pooling for Deep Speaker Embedding.” Interspeech 2018, 2018, pp. 2252–2256. 10.21437/INTERSPEECH.2018-993.
XXIII. Okada, H., et al. “Improvements in i-Vector PLDA Scoring and Calibration for NIST SREs.” IEEE workshop/conference paper.
XXIV. Povey, Daniel, et al. “The Kaldi Speech Recognition Toolkit.” 2011 IEEE Workshop on Automatic Speech Recognition and Understanding, 2011.
XXV. Ravanelli, Mirco, and Yoshua Bengio. “Speaker Recognition from Raw Waveform with SincNet.” Interspeech 2018, 2018, pp. 3698–3702. 10.21437/Interspeech.2018-2492.
XXVI. Reynolds, Douglas A., Thomas F. Quatieri, and Robert B. Dunn. “Speaker Verification Using Adapted Gaussian Mixture Models.” Digital Signal Processing, vol. 10, nos. 1–3, 2000, pp. 19–41. 10.1006/dspr.1999.0362.
XXVII. Sigona, Francesco, et al. “Validation of an ECAPA-TDNN System for Forensic Automatic Speaker Recognition.” Speech Communication, 2024. 10.1016/j.specom.2024.103045.
XXVIII. Snyder, David, et al. “Spoken Language Recognition Using X-Vectors.” Odyssey 2018: The Speaker and Language Recognition Workshop, 2018, pp. 105–111. 10.21437/Odyssey.2018-15.
XXIX. Snyder, David, Daniel Garcia-Romero, Gregory Sell, Daniel Povey, and Sanjeev Khudanpur. “X-Vectors: Robust DNN Embeddings for Speaker Recognition.” 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2018, pp. 5329–5333. 10.1109/ICASSP.2018.8461375.
XXX. Sztahó, Dániel, Gábor Szaszák, and Anna Beke. “Deep Learning Methods in Speaker Recognition: A Review.” Periodica Polytechnica Electrical Engineering and Computer Science, vol. 65, no. 4, 2021, pp. 310–328. 10.3311/PPee.17024.
XXXI. Variani, Ehsan, Xin Lei, Erik McDermott, Ignacio Lopez Moreno, and Javier Gonzalez-Dominguez. “Deep Neural Networks for Small-Footprint Text-Dependent Speaker Verification.” 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2014, pp. 4052–4056. 10.1109/ICASSP.2014.6854363.
XXXII. Villalba, Jesús, and Niko Brümmer. “Towards Fully Bayesian Speaker Recognition: Integrating Out the Between-Speaker Covariance.” Interspeech 2011, 2011, pp. 505–508. 10.21437/Interspeech.2011-142.
XXXIII. Villalba, Jesús, et al. “Domain Adaptation Techniques for PLDA and Neural Back-Ends.” ICASSP/Interspeech proceedings.
XXXIV. Wan, Li, Quan Wang, Alan Papir, and Ignacio Lopez Moreno. “Generalized End-to-End Loss for Speaker Verification.” 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2018, pp. 4879–4883. 10.1109/ICASSP.2018.8462665.
XXXV. Wu, Zhizheng, Tomi Kinnunen, Nicholas Evans, Junichi Yamagishi, Cemal Hanilçi, Md. Sahidullah, and Aleksandr Sizov. “ASVspoof 2015: The First Automatic Speaker Verification Spoofing and Countermeasures Challenge.” Interspeech 2015, 2015, pp. 2037–2041. 10.21437/Interspeech.2015-462.
XXXVI. Wu, Zhizheng, et al. “ASVspoof 2017/2019 Challenge Series: Challenge Overview and Datasets.” Interspeech Proceedings.

View Download