{# Audit 04/10/2026 : « autre » n'est pas un code de langue ; SPHAERO n'est pas l'éditeur des documents qu'elle héberge ou référence. #} {# citation_pdf_url doit mener à un PDF : un lien vers une page DOI est pénalisé par Google Scholar (avant : tout lien externe). #}
Accès ouvert · CC BY

Fusion of Cochleogram and Mel Spectrogram Features for Deep Learning Based Speaker Recognition

Article scientifique 2022 Anglais

Résumé

Abstract Speaker recognition has crucial application in forensic science, financial areas, access control, surveillance and law enforcement. The performance of speaker recognition get degraded with the noise, speakers physical and behavioral changes. Fusion of Mel Frequency Cepstral Coefficient (MFCC) and Gammatone Frequency Cepstral Coefficient (GFCC) features are used to improve the performance of machine learning based speaker recognition systems in the noisy condition. Deep learning models, especially Convolutional Neural Network (CNN) and its hybrid approaches outperform machine learning approaches in speaker recognition. Previous CNN based speaker recognition models has used Mel Spectrogram features as an input. Even though, Mel Spectrogram features show better performance compared to the handcrafted features, its performance get degraded with noise and behavioral changes of speaker. In this work, a CNN based speaker recognition model is developed using fusion of Mel Spectrogram and Cochleogram feature as input. The speaker recognition performance of the fusion of Mel Spectrogram and Cochleogram features is compared with the performance of Mel Spectrogram and Cochleogram features without fusing. The train-clean-100 part of the LibriSpeech dataset, which consists of 251 speakers (126 male and 125 female speakers) and 28,539 utterances is used for the experiment of proposed model. CNN model is trained and evaluated for 20 epochs using training and validation data respectively. Proposed speaker recognition model which uses fusion of Mel Spectrogram and Cochleogram as input for CNN has accuracy of 99.56%. Accuracy of CNN based speaker recognition with Mel Spectrogram is 98.15% and Cochleogram features is 97.43%. The results show that fusion of Mel Spectrogram and Cochleogram features improve the performance of speaker recognition.

Citer ce document

Lambamo, W., Srinivasa, R., & Jifara, W. (2022). Fusion of Cochleogram and Mel Spectrogram Features for Deep Learning Based Speaker Recognition. Research Square. https://doi.org/10.21203/rs.3.rs-2139057/v1

Exporter : BibTeX · RIS (Zotero, Mendeley, EndNote)

Accès au document

Texte intégral en lecture en ligne, réservé aux abonnés SPHAERO et aux membres de l'institution. Se connecter

Voir l'article sur le site de la revue

Licence et provenance

Licence : CC BY

Notice moissonnée depuis OpenAlex le 09/09/2026. Le document reste hébergé par sa source.
Voir le document à la source →

Statistiques

Consultations : 1

Téléchargements : 0