Deep Neural Network based Speaker Identification, Comparing Three Different Models.

Loading...
Thumbnail Image

Issue Date

Language

en

Document type

Keywords

Publisher

Alternative Title

Title

ISSN

Volume

Issue

Startpage

Endpage

DOI

Abstract

Speaker identification is a classification task which is a popular authentication model. The current state of the art offers a variety of different approaches for both non Deep Neural Network-based as well as Deep Neural Network-based methods. However most of the current Deep Neural Network-based methods consist of either Convolutional Neural Networks or Recurrent Neural Networks. The combination of these models seems to be a rare case. By using the key benefits of both, this research aims to explore the use of Convolutional Recurrent Neural Networks in speaker identification tasks as well as compare it to other types of models. Three models have been created for this purpose, a Feedforward Neural Network, a Convolutional Neural Network and a Convolutional Recurrent Neural Network. These networks were created and tuned with an English database, and later also tested with a mixed database consisting of five considerably different languages. The experimental research showed that the Feedforward and Convolutional Neural Networks performed excellently, whereas the Convolutional Recurrent Neural Network did not perform well. Despite the bad performance over all, when dealing with more complex data, the mixed database, the performance of the Convolutional Recurrent Neural Network increased with from a mere 29% to 45%. It should be noted that the scope of this research was very limited, since the lack of computational resources was a significant constraint. With better resources, application of Convolutional Recurrent Neural Networks in the field of speaker identification and other related fields could still be beneficial. Especially with respect to real-world mimicking scenarios such as language independence and dealing with noise. Hence through further optimisation and investigation, this method could become a robust and reliable approach. Keywords: Speaker identification, speaker recognition, voice identification, deep learning, convolutional neural network, recurrent neural network, Mel spectrograms

Description

Citation

Faculty

Faculteit der Sociale Wetenschappen

License

PubMed ID

EISSN

Endorsement

Review

Supplemented By

Referenced By