Deep Neural Network based Speaker Identification, Comparing Three Different Models.
Loading...
Authors
Issue Date
Language
en
Document type
Keywords
Publisher
Alternative Title
Title
ISSN
Volume
Issue
Startpage
Endpage
DOI
Abstract
Speaker identification is a classification task which is a popular authentication model. The current state of the art offers a variety of different approaches for both non Deep Neural Network-based as well as Deep Neural Network-based methods. However most of the current Deep Neural Network-based methods consist of either Convolutional Neural Networks or Recurrent Neural Networks. The combination of these models seems to be a rare case. By using the key benefits of both, this research aims to explore the use of Convolutional Recurrent Neural Networks in speaker identification tasks as well as compare it to other types of models. Three models have been created for this purpose, a Feedforward Neural Network, a Convolutional Neural Network and a Convolutional Recurrent Neural Network. These networks were created and tuned with an English database, and later also tested with a mixed database consisting of five considerably different languages. The experimental research showed that the Feedforward and Convolutional Neural Networks performed excellently, whereas the Convolutional Recurrent Neural Network did not perform well. Despite the bad performance over all, when dealing with more complex data, the mixed database, the performance of the Convolutional Recurrent Neural Network increased with from a mere 29% to 45%. It should be noted that the scope of this research was very limited, since the lack of computational resources was a significant constraint. With better resources, application of Convolutional Recurrent Neural Networks in the field of speaker identification and other related fields could still be beneficial. Especially with respect to real-world mimicking scenarios such as language independence and dealing with noise. Hence through further optimisation and investigation, this method could become a robust and reliable approach.
Keywords: Speaker identification, speaker recognition, voice identification, deep learning, convolutional neural network, recurrent neural network, Mel spectrograms
Description
Citation
Supervisor
Faculty
Faculteit der Sociale Wetenschappen
