Georgia Institute of Technology
Automatic speaker verification and diarization on VoxCeleb data collection
Abstract
dc:description.abstractAutomatic speaker verification (ASV) is increasingly getting more attention in speech research field in recent years. Because of the importance of cyber-security and personal property security, ASV can be used in many fields in the future in addition to fingerprint and face information. In ASV research, a variety of datasets are needed to train good models. Current datasets include NIST SRE, VoxCeleb, etc. In this work, to collect a non-English speaking dataset, the pipeline of VoxCeleb data collection is adopted to collect an East Asian language-speaking Celebrities (EACeleb) dataset. To remove some noisy segments of the output and make the dataset cleaner, speaker diarization is used in this research and the collected data is filtered. Due to the lack of ground truth labels of the collected data, ASV is used to measure the data cleanness improvement of our dataset. Equal error rate (EER) can be lowered by 25.63% after speaker diarization compared to the original EACeleb using a pretrained x-vector model for measurement. Also, by training the speaker verification using EACeleb data, when testing the EER performance, EACeleb after diarization can outperform VoxCeleb by 36.78%.
Degree
thesis:*- Level thesis:degree_level
- Masters
- Department dc:contributor.department
- Electrical and Computer Engineering
- Grantor dc:publisher
- Georgia Institute of Technology
- Year dc:date.issued
- 2020
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Yang, Yufeng
- Advisor dc:contributor.advisor
-
- Anderson, David V.
- Committee members dc:contributor.committeemember
-
- Davenport, Mark
- Bloch, Matthieu
Subjects
dc:subject × 4Rights
- Language dc:language.iso
- en_US
Identifiers
dc:identifier.*- Handle dc:identifier.uri
- http://hdl.handle.net/1853/62830
- OAI identifier oai:identifier
- oai:repository.gatech.edu:1853/62830