University of Illinois at Urbana-Champaign
Extracting single talker segments from audio mixtures in reverberant environments
Abstract
dc:descriptionClassifying the number of simultaneous speakers in a reverberant environment remains a difficult problem to solve. The less complicated but no less important problem of detecting whether a single speaker or multiple speakers are present also poses a challenge, as multipath often distorts or alters the features commonly used in speaker number classification. When this classification system is used as a front-end to a learning-based system, then it becomes imperative that the classifier be accurate. This thesis presents SinguDetect, an end-to-end system which uses the inter-aural phase differences between a pair of in-ear microphones to classify each 64 ms time window of audio as belonging to one speaker or not. Even in heavily reverberant environments (RT60 = 2.0 s), SinguDetect is still able to maintain a precision of around 0.8 where the number of simultaneous sources is not higher than three, whereas other state-of-the-art algorithms tend to suffer decreases in precision and f-score in this range.
Degree
thesis:*- Name thesis:degree_name
- M.S.
- Level thesis:degree_level
- Thesis
- Discipline thesis:degree_discipline
- Electrical & Computer Engr
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2023
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Subramaniam, Avinash
- Contributors dc:contributor
-
- Choudhury, Romit Roy
Subjects
dc:subject × 3Rights
dc:rights- Statement dc:rights
-
- Copyright 2023 Avinash Subramaniam
- Language dc:language
- en, eng
Identifiers
dc:identifier.*- Handle dc:identifier
- https://hdl.handle.net/2142/121540