{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/105158"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/105158","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Acoustic event, spoken keyword and emotional outburst detection","abstract":"This thesis presents work in research topics of audio detection. It first describes a system for large-scale multi-label acoustic event detection (AED) in YouTube videos. It explores the potential of the state-of-the-art deep learning classifiers for AED, describes both qualitative and quantitative results (Hit@1 is 47.9%) and presents the pre-trained embedding model as a powerful feature extractor to be adapted to new domains with limited data and improve the detection accuracy (Hit@1 is 58.1%). Second, the thesis focuses on the speech acoustic events and the spoken keyword spotting task for speech. It presents a phonetic keyword spotter as a lightweight alternative to full speech recognition (3x faster, with comparable detection rates and that addresses automatic speech recognition problems). It also explores cross-lingual keyword spotting to support low resource languages and finds that the acoustic model is dominant in determining the cross-lingual keyword search performance. Third, the thesis further presents the emotional outburst detection for infant nonspeech acoustic events. It reports on the efforts to manually code child utterances as being of type “laugh,” “cry,” “fuss,” “babble,” and “hiccup” and to develop the algorithms capable of performing the same task automatically.","abstract_html":"This thesis presents work in research topics of audio detection. It first describes a system for large-scale multi-label acoustic event detection (AED) in YouTube videos. It explores the potential of the state-of-the-art deep learning classifiers for AED, describes both qualitative and quantitative results (Hit@1 is 47.9%) and presents the pre-trained embedding model as a powerful feature extractor to be adapted to new domains with limited data and improve the detection accuracy (Hit@1 is 58.1%). Second, the thesis focuses on the speech acoustic events and the spoken keyword spotting task for speech. It presents a phonetic keyword spotter as a lightweight alternative to full speech recognition (3x faster, with comparable detection rates and that addresses automatic speech recognition problems). It also explores cross-lingual keyword spotting to support low resource languages and finds that the acoustic model is dominant in determining the cross-lingual keyword search performance. Third, the thesis further presents the emotional outburst detection for infant nonspeech acoustic events. It reports on the efforts to manually code child utterances as being of type “laugh,” “cry,” “fuss,” “babble,” and “hiccup” and to develop the algorithms capable of performing the same task automatically.","abstract_has_math":false,"creators":["Xu, Yijia"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Hasegawa-Johnson, Mark Allen"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019-08-23T20:44:38Z","date_published":"2019-08-23T20:44:38Z","updated_at":"2026-07-22T22:24:44Z","subjects":["audio event detection","spoken keyword detection","emotion detection","speech recognition","convolutional neural network","hidden Markov model","phonetic keywork spotter"],"languages":["en"],"rights":["Copyright 2019 Yijia Xu"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/105158","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Hasegawa-Johnson, Mark Allen"]},{"key":"dc:creator","label":"Author","values":["Xu, Yijia"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2019-08-23T20:44:38Z","2021-08-24T09:15:38Z","2019-04-03","2019-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["audio event detection","spoken keyword detection","emotion detection","speech recognition","convolutional neural network","hidden Markov model","phonetic keywork spotter"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2019 Yijia Xu"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/105158"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["This thesis presents work in research topics of audio detection. It first describes a system for large-scale multi-label acoustic event detection (AED) in YouTube videos. It explores the potential of the state-of-the-art deep learning classifiers for AED, describes both qualitative and quantitative results (Hit@1 is 47.9%) and presents the pre-trained embedding model as a powerful feature extractor to be adapted to new domains with limited data and improve the detection accuracy (Hit@1 is 58.1%). Second, the thesis focuses on the speech acoustic events and the spoken keyword spotting task for speech. It presents a phonetic keyword spotter as a lightweight alternative to full speech recognition (3x faster, with comparable detection rates and that addresses automatic speech recognition problems). It also explores cross-lingual keyword spotting to support low resource languages and finds that the acoustic model is dominant in determining the cross-lingual keyword search performance. Third, the thesis further presents the emotional outburst detection for infant nonspeech acoustic events. It reports on the efforts to manually code child utterances as being of type “laugh,” “cry,” “fuss,” “babble,” and “hiccup” and to develop the algorithms capable of performing the same task automatically.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2021-05-01","The student, Yijia Xu, accepted the attached license on 2019-04-03 at 10:56.","The student, Yijia Xu, submitted this Thesis for approval on 2019-04-03 at 11:12.","This Thesis was approved for publication on 2019-04-03 at 15:29.","DSpace SAF Submission Ingestion Package generated from Vireo submission #13484 on 2019-08-22 at 16:20:39","Made available in DSpace on 2019-08-23T20:44:38Z (GMT). No. of bitstreams: 2 XU-THESIS-2019.pdf: 590207 bytes, checksum: 19cae6b0b3dc905564d87371d9b4499e (MD5) LICENSE.txt: 4205 bytes, checksum: c212041f120bb11d1bebaaf39ac769dd (MD5) Previous issue date: 2019-04-03","Embargo set by: Seth Robbins for item 112277 Lift date: 2021-08-23T20:44:50Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 112277 Lift date: 2021-08-23T20:46:41Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 112277 Lift date: 2021-08-23T20:47:38Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 112277 Lift date: 2021-08-23T20:48:32Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 112277 on 2021-08-24T09:15:38Z."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Acoustic event, spoken keyword and emotional outburst detection"]}]}],"canonical_facts":{"dc:contributor":["Hasegawa-Johnson, Mark Allen"],"dc:creator":["Xu, Yijia"],"dc:date":["2019-08-23T20:44:38Z","2021-08-24T09:15:38Z","2019-04-03","2019-05"],"dc:description":["This thesis presents work in research topics of audio detection. It first describes a system for large-scale multi-label acoustic event detection (AED) in YouTube videos. It explores the potential of the state-of-the-art deep learning classifiers for AED, describes both qualitative and quantitative results (Hit@1 is 47.9%) and presents the pre-trained embedding model as a powerful feature extractor to be adapted to new domains with limited data and improve the detection accuracy (Hit@1 is 58.1%). Second, the thesis focuses on the speech acoustic events and the spoken keyword spotting task for speech. It presents a phonetic keyword spotter as a lightweight alternative to full speech recognition (3x faster, with comparable detection rates and that addresses automatic speech recognition problems). It also explores cross-lingual keyword spotting to support low resource languages and finds that the acoustic model is dominant in determining the cross-lingual keyword search performance. Third, the thesis further presents the emotional outburst detection for infant nonspeech acoustic events. It reports on the efforts to manually code child utterances as being of type “laugh,” “cry,” “fuss,” “babble,” and “hiccup” and to develop the algorithms capable of performing the same task automatically.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2021-05-01","The student, Yijia Xu, accepted the attached license on 2019-04-03 at 10:56.","The student, Yijia Xu, submitted this Thesis for approval on 2019-04-03 at 11:12.","This Thesis was approved for publication on 2019-04-03 at 15:29.","DSpace SAF Submission Ingestion Package generated from Vireo submission #13484 on 2019-08-22 at 16:20:39","Made available in DSpace on 2019-08-23T20:44:38Z (GMT). No. of bitstreams: 2 XU-THESIS-2019.pdf: 590207 bytes, checksum: 19cae6b0b3dc905564d87371d9b4499e (MD5) LICENSE.txt: 4205 bytes, checksum: c212041f120bb11d1bebaaf39ac769dd (MD5) Previous issue date: 2019-04-03","Embargo set by: Seth Robbins for item 112277 Lift date: 2021-08-23T20:44:50Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 112277 Lift date: 2021-08-23T20:46:41Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 112277 Lift date: 2021-08-23T20:47:38Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 112277 Lift date: 2021-08-23T20:48:32Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 112277 on 2021-08-24T09:15:38Z."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/105158"],"dc:language":["en"],"dc:rights":["Copyright 2019 Yijia Xu"],"dc:subject":["audio event detection","spoken keyword detection","emotion detection","speech recognition","convolutional neural network","hidden Markov model","phonetic keywork spotter"],"dc:title":["Acoustic event, spoken keyword and emotional outburst detection"],"dc:type":["text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:44Z"}