{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/29841"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/29841","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Non-speech Acoustic Event Detection Using Multimodal Information","abstract":"Non-speech acoustic event detection (AED) aims to recognize events that are relevant to human activities associated with audio information. Much previous research has been focused on restricted highlight events, and highly relied on ad-hoc detectors for these events. This thesis focuses on using multimodal data in order to make non-speech acoustic event detection and classification tasks more robust, requiring no expensive annotation. To be specific, the thesis emphasizes designing suitable feature representations for different modalities and fusing the information properly. Two cases are studied in this thesis: (1) Acoustic event detection in a meeting room scenario using single-microphone audio cues and single-camera visual cues. Non-speech event cues often exist in both audio and vision, but not necessarily in a synchronized fashion. We jointly model audio and visual cues in order to improve event detection using multistream HMMs and coupled HMMs (CHMM). Spatial pyramid histograms based on the optical flow are proposed as a generalizable visual representation that does not require training on labeled video data. In a multimedia meeting room non-speech event detection task, the proposed methods outperform previously reported systems leveraging ad-hoc visual object detectors and sound localization information obtained from multiple microphones. (2) Multimodal feature representation for person detection at border crossings. Based on phenomenology of the differences between humans and four-legged animals, we propose using enhanced autocorrelation pattern for feature extraction for seismic sensors, and an exemplar selection framework for acoustic sensors. We also propose using temporal pattens from ultrasonic sensors. We perform decision and feature fusion to combine the information from all three modalities. From experimental results, we show that our proposed methods improve the robustness of the system.","abstract_html":"Non-speech acoustic event detection (AED) aims to recognize events that are relevant to human activities associated with audio information. Much previous research has been focused on restricted highlight events, and highly relied on ad-hoc detectors for these events. This thesis focuses on using multimodal data in order to make non-speech acoustic event detection and classification tasks more robust, requiring no expensive annotation. To be specific, the thesis emphasizes designing suitable feature representations for different modalities and fusing the information properly. Two cases are studied in this thesis: (1) Acoustic event detection in a meeting room scenario using single-microphone audio cues and single-camera visual cues. Non-speech event cues often exist in both audio and vision, but not necessarily in a synchronized fashion. We jointly model audio and visual cues in order to improve event detection using multistream HMMs and coupled HMMs (CHMM). Spatial pyramid histograms based on the optical flow are proposed as a generalizable visual representation that does not require training on labeled video data. In a multimedia meeting room non-speech event detection task, the proposed methods outperform previously reported systems leveraging ad-hoc visual object detectors and sound localization information obtained from multiple microphones. (2) Multimodal feature representation for person detection at border crossings. Based on phenomenology of the differences between humans and four-legged animals, we propose using enhanced autocorrelation pattern for feature extraction for seismic sensors, and an exemplar selection framework for acoustic sensors. We also propose using temporal pattens from ultrasonic sensors. We perform decision and feature fusion to combine the information from all three modalities. From experimental results, we show that our proposed methods improve the robustness of the system.","abstract_has_math":false,"creators":["Huang, Po-Sen"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Hasegawa-Johnson, Mark A."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2012,"date_issued":"2012-02-06T20:21:11Z","date_published":"2012-02-06T20:21:11Z","updated_at":"2026-07-22T22:25:29Z","subjects":["Acoustic Event Detection","Optical Flow","Hidden Markov Models","Multistream Hidden Markov Models","Coupled Hidden Markov Models","Gaussian Mixture Models","Support Vector Machines","Sensor Fusion","Footstep Detection","Person Detection"],"languages":["en"],"rights":["Copyright 2011 Po-Sen Huang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/29841","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Hasegawa-Johnson, Mark A."]},{"key":"dc:creator","label":"Author","values":["Huang, Po-Sen"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2012-02-06T20:21:11Z","2011-12"]},{"key":"dc:type","label":"Dc Type","values":["Dissertation / Thesis","text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Acoustic Event Detection","Optical Flow","Hidden Markov Models","Multistream Hidden Markov Models","Coupled Hidden Markov Models","Gaussian Mixture Models","Support Vector Machines","Sensor Fusion","Footstep Detection","Person Detection"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2011 Po-Sen Huang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/29841"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Non-speech acoustic event detection (AED) aims to recognize events that are relevant to human activities associated with audio information. Much previous research has been focused on restricted highlight events, and highly relied on ad-hoc detectors for these events. This thesis focuses on using multimodal data in order to make non-speech acoustic event detection and classification tasks more robust, requiring no expensive annotation. To be specific, the thesis emphasizes designing suitable feature representations for different modalities and fusing the information properly. Two cases are studied in this thesis: (1) Acoustic event detection in a meeting room scenario using single-microphone audio cues and single-camera visual cues. Non-speech event cues often exist in both audio and vision, but not necessarily in a synchronized fashion. We jointly model audio and visual cues in order to improve event detection using multistream HMMs and coupled HMMs (CHMM). Spatial pyramid histograms based on the optical flow are proposed as a generalizable visual representation that does not require training on labeled video data. In a multimedia meeting room non-speech event detection task, the proposed methods outperform previously reported systems leveraging ad-hoc visual object detectors and sound localization information obtained from multiple microphones. (2) Multimodal feature representation for person detection at border crossings. Based on phenomenology of the differences between humans and four-legged animals, we propose using enhanced autocorrelation pattern for feature extraction for seismic sensors, and an exemplar selection framework for acoustic sensors. We also propose using temporal pattens from ultrasonic sensors. We perform decision and feature fusion to combine the information from all three modalities. From experimental results, we show that our proposed methods improve the robustness of the system.","Item withdrawn by Alexis Thompson (athmpsn1@illinois.edu) on 2011-12-01T22:43:19Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 2 MS_thesis.zip: 10007868 bytes, checksum: f9b3438431cf41d8fbd4c10404190659 (MD5) Huang_Po-Sen.pdf: 1577256 bytes, checksum: d07559064aec14eb4b5ccef5bcc3a9b9 (MD5)","Made available in DSpace on 2012-02-06T20:21:11Z (GMT). No. of bitstreams: 3 Huang_Po-Sen.pdf: 1577256 bytes, checksum: d07559064aec14eb4b5ccef5bcc3a9b9 (MD5) license.txt: 4058 bytes, checksum: 5d6f02a759b6df660f7f35e5cdfcad20 (MD5) MS_thesis.zip: 10007868 bytes, checksum: f9b3438431cf41d8fbd4c10404190659 (MD5)"]},{"key":"dc:title","label":"Title","values":["Non-speech Acoustic Event Detection Using Multimodal Information"]}]}],"canonical_facts":{"dc:contributor":["Hasegawa-Johnson, Mark A."],"dc:creator":["Huang, Po-Sen"],"dc:date":["2012-02-06T20:21:11Z","2011-12"],"dc:description":["Non-speech acoustic event detection (AED) aims to recognize events that are relevant to human activities associated with audio information. Much previous research has been focused on restricted highlight events, and highly relied on ad-hoc detectors for these events. This thesis focuses on using multimodal data in order to make non-speech acoustic event detection and classification tasks more robust, requiring no expensive annotation. To be specific, the thesis emphasizes designing suitable feature representations for different modalities and fusing the information properly. Two cases are studied in this thesis: (1) Acoustic event detection in a meeting room scenario using single-microphone audio cues and single-camera visual cues. Non-speech event cues often exist in both audio and vision, but not necessarily in a synchronized fashion. We jointly model audio and visual cues in order to improve event detection using multistream HMMs and coupled HMMs (CHMM). Spatial pyramid histograms based on the optical flow are proposed as a generalizable visual representation that does not require training on labeled video data. In a multimedia meeting room non-speech event detection task, the proposed methods outperform previously reported systems leveraging ad-hoc visual object detectors and sound localization information obtained from multiple microphones. (2) Multimodal feature representation for person detection at border crossings. Based on phenomenology of the differences between humans and four-legged animals, we propose using enhanced autocorrelation pattern for feature extraction for seismic sensors, and an exemplar selection framework for acoustic sensors. We also propose using temporal pattens from ultrasonic sensors. We perform decision and feature fusion to combine the information from all three modalities. From experimental results, we show that our proposed methods improve the robustness of the system.","Item withdrawn by Alexis Thompson (athmpsn1@illinois.edu) on 2011-12-01T22:43:19Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 2 MS_thesis.zip: 10007868 bytes, checksum: f9b3438431cf41d8fbd4c10404190659 (MD5) Huang_Po-Sen.pdf: 1577256 bytes, checksum: d07559064aec14eb4b5ccef5bcc3a9b9 (MD5)","Made available in DSpace on 2012-02-06T20:21:11Z (GMT). No. of bitstreams: 3 Huang_Po-Sen.pdf: 1577256 bytes, checksum: d07559064aec14eb4b5ccef5bcc3a9b9 (MD5) license.txt: 4058 bytes, checksum: 5d6f02a759b6df660f7f35e5cdfcad20 (MD5) MS_thesis.zip: 10007868 bytes, checksum: f9b3438431cf41d8fbd4c10404190659 (MD5)"],"dc:identifier":["http://hdl.handle.net/2142/29841"],"dc:language":["en"],"dc:rights":["Copyright 2011 Po-Sen Huang"],"dc:subject":["Acoustic Event Detection","Optical Flow","Hidden Markov Models","Multistream Hidden Markov Models","Coupled Hidden Markov Models","Gaussian Mixture Models","Support Vector Machines","Sensor Fusion","Footstep Detection","Person Detection"],"dc:title":["Non-speech Acoustic Event Detection Using Multimodal Information"],"dc:type":["Dissertation / Thesis","text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:29Z"}