{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/80822"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/80822","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Computational Models for Binaural Sound Source Localization and Sound Understanding","abstract":"As one of humans' primary sensors, the auditory system plays an important role in language acquisition. Computational models for binaural sound source localization and sound source understanding are proposed in this thesis. The models build a fundamental auditory system for a mobile robot that will automatically learn language through multisensory inputs and interaction with the external environment. A hypothesis-driven approach is followed for the localization model. Using only binaural inputs, it enables three-dimensional (3D) localization by combining multiple cues. Two binaural localization cues, interaural time differences (ITDs) and interaural intensity differences (IIDs), and one monoaural localization cue, spectral cues, are extracted from the input sounds. A Bayes rule-based hierarchical framework is applied for decision making. Simulations show the effectiveness of the model. A robust ITD estimation algorithm is introduced and implemented on the robot. Satisfactory results are achieved under real-world environments. A multimodal learning scheme is proposed with the aid of vision to realize autonomous learning for the 3D binaural localization. No human instructors need to be involved. A generic model is presented for sound source understanding. No labelled training data is required to build the model. A histogram is employed as the sound representation, where the time-varying characteristics of sound can be preserved. Histogram intersection is used as the similarity measurement between different sounds. The model is successfully applied to content-based audio information retrieval and automatic audio indexing systems.","abstract_html":"As one of humans&#x27; primary sensors, the auditory system plays an important role in language acquisition. Computational models for binaural sound source localization and sound source understanding are proposed in this thesis. The models build a fundamental auditory system for a mobile robot that will automatically learn language through multisensory inputs and interaction with the external environment. A hypothesis-driven approach is followed for the localization model. Using only binaural inputs, it enables three-dimensional (3D) localization by combining multiple cues. Two binaural localization cues, interaural time differences (ITDs) and interaural intensity differences (IIDs), and one monoaural localization cue, spectral cues, are extracted from the input sounds. A Bayes rule-based hierarchical framework is applied for decision making. Simulations show the effectiveness of the model. A robust ITD estimation algorithm is introduced and implemented on the robot. Satisfactory results are achieved under real-world environments. A multimodal learning scheme is proposed with the aid of vision to realize autonomous learning for the 3D binaural localization. No human instructors need to be involved. A generic model is presented for sound source understanding. No labelled training data is required to build the model. A histogram is employed as the sound representation, where the time-varying characteristics of sound can be preserved. Histogram intersection is used as the similarity measurement between different sounds. The model is successfully applied to content-based audio information retrieval and automatic audio indexing systems.","abstract_has_math":false,"creators":["Li, Danfeng"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical Engineering","degree_department":null,"school":null,"contributors":["Levinson, Stephen E."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2015,"date_issued":"2015-09-25T20:08:20Z","date_published":"2015-09-25T20:08:20Z","updated_at":"2026-07-22T22:26:15Z","subjects":["Computer Science"],"languages":["eng"],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["(MiAaPQ)AAI3086118"],"render_values":[{"text":"(MiAaPQ)AAI3086118","href":null,"code":true}]}]},"links":{"outbound_url":"http://hdl.handle.net/2142/80822","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Levinson, Stephen E."]},{"key":"dc:creator","label":"Author","values":["Li, Danfeng"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2015-09-25T20:08:20Z","10000-01-01","2003"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical Engineering"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer Science"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/80822","(MiAaPQ)AAI3086118"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["As one of humans' primary sensors, the auditory system plays an important role in language acquisition. Computational models for binaural sound source localization and sound source understanding are proposed in this thesis. The models build a fundamental auditory system for a mobile robot that will automatically learn language through multisensory inputs and interaction with the external environment. A hypothesis-driven approach is followed for the localization model. Using only binaural inputs, it enables three-dimensional (3D) localization by combining multiple cues. Two binaural localization cues, interaural time differences (ITDs) and interaural intensity differences (IIDs), and one monoaural localization cue, spectral cues, are extracted from the input sounds. A Bayes rule-based hierarchical framework is applied for decision making. Simulations show the effectiveness of the model. A robust ITD estimation algorithm is introduced and implemented on the robot. Satisfactory results are achieved under real-world environments. A multimodal learning scheme is proposed with the aid of vision to realize autonomous learning for the 3D binaural localization. No human instructors need to be involved. A generic model is presented for sound source understanding. No labelled training data is required to build the model. A histogram is employed as the sound representation, where the time-varying characteristics of sound can be preserved. Histogram intersection is used as the similarity measurement between different sounds. The model is successfully applied to content-based audio information retrieval and automatic audio indexing systems.","Made available in DSpace on 2015-09-25T20:08:20Z (GMT). No. of bitstreams: 2 license.txt: 4848 bytes, checksum: 96035ab3f5e1c23cc7138a224ce498bd (MD5) 3086118.pdf: 5744066 bytes, checksum: 56d8f1343e0af6494c52eb02249a8d82 (MD5) Previous issue date: 2003","Embargo set by: Seth Robbins for item 82104 Lift date: Forever Reason: Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","U of I Only","108 p.","Thesis (Ph.D.)--University of Illinois at Urbana-Champaign, 2003."]},{"key":"dc:title","label":"Title","values":["Computational Models for Binaural Sound Source Localization and Sound Understanding"]}]}],"canonical_facts":{"dc:contributor":["Levinson, Stephen E."],"dc:creator":["Li, Danfeng"],"dc:date":["2015-09-25T20:08:20Z","10000-01-01","2003"],"dc:description":["As one of humans' primary sensors, the auditory system plays an important role in language acquisition. Computational models for binaural sound source localization and sound source understanding are proposed in this thesis. The models build a fundamental auditory system for a mobile robot that will automatically learn language through multisensory inputs and interaction with the external environment. A hypothesis-driven approach is followed for the localization model. Using only binaural inputs, it enables three-dimensional (3D) localization by combining multiple cues. Two binaural localization cues, interaural time differences (ITDs) and interaural intensity differences (IIDs), and one monoaural localization cue, spectral cues, are extracted from the input sounds. A Bayes rule-based hierarchical framework is applied for decision making. Simulations show the effectiveness of the model. A robust ITD estimation algorithm is introduced and implemented on the robot. Satisfactory results are achieved under real-world environments. A multimodal learning scheme is proposed with the aid of vision to realize autonomous learning for the 3D binaural localization. No human instructors need to be involved. A generic model is presented for sound source understanding. No labelled training data is required to build the model. A histogram is employed as the sound representation, where the time-varying characteristics of sound can be preserved. Histogram intersection is used as the similarity measurement between different sounds. The model is successfully applied to content-based audio information retrieval and automatic audio indexing systems.","Made available in DSpace on 2015-09-25T20:08:20Z (GMT). No. of bitstreams: 2 license.txt: 4848 bytes, checksum: 96035ab3f5e1c23cc7138a224ce498bd (MD5) 3086118.pdf: 5744066 bytes, checksum: 56d8f1343e0af6494c52eb02249a8d82 (MD5) Previous issue date: 2003","Embargo set by: Seth Robbins for item 82104 Lift date: Forever Reason: Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","U of I Only","108 p.","Thesis (Ph.D.)--University of Illinois at Urbana-Champaign, 2003."],"dc:identifier":["http://hdl.handle.net/2142/80822","(MiAaPQ)AAI3086118"],"dc:language":["eng"],"dc:subject":["Computer Science"],"dc:title":["Computational Models for Binaural Sound Source Localization and Sound Understanding"],"dc:type":["text"],"thesis:degree_discipline":["Electrical Engineering"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:26:15Z"}