{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/117723"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/117723","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Efficient audio-visual representations for reasoning and synthesis tasks","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2023-04-12 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2023-04-12 without embargo terms","abstract_has_math":false,"creators":["Chatterjee, Moitreya"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Ahuja, Narendra","Hasegawa-Johnson, Mark A","Do, Minh N","Gupta, Saurabh","Wang, Yuxiong","Owens, Andrew","Harwath, David"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-12","date_published":"2022-12","updated_at":"2026-07-22T22:24:56Z","subjects":["Audio-visual","Scene Understanding","Multimodal","Frame Generation","Audio Source Separation","Machine Learning","Computer Vision"],"languages":["en","eng"],"rights":["Copyright 2022 Moitreya Chatterjee"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/117723","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Ahuja, Narendra","Hasegawa-Johnson, Mark A","Do, Minh N","Gupta, Saurabh","Wang, Yuxiong","Owens, Andrew","Harwath, David"]},{"key":"dc:creator","label":"Author","values":["Chatterjee, Moitreya"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022-12","2022-10-28"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Audio-visual","Scene Understanding","Multimodal","Frame Generation","Audio Source Separation","Machine Learning","Computer Vision"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2022 Moitreya Chatterjee"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/117723"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2023-04-12 without embargo terms","The student, Moitreya Chatterjee, accepted the attached license on 2022-09-15 at 00:22.","The student, Moitreya Chatterjee, submitted this Dissertation for approval on 2022-09-15 at 00:41.","This Dissertation was approved for publication on 2022-10-28 at 09:47.","DSpace SAF Submission Ingestion Package generated from Vireo submission #18490 on 2023-04-12 at 07:23:33","Events that occur in the real world often leave acoustic and visual imprints. Highly evolved organisms, such as human beings, have capabilities to perceive such events jointly across multiple modalities (such as audio and video). This allows for faster and more accurate perception with the possibility of supplementing information lacunae in one modality by another. As we seek to enable artificial intelligence (AI) systems to complement humans in their endeavors, it is therefore critical that they be equipped with such mulitmodal reasoning capabilities as well and be able to undertake tasks that humans usually solve effortlessly. Towards this end, this dissertation explores reasoning and synthesis tasks in the audio-visual space, particularly, the tasks of: (a) disambiguating mixed/noisy audio by leveraging video, and (b) being able to predict a video from audio. This dissertation proposes methods to tackle these challenges while striving to ensure that the proposed solutions be deployable in resource-constrained environments where there might be an inadequacy of high-performance computing resources or a paucity of training data. Concretely this dissertation makes the following four contributions: (i) a novel geometry-aware sparse scene graph based representation is proposed to undertake audio-source separation given videos of the audio sources in their natural settings; (ii) a multimodal, variational encoder-decoder model, called Sound2Sight, is introduced to synthesize the frames of a video given the audio coherently; (iii) an improved regime for training, unimodal, and audio-conditioned frame prediction systems, which factors in the predictive uncertainty of the model, is put forth. This adaptation results in such models requiring lesser data and fewer epochs for training; (iv) finally, a method to compress the dominant image representation tool in multimodal deep neural networks, Convolutional Neural Network (CNN), is presented so that they might operate in environments with lower computation capacity by exploiting filter activation patterns and inter-filter weight dependencies."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Efficient audio-visual representations for reasoning and synthesis tasks"]}]}],"canonical_facts":{"dc:contributor":["Ahuja, Narendra","Hasegawa-Johnson, Mark A","Do, Minh N","Gupta, Saurabh","Wang, Yuxiong","Owens, Andrew","Harwath, David"],"dc:creator":["Chatterjee, Moitreya"],"dc:date":["2022-12","2022-10-28"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2023-04-12 without embargo terms","The student, Moitreya Chatterjee, accepted the attached license on 2022-09-15 at 00:22.","The student, Moitreya Chatterjee, submitted this Dissertation for approval on 2022-09-15 at 00:41.","This Dissertation was approved for publication on 2022-10-28 at 09:47.","DSpace SAF Submission Ingestion Package generated from Vireo submission #18490 on 2023-04-12 at 07:23:33","Events that occur in the real world often leave acoustic and visual imprints. Highly evolved organisms, such as human beings, have capabilities to perceive such events jointly across multiple modalities (such as audio and video). This allows for faster and more accurate perception with the possibility of supplementing information lacunae in one modality by another. As we seek to enable artificial intelligence (AI) systems to complement humans in their endeavors, it is therefore critical that they be equipped with such mulitmodal reasoning capabilities as well and be able to undertake tasks that humans usually solve effortlessly. Towards this end, this dissertation explores reasoning and synthesis tasks in the audio-visual space, particularly, the tasks of: (a) disambiguating mixed/noisy audio by leveraging video, and (b) being able to predict a video from audio. This dissertation proposes methods to tackle these challenges while striving to ensure that the proposed solutions be deployable in resource-constrained environments where there might be an inadequacy of high-performance computing resources or a paucity of training data. Concretely this dissertation makes the following four contributions: (i) a novel geometry-aware sparse scene graph based representation is proposed to undertake audio-source separation given videos of the audio sources in their natural settings; (ii) a multimodal, variational encoder-decoder model, called Sound2Sight, is introduced to synthesize the frames of a video given the audio coherently; (iii) an improved regime for training, unimodal, and audio-conditioned frame prediction systems, which factors in the predictive uncertainty of the model, is put forth. This adaptation results in such models requiring lesser data and fewer epochs for training; (iv) finally, a method to compress the dominant image representation tool in multimodal deep neural networks, Convolutional Neural Network (CNN), is presented so that they might operate in environments with lower computation capacity by exploiting filter activation patterns and inter-filter weight dependencies."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/117723"],"dc:language":["en","eng"],"dc:rights":["Copyright 2022 Moitreya Chatterjee"],"dc:subject":["Audio-visual","Scene Understanding","Multimodal","Frame Generation","Audio Source Separation","Machine Learning","Computer Vision"],"dc:title":["Efficient audio-visual representations for reasoning and synthesis tasks"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:56Z"}