{"id":{"repo_id":"buffalo","oai_identifier":"oai:ubir.buffalo.edu:10477/80964"},"canonical_url":"https://search.dev.ndltd.org/etd/buffalo/oai:ubir.buffalo.edu:10477/80964","repository":{"repo_id":"buffalo","name":"Buffalo","base_url":"https://ubir.buffalo.edu/oai/request"},"display":{"title":"Learning to Build Multimodal Intelligence Across Vision, Language and Speech","abstract":"Ph.D.","abstract_html":"Ph.D.","abstract_has_math":false,"creators":["Ma, Shuang"],"institution":"State University of New York at Buffalo","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":["Chen, Chang Wen","Computer Science and Engineering"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019-10-29T16:48:33Z","date_published":"2019-10-29T16:48:33Z","updated_at":"2026-07-27T19:05:28Z","subjects":["computer science"],"languages":["eng"],"rights":["Users of works found in University at Buffalo Institutional Repository (UBIR) are responsible for identifying and contacting the copyright owner for permission to reuse. University at Buffalo Libraries do not manage rights for copyright-protected works and cannot assist with permissions.","Copyright retained by author."],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/10477/80964","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Chen, Chang Wen","Computer Science and Engineering"]},{"key":"dc:creator","label":"Author","values":["Ma, Shuang"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2019-10-29T16:48:33Z","2019","2019-08-19 14:32:13"]},{"key":"dc:publisher","label":"Institution","values":["State University of New York at Buffalo"]},{"key":"dc:type","label":"Dc Type","values":["Text","Dissertation"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["computer science"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Users of works found in University at Buffalo Institutional Repository (UBIR) are responsible for identifying and contacting the copyright owner for permission to reuse. University at Buffalo Libraries do not manage rights for copyright-protected works and cannot assist with permissions.","Copyright retained by author."]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/10477/80964"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Ph.D.","Artificial intelligence can already do many things that humans cannot. But how far away are we from building ``human-like\" AI? What are the key problems that we need to solve before we get there?Human gain and evolve intelligence by absorbing and accumulating knowledge from multiple sources, which can also be applied for how to build machine intelligence. In this research, we seek to enable machine to have multimodal intelligence like human beings.One key obstacle for building Multimodal Intelligence is that data from different modalities does not share the same representation. As this dissertation focuses on Machine Learning and Artificial Intelligence research, we seek for building the cornerstone with the universal perceptive knowledge, which directly represents data from different modalities into a common space. In such a system, multimodal data can be perceived, understood and fused into the intelligence building process.We accomplish the goal of building multimodal intelligence by mainly focusing on three steps, i.e. learning to perceive knowledge from a single modal, learning to align knowledge across modalities and further learning to fuse knowledge from multiple modalities.In this dissertation, we present novel techniques that are developed by learning from Vision, Language and Speech. We also show how such techniques can effectively resolve many real world problems across different modalities.It is very challenging to achieve the above goals:first, to make machines have the ability to understand, especially for images, the performance of deep CNN methods is often compromised by the constraint that the neural network only takes the fixed-size input. To accommodate this requirement, input images need to be properly transformed. Thus the high level information of the original images is impaired because of potential loss of fine grained details and holistic image layout."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Learning to Build Multimodal Intelligence Across Vision, Language and Speech"]}]}],"canonical_facts":{"dc:contributor":["Chen, Chang Wen","Computer Science and Engineering"],"dc:creator":["Ma, Shuang"],"dc:date":["2019-10-29T16:48:33Z","2019","2019-08-19 14:32:13"],"dc:description":["Ph.D.","Artificial intelligence can already do many things that humans cannot. But how far away are we from building ``human-like\" AI? What are the key problems that we need to solve before we get there?Human gain and evolve intelligence by absorbing and accumulating knowledge from multiple sources, which can also be applied for how to build machine intelligence. In this research, we seek to enable machine to have multimodal intelligence like human beings.One key obstacle for building Multimodal Intelligence is that data from different modalities does not share the same representation. As this dissertation focuses on Machine Learning and Artificial Intelligence research, we seek for building the cornerstone with the universal perceptive knowledge, which directly represents data from different modalities into a common space. In such a system, multimodal data can be perceived, understood and fused into the intelligence building process.We accomplish the goal of building multimodal intelligence by mainly focusing on three steps, i.e. learning to perceive knowledge from a single modal, learning to align knowledge across modalities and further learning to fuse knowledge from multiple modalities.In this dissertation, we present novel techniques that are developed by learning from Vision, Language and Speech. We also show how such techniques can effectively resolve many real world problems across different modalities.It is very challenging to achieve the above goals:first, to make machines have the ability to understand, especially for images, the performance of deep CNN methods is often compromised by the constraint that the neural network only takes the fixed-size input. To accommodate this requirement, input images need to be properly transformed. Thus the high level information of the original images is impaired because of potential loss of fine grained details and holistic image layout."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/10477/80964"],"dc:language":["eng"],"dc:publisher":["State University of New York at Buffalo"],"dc:rights":["Users of works found in University at Buffalo Institutional Repository (UBIR) are responsible for identifying and contacting the copyright owner for permission to reuse. University at Buffalo Libraries do not manage rights for copyright-protected works and cannot assist with permissions.","Copyright retained by author."],"dc:subject":["computer science"],"dc:title":["Learning to Build Multimodal Intelligence Across Vision, Language and Speech"],"dc:type":["Text","Dissertation"]},"updated_at":"2026-07-27T19:05:28Z"}