{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/121435"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/121435","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Event-centric multimodal knowledge acquisition","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2023-12-04 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2023-12-04 without embargo terms","abstract_has_math":false,"creators":["Li, Manling"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Ji, Heng","Han, Jiawei","Zhai, Chengxiang","Chang, Shih-Fu","Cho, Kyunghyun"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2023,"date_issued":"2023-08","date_published":"2023-08","updated_at":"2026-07-22T22:24:57Z","subjects":["Multimodal","Knowledge","Event-centric"],"languages":["en","eng"],"rights":["Copyright 2023 Manling Li"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/121435","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Ji, Heng","Han, Jiawei","Zhai, Chengxiang","Chang, Shih-Fu","Cho, Kyunghyun"]},{"key":"dc:creator","label":"Author","values":["Li, Manling"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2023-08","2023-07-10"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Multimodal","Knowledge","Event-centric"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2023 Manling Li"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/121435"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2023-12-04 without embargo terms","The student, Manling Li, accepted the attached license on 2023-07-07 at 10:13.","The student, Manling Li, submitted this Dissertation for approval on 2023-07-07 at 10:41.","This Dissertation was approved for publication on 2023-07-10 at 08:57.","DSpace SAF Submission Ingestion Package generated from Vireo submission #19467 on 2023-12-04 at 17:00:14","What happened? Who? When? Where? Why? What will happen next? are the fundamental questions asked to comprehend the overwhelming amount of information. Answers to these questions are the core knowledge communicated through multiple forms of information, regardless of whether presented as text, images, videos, audio, or other modalities. To obtain such knowledge from multimodal data, this dissertation focuses on Multimodal Information Extraction (IE), and propose Event-Centric Multimodal Knowledge Acquisition to evolve traditional Entity-centric Single-modality knowledge into Event-centric Multi-modality knowledge. Traditional entity-centric approaches to consuming multimodal information focus on concrete concepts (such as objects, object types, physical relations, e.g., a person in a car), while this dissertation endows machines to understand complex abstract semantic structures that are difficult to ground into image regions but are essential knowledge (such as events and semantic roles of objects, e.g., driver, passenger, passerby, salesperson). It is able to consolidate complex semantic structures of multiple modalities, providing a major benefit over recent research advances in single-modality (text-only or vision-only) knowledge. Such a transformation poses significant challenges in terms of understanding multimodal semantic structures (such as semantic roles) and temporal dynamics (such as future participants and their roles): - Understanding Multimodal Semantic Structures to answer What happened?, Who?, Where?, and When? (Knowledge Extraction): Due to the structural nature and lack of anchoring in a specific image region, abstract semantic structures are difficult to synthesize between text and vision modalities through general large-scale pretraining. We introduce complex event semantic structures into vision-language pretraining (CLIP-Event), and propose a zero-shot cross-modal transfer of semantic understanding abilities from language to vision, which resolves the poor portability issue of IE and supports Zero-shot Multimodal Event Extraction (M2E2) for the first time. We also release an open-source Multimodal IE system GAIA to serve as an off-the-shelf tool for the research community. - Understanding Temporal Dynamics to answer What will happen next?, Who will participant? and Why? (Knowledge Reasoning): The significance of capturing temporal dynamics has led to recent advances in script knowledge learning, however, which has been overly simplified to be local and sequential. We propose Event Graph Schema, which open doors to a global event graph context to enable alternative predictions, along with structural justifications including location-, attribute-, and participant-specific details. - Generating truthfully with Event-Centric Knowledge Facts (Knowledge Driven Applications): Our work has shown positive results on long-standing open problems, such as Timeline Summarization, Meeting Summarization, and Multimedia News Question Answering, Report Generation, etc. This work on Multimedia Event Knowledge Graphs aims to open doors to the next generation of information access, in order to equip machines with factual knowledge discovery and reasoning from diverse sources of information, so that we can lay a foundation for promoting factuality and truthfulness in information access, through a structured knowledge view that is easily explainable, highly compositional, and capable of long-horizon reasoning."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Event-centric multimodal knowledge acquisition"]}]}],"canonical_facts":{"dc:contributor":["Ji, Heng","Han, Jiawei","Zhai, Chengxiang","Chang, Shih-Fu","Cho, Kyunghyun"],"dc:creator":["Li, Manling"],"dc:date":["2023-08","2023-07-10"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2023-12-04 without embargo terms","The student, Manling Li, accepted the attached license on 2023-07-07 at 10:13.","The student, Manling Li, submitted this Dissertation for approval on 2023-07-07 at 10:41.","This Dissertation was approved for publication on 2023-07-10 at 08:57.","DSpace SAF Submission Ingestion Package generated from Vireo submission #19467 on 2023-12-04 at 17:00:14","What happened? Who? When? Where? Why? What will happen next? are the fundamental questions asked to comprehend the overwhelming amount of information. Answers to these questions are the core knowledge communicated through multiple forms of information, regardless of whether presented as text, images, videos, audio, or other modalities. To obtain such knowledge from multimodal data, this dissertation focuses on Multimodal Information Extraction (IE), and propose Event-Centric Multimodal Knowledge Acquisition to evolve traditional Entity-centric Single-modality knowledge into Event-centric Multi-modality knowledge. Traditional entity-centric approaches to consuming multimodal information focus on concrete concepts (such as objects, object types, physical relations, e.g., a person in a car), while this dissertation endows machines to understand complex abstract semantic structures that are difficult to ground into image regions but are essential knowledge (such as events and semantic roles of objects, e.g., driver, passenger, passerby, salesperson). It is able to consolidate complex semantic structures of multiple modalities, providing a major benefit over recent research advances in single-modality (text-only or vision-only) knowledge. Such a transformation poses significant challenges in terms of understanding multimodal semantic structures (such as semantic roles) and temporal dynamics (such as future participants and their roles): - Understanding Multimodal Semantic Structures to answer What happened?, Who?, Where?, and When? (Knowledge Extraction): Due to the structural nature and lack of anchoring in a specific image region, abstract semantic structures are difficult to synthesize between text and vision modalities through general large-scale pretraining. We introduce complex event semantic structures into vision-language pretraining (CLIP-Event), and propose a zero-shot cross-modal transfer of semantic understanding abilities from language to vision, which resolves the poor portability issue of IE and supports Zero-shot Multimodal Event Extraction (M2E2) for the first time. We also release an open-source Multimodal IE system GAIA to serve as an off-the-shelf tool for the research community. - Understanding Temporal Dynamics to answer What will happen next?, Who will participant? and Why? (Knowledge Reasoning): The significance of capturing temporal dynamics has led to recent advances in script knowledge learning, however, which has been overly simplified to be local and sequential. We propose Event Graph Schema, which open doors to a global event graph context to enable alternative predictions, along with structural justifications including location-, attribute-, and participant-specific details. - Generating truthfully with Event-Centric Knowledge Facts (Knowledge Driven Applications): Our work has shown positive results on long-standing open problems, such as Timeline Summarization, Meeting Summarization, and Multimedia News Question Answering, Report Generation, etc. This work on Multimedia Event Knowledge Graphs aims to open doors to the next generation of information access, in order to equip machines with factual knowledge discovery and reasoning from diverse sources of information, so that we can lay a foundation for promoting factuality and truthfulness in information access, through a structured knowledge view that is easily explainable, highly compositional, and capable of long-horizon reasoning."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/121435"],"dc:language":["en","eng"],"dc:rights":["Copyright 2023 Manling Li"],"dc:subject":["Multimodal","Knowledge","Event-centric"],"dc:title":["Event-centric multimodal knowledge acquisition"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:57Z"}