{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/49376"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/49376","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Mining novel sources of knowledge to identify causal information in text","abstract":"The abundance of information on the internet has impacted the lives of people to a great extent. People take advantage of the internet to acquire information for several day to day social and political activities. Though the plenty of information on the internet is of great use, it takes lot of time to go through a number of text articles to understand events and the causal relations between events that build a particular social or political news story. In this thesis, we focus on the problem of automated extraction of causal information in text. This can be of great assistance to the people who strive to acquire the flow of events in text to make various decisions and predict consequences of their decisions. In natural language, causal relations can be encoded using various linguistic constructions. Each construction with its own semantics can pose various challenges for the problem of identifying causality. In this thesis, we address the tasks of identifying causality between two verbs and a verb and a noun by deeply analyzing semantics of these constructions. After the successful use of linguistic features for various Natural Language Processing (NLP) tasks, several approaches have been proposed to identify causality using such features in the framework of supervised learning. However, it is not practical to depend merely on these features because there are many factors involved in identifying causality such as background knowledge, semantic and pragmatic features of events, world knowledge, etc. In addition to the above, the supervised learning approaches are sensitive to the size of training corpus and the type of contexts of training instances. For example, the unambiguous training instances do not provide a better supervision for the ambiguous and implicit instances of semantic relations including causality [Sporleder and Lascarides 2008]. Therefore, in this work instead of merely relying on the linguistic features extracted from the contexts of training instances, we propose an approach to derive novel sources of knowledge for identifying causal information in text. In the first part of this thesis, we introduce methods to acquire background knowledge and the knowledge of causal semantics of verbs for the task of identifying causality between the two state of affairs represented by verbs. After the knowledge acquisition step, we integrate the above types of knowledge with a supervised classifier employing linguistic features to obtain optimal predictions for the current task. Similarly, in the second part of this thesis, we propose methods to acquire and employ the knowledge of causal semantics of nouns, verbs and verb frames to identify causality between the two state of affairs represented by verbs and nouns. With the addition of novel sources of knowledge, our models for the current tasks gain lots of progress in performance over the baseline of supervised classifiers relying merely on linguistics features. Moreover, in comparison with these supervised classifiers, performance of our models is more robust on all types of context - i.e., unambiguous, ambiguous and implicit contexts.","abstract_html":"The abundance of information on the internet has impacted the lives of people to a great extent. People take advantage of the internet to acquire information for several day to day social and political activities. Though the plenty of information on the internet is of great use, it takes lot of time to go through a number of text articles to understand events and the causal relations between events that build a particular social or political news story. In this thesis, we focus on the problem of automated extraction of causal information in text. This can be of great assistance to the people who strive to acquire the flow of events in text to make various decisions and predict consequences of their decisions. In natural language, causal relations can be encoded using various linguistic constructions. Each construction with its own semantics can pose various challenges for the problem of identifying causality. In this thesis, we address the tasks of identifying causality between two verbs and a verb and a noun by deeply analyzing semantics of these constructions. After the successful use of linguistic features for various Natural Language Processing (NLP) tasks, several approaches have been proposed to identify causality using such features in the framework of supervised learning. However, it is not practical to depend merely on these features because there are many factors involved in identifying causality such as background knowledge, semantic and pragmatic features of events, world knowledge, etc. In addition to the above, the supervised learning approaches are sensitive to the size of training corpus and the type of contexts of training instances. For example, the unambiguous training instances do not provide a better supervision for the ambiguous and implicit instances of semantic relations including causality [Sporleder and Lascarides 2008]. Therefore, in this work instead of merely relying on the linguistic features extracted from the contexts of training instances, we propose an approach to derive novel sources of knowledge for identifying causal information in text. In the first part of this thesis, we introduce methods to acquire background knowledge and the knowledge of causal semantics of verbs for the task of identifying causality between the two state of affairs represented by verbs. After the knowledge acquisition step, we integrate the above types of knowledge with a supervised classifier employing linguistic features to obtain optimal predictions for the current task. Similarly, in the second part of this thesis, we propose methods to acquire and employ the knowledge of causal semantics of nouns, verbs and verb frames to identify causality between the two state of affairs represented by verbs and nouns. With the addition of novel sources of knowledge, our models for the current tasks gain lots of progress in performance over the baseline of supervised classifiers relying merely on linguistics features. Moreover, in comparison with these supervised classifiers, performance of our models is more robust on all types of context - i.e., unambiguous, ambiguous and implicit contexts.","abstract_has_math":false,"creators":["Riaz, Mehwish"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Girju, Roxana","Zhai, ChengXiang","Hockenmaier, Julia C.","Di Eugenio, Barbara"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2014,"date_issued":"2014-05-30T16:40:52Z","date_published":"2014-05-30T16:40:52Z","updated_at":"2026-07-22T22:25:38Z","subjects":["Causality","Events","Discourse Processing","Natural Language Semantics"],"languages":["en"],"rights":["2014 by Mehwish Riaz"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/49376","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Girju, Roxana","Zhai, ChengXiang","Hockenmaier, Julia C.","Di Eugenio, Barbara"]},{"key":"dc:creator","label":"Author","values":["Riaz, Mehwish"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2014-05-30T16:40:52Z","2014-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Causality","Events","Discourse Processing","Natural Language Semantics"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["2014 by Mehwish Riaz"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/49376"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The abundance of information on the internet has impacted the lives of people to a great extent. People take advantage of the internet to acquire information for several day to day social and political activities. Though the plenty of information on the internet is of great use, it takes lot of time to go through a number of text articles to understand events and the causal relations between events that build a particular social or political news story. In this thesis, we focus on the problem of automated extraction of causal information in text. This can be of great assistance to the people who strive to acquire the flow of events in text to make various decisions and predict consequences of their decisions. In natural language, causal relations can be encoded using various linguistic constructions. Each construction with its own semantics can pose various challenges for the problem of identifying causality. In this thesis, we address the tasks of identifying causality between two verbs and a verb and a noun by deeply analyzing semantics of these constructions. After the successful use of linguistic features for various Natural Language Processing (NLP) tasks, several approaches have been proposed to identify causality using such features in the framework of supervised learning. However, it is not practical to depend merely on these features because there are many factors involved in identifying causality such as background knowledge, semantic and pragmatic features of events, world knowledge, etc. In addition to the above, the supervised learning approaches are sensitive to the size of training corpus and the type of contexts of training instances. For example, the unambiguous training instances do not provide a better supervision for the ambiguous and implicit instances of semantic relations including causality [Sporleder and Lascarides 2008]. Therefore, in this work instead of merely relying on the linguistic features extracted from the contexts of training instances, we propose an approach to derive novel sources of knowledge for identifying causal information in text. In the first part of this thesis, we introduce methods to acquire background knowledge and the knowledge of causal semantics of verbs for the task of identifying causality between the two state of affairs represented by verbs. After the knowledge acquisition step, we integrate the above types of knowledge with a supervised classifier employing linguistic features to obtain optimal predictions for the current task. Similarly, in the second part of this thesis, we propose methods to acquire and employ the knowledge of causal semantics of nouns, verbs and verb frames to identify causality between the two state of affairs represented by verbs and nouns. With the addition of novel sources of knowledge, our models for the current tasks gain lots of progress in performance over the baseline of supervised classifiers relying merely on linguistics features. Moreover, in comparison with these supervised classifiers, performance of our models is more robust on all types of context - i.e., unambiguous, ambiguous and implicit contexts.","Item withdrawn by Alexis Thompson (athmpsn1@illinois.edu) on 2014-04-23T16:33:58Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Riaz_Mehwish.pdf: 15246857 bytes, checksum: 8cd372921b78dbeb591106512a43ae86 (MD5)","Made available in DSpace on 2014-05-30T16:40:52Z (GMT). No. of bitstreams: 2 Mehwish_Riaz.pdf: 15246857 bytes, checksum: 8cd372921b78dbeb591106512a43ae86 (MD5) license.txt: 4060 bytes, checksum: a09cc9c200e5b38280305ab78c125c68 (MD5)"]},{"key":"dc:title","label":"Title","values":["Mining novel sources of knowledge to identify causal information in text"]}]}],"canonical_facts":{"dc:contributor":["Girju, Roxana","Zhai, ChengXiang","Hockenmaier, Julia C.","Di Eugenio, Barbara"],"dc:creator":["Riaz, Mehwish"],"dc:date":["2014-05-30T16:40:52Z","2014-05"],"dc:description":["The abundance of information on the internet has impacted the lives of people to a great extent. People take advantage of the internet to acquire information for several day to day social and political activities. Though the plenty of information on the internet is of great use, it takes lot of time to go through a number of text articles to understand events and the causal relations between events that build a particular social or political news story. In this thesis, we focus on the problem of automated extraction of causal information in text. This can be of great assistance to the people who strive to acquire the flow of events in text to make various decisions and predict consequences of their decisions. In natural language, causal relations can be encoded using various linguistic constructions. Each construction with its own semantics can pose various challenges for the problem of identifying causality. In this thesis, we address the tasks of identifying causality between two verbs and a verb and a noun by deeply analyzing semantics of these constructions. After the successful use of linguistic features for various Natural Language Processing (NLP) tasks, several approaches have been proposed to identify causality using such features in the framework of supervised learning. However, it is not practical to depend merely on these features because there are many factors involved in identifying causality such as background knowledge, semantic and pragmatic features of events, world knowledge, etc. In addition to the above, the supervised learning approaches are sensitive to the size of training corpus and the type of contexts of training instances. For example, the unambiguous training instances do not provide a better supervision for the ambiguous and implicit instances of semantic relations including causality [Sporleder and Lascarides 2008]. Therefore, in this work instead of merely relying on the linguistic features extracted from the contexts of training instances, we propose an approach to derive novel sources of knowledge for identifying causal information in text. In the first part of this thesis, we introduce methods to acquire background knowledge and the knowledge of causal semantics of verbs for the task of identifying causality between the two state of affairs represented by verbs. After the knowledge acquisition step, we integrate the above types of knowledge with a supervised classifier employing linguistic features to obtain optimal predictions for the current task. Similarly, in the second part of this thesis, we propose methods to acquire and employ the knowledge of causal semantics of nouns, verbs and verb frames to identify causality between the two state of affairs represented by verbs and nouns. With the addition of novel sources of knowledge, our models for the current tasks gain lots of progress in performance over the baseline of supervised classifiers relying merely on linguistics features. Moreover, in comparison with these supervised classifiers, performance of our models is more robust on all types of context - i.e., unambiguous, ambiguous and implicit contexts.","Item withdrawn by Alexis Thompson (athmpsn1@illinois.edu) on 2014-04-23T16:33:58Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Riaz_Mehwish.pdf: 15246857 bytes, checksum: 8cd372921b78dbeb591106512a43ae86 (MD5)","Made available in DSpace on 2014-05-30T16:40:52Z (GMT). No. of bitstreams: 2 Mehwish_Riaz.pdf: 15246857 bytes, checksum: 8cd372921b78dbeb591106512a43ae86 (MD5) license.txt: 4060 bytes, checksum: a09cc9c200e5b38280305ab78c125c68 (MD5)"],"dc:identifier":["http://hdl.handle.net/2142/49376"],"dc:language":["en"],"dc:rights":["2014 by Mehwish Riaz"],"dc:subject":["Causality","Events","Discourse Processing","Natural Language Semantics"],"dc:title":["Mining novel sources of knowledge to identify causal information in text"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:38Z"}