University of Missouri -- Kansas City
Deep Assertion discovery using word embeddings
Abstract
dc:description.abstractIn recent years, there has been explosive growth in the amount of biomedical data (e.g., publications, notes from EHRs, clinical trial results), with the majority being unstructured data. As the volume of data is increasing faster, the demand for extracting knowledge from unstructured data increases in biomedical research. Many information extraction techniques like OpenIE, Ollie, ReVerb have been proposed to extract binary relations from plain text. The result of these techniques is triplets containing two argument phrases from the input text and a relation between these arguments which are a trade-off between relevance and completeness. These existing IE frameworks are unable to deal with big data. In this thesis, we propose an assertion framework that dynamically refines the triplets to form factual tuples called ‘assertions’ from multiple free-text sources. The framework is built using a pipeline approach to discover the significant features (words) and map them to the relation tuples extracted using the open information extraction technique. We have implemented three different feature discovery techniques – using the numerical statistic like TF-IDF (Term Frequency-Inverse Document Frequency), using word embedding approaches like Word2Vec in Deep Learning frameworks. We have shown that the assertions discovered using our proposed framework can be effectively used for the topic-based classification of the biomedical dataset containing abstracts for three categories such as obesity, heart disease and diabetes from NCBI PubMed and a benchmark dataset, 20 Newsgroups. The knowledge discovery from data in the proposed framework was verified through the classification using a five-layered Convolutional Neural Network. We show that the classification accuracy with assertions derived from the data improves the accuracy of the existing approach like deep belief networks by more than 10% on the benchmark dataset like 20 Newsgroups.
Degree
thesis:*- Name thesis:degree_name
- M.S.
- Level thesis:degree_level
- Masters
- Discipline thesis:degree_discipline
- Computer Science (UMKC)
- Grantor dc:publisher
- University of Missouri -- Kansas City
- Year dc:date.issued
- 2018
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Nagabhushan, Megha
- Advisor dc:contributor.advisor
-
- Lee, Yugyung, 1960-
Rights
- Language dc:language.iso
- en_US
Identifiers
dc:identifier.*- Handle dc:identifier.uri
- https://hdl.handle.net/10355/67047
- OAI identifier oai:identifier
- oai:mospace.umsystem.edu:10355/67047