Back to results

NJIT

Concept graphs: Applications to biomedical text categorization and concept extraction

Abstract

dc:description.abstract

As science advances, the underlying literature grows rapidly providing valuable knowledge mines for researchers and practitioners. The text content that makes up these knowledge collections is often unstructured and, thus, extracting relevant or novel information could be nontrivial and costly. In addition, human knowledge and expertise are being transformed into structured digital information in the form of vocabulary databases and ontologies. These knowledge bases hold substantial hierarchical and semantic relationships of common domain concepts. Consequently, automating learning tasks could be reinforced with those knowledge bases through constructing human-like representations of knowledge. This allows developing algorithms that simulate the human reasoning tasks of content perception, concept identification, and classification. This study explores the representation of text documents using concept graphs that are constructed with the help of a domain ontology. In particular, the target data sets are collections of biomedical text documents, and the domain ontology is a collection of predefined biomedical concepts and relationships among them. The proposed representation preserves those relationships and allows using the structural features of graphs in text mining and learning algorithms. Those features emphasize the significance of the underlying relationship information that exists in the text content behind the interrelated topics and concepts of a text document. The experiments presented in this study include text categorization and concept extraction applied on biomedical data sets. The experimental results demonstrate how the relationships extracted from text and captured in graph structures can be used to improve the performance of the aforementioned applications. The discussed techniques can be used in creating and maintaining digital libraries through enhancing indexing, retrieval, and management of documents as well as in a broad range of domain-specific applications such as drug discovery, hypothesis generation, and the analysis of molecular structures in chemoinformatics.

Degree

thesis:*
Name thesis:degree_name
Doctor of Philosophy in Information Systems - (Ph.D.)
Discipline thesis:degree_discipline
Information Systems
Year
2013

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Bleik, Said
Contributors dc:contributor
  • Min Song
  • Fadi P. Deek
  • James Geller

Subjects

dc:subject × 8

Identifiers

dc:identifier.*
Repository record dc:identifier
https://digitalcommons.njit.edu/dissertations/360
OAI identifier oai:identifier
oai:digitalcommons.njit.edu:dissertations-1415

Chain of custody

source
Harvested from
NJIT
Base URL
digitalcommons.njit.edu/do/oai/
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Bleik, Said. Concept graphs: Applications to biomedical text categorization and concept extraction. 2013. https://digitalcommons.njit.edu/dissertations/360