Back to results

Temple University. Libraries

Information Extraction from Scientific Literature

Abstract

dc:description.abstract

The exponential growth of scientific literature, with millions of new articles published annually, has created an unsustainable discovery bottleneck across research communities. Manual extraction of critical information---including methodologies, datasets, and domain-specific terminologies---now consumes a substantial proportion of researchers' literature review time, particularly impacting time-sensitive fields like climate science and biomedical research where delayed insights hinder urgent policy decisions or therapeutic developments. Automated information extraction systems have transitioned from supplemental tools to essential infrastructure, addressing three critical imperatives: preserving collective understanding through cross-publication discovery linking, enabling real-time knowledge synthesis in rapidly evolving domains, and democratizing access to specialized findings via structured knowledge representation. Without robust frameworks, the scientific community risks perpetuating redundant investigations, overlooking critical interdisciplinary connections, and failing to transform publication volume into actionable insight networks. Current information extraction paradigms face four fundamental technical challenges rooted in scientific communication's unique characteristics. First, terminological instability arises from continuous conceptual evolution, where emerging constructs like ``attribution-based climate models'' and ``GPT-4.5'' outpace standardized taxonomies, generating persistent errors in entity disambiguation. Second, structural heterogeneity manifests through hundreds of distinct methodological description formats observed even within focused disciplines like materials science, complicating pattern generalization. Third, contextual dependency demands adaptive interpretation of concepts such as ``deep learning,'' whose technical meanings diverge fundamentally between protein folding architectures and geospatial mapping applications. Fourth, the scalability--accuracy tradeoff forces untenable compromises between precision (evidenced by frequent LLM hallucinations) and coverage (marked by traditional NLP's oversight of domain-specific abbreviations). These technical barriers compound with systemic data limitations---existing corpora cover only a small fraction of specialized domains while exhibiting annotation inconsistencies that undermine model reliability. Emerging paradigms in hierarchical relationship modeling hint at potential resolutions through hybrid neural-symbolic architectures. This research advances scientific information extraction through three interconnected contributions: the development of domain-annotated corpora spanning climate science and computer science; systematic evaluation of machine learning architectures across extraction tasks and disciplinary contexts; and demonstrated pathways for transforming extracted entities into evolvable knowledge graphs. By creating structured repositories that capture methodological lineages, dataset dependencies, and conceptual evolution patterns, our work provides researchers with interoperable frameworks for mapping relationships across fragmented scientific domains. The resulting infrastructure enables both precision-focused analysis within specialized fields and cross-domain knowledge discovery, offering scalable solutions to organize literature at scale while preserving disciplinary nuance. These contributions collectively address the dual challenges of maintaining taxonomic rigor and enabling adaptive knowledge synthesis in modern scientific communication ecosystems.

Degree

thesis:*
Grantor dc:publisher
Temple University. Libraries
Year dc:date.issued
2025

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Pan, Huitong
Advisor dc:contributor.advisor
  • Latecki, Longin
Committee members dc:contributor.committeemember
  • Latecki, Longin
  • Dragut, Eduard Constantin
  • Gao, Hongchang
  • Caragea, Cornelia

Subjects

dc:subject × 5

Rights

dc:rights
Statement dc:rights
  • IN COPYRIGHT- This Rights Statement can be used for an Item that is in copyright. Using this statement implies that the organization making this Item available has determined that the Item is in copyright and either is the rights-holder, has obtained permission from the rights-holder(s) to make their Work(s) available, or makes the Item available under an exception or limitation to copyright (including Fair Use) that entitles it to make the Item available.
Language dc:language.iso
eng

Identifiers

dc:identifier.*
Repository record dc:identifier.uri
https://scholarshare.temple.edu/handle/20.500.12613/11150
OAI identifier oai:identifier
oai:scholarshare.temple.edu:20.500.12613/11150

Chain of custody

source
Harvested from
Temple University
Base URL
scholarshare.temple.edu/server/oai/request
Last updated
2026-07-27
Source record
OAI-PMH GetRecord
citation

Pan, Huitong. Information Extraction from Scientific Literature. Temple University. Libraries, 2025. https://scholarshare.temple.edu/handle/20.500.12613/11150