{"id":{"repo_id":"temple","oai_identifier":"oai:scholarshare.temple.edu:20.500.12613/12237"},"canonical_url":"https://search.dev.ndltd.org/etd/temple/oai:scholarshare.temple.edu:20.500.12613/12237","repository":{"repo_id":"temple","name":"Temple University","base_url":"https://scholarshare.temple.edu/server/oai/request"},"display":{"title":"Scalable information extraction with large language models","abstract":"Information extraction (IE) transforms unstructured text into structured knowledge such as entities and relations, and is fundamental to applications including knowledge graph construction, information retrieval, question answering, and domain-specific document understanding. Although large language models (LLMs) have broadened the scope of IE through zero-shot and in-context extraction, scalable IE remains challenging in realistic settings, particularly for scientific and other specialized domains where labeled data is scarce, expensive, and difficult to curate. This dissertation studies scalable information extraction with large language models from a data-centric perspective, arguing that progress requires advances in benchmark construction, noise-aware supervision, and LLM-based annotation. This dissertation makes three main contributions. First, it introduces a full-text benchmark for scientific information extraction that supports the extraction of datasets, methods, tasks, and their relations from scientific publications. The benchmark contains 106 manually annotated full-text papers with more than 24,000 entity mentions and 12,000 relations, providing a more realistic testbed than prior resources limited to abstracts or selected paragraphs. Second, it proposes DynClean, a training dynamics-based label cleaning framework for distantly supervised named entity recognition. By locating false positive and false negative annotations in weakly supervised data, DynClean improves downstream F1 by 3.19% to 8.95% across four benchmark datasets and outperforms prior distantly supervised NER methods by up to 4.53 F1 points. Third, it presents a systematic study of many-shot in-context learning for named entity recognition and develops an in-context annotation framework for low-resource settings. The results show that using around 100 human-labeled examples, LLMs can generate high-quality labeled corpora for training smaller models, yielding improvements of approximately 10 absolute F1 points over strong baselines in low-resource domain-specific NER. Taken together, these contributions support a unified view of scalable information extraction in which realistic benchmarks define meaningful tasks, noise-aware supervision improves the quality of automatically generated labels, and LLMs act as annotation engines that amplify limited human supervision. More broadly, this dissertation shows that scalable IE is not only a modeling problem, but also a data-centric systems problem requiring the joint design of resources, supervision, and adaptation mechanisms.","abstract_html":"Information extraction (IE) transforms unstructured text into structured knowledge such as entities and relations, and is fundamental to applications including knowledge graph construction, information retrieval, question answering, and domain-specific document understanding. Although large language models (LLMs) have broadened the scope of IE through zero-shot and in-context extraction, scalable IE remains challenging in realistic settings, particularly for scientific and other specialized domains where labeled data is scarce, expensive, and difficult to curate. This dissertation studies scalable information extraction with large language models from a data-centric perspective, arguing that progress requires advances in benchmark construction, noise-aware supervision, and LLM-based annotation. This dissertation makes three main contributions. First, it introduces a full-text benchmark for scientific information extraction that supports the extraction of datasets, methods, tasks, and their relations from scientific publications. The benchmark contains 106 manually annotated full-text papers with more than 24,000 entity mentions and 12,000 relations, providing a more realistic testbed than prior resources limited to abstracts or selected paragraphs. Second, it proposes DynClean, a training dynamics-based label cleaning framework for distantly supervised named entity recognition. By locating false positive and false negative annotations in weakly supervised data, DynClean improves downstream F1 by 3.19% to 8.95% across four benchmark datasets and outperforms prior distantly supervised NER methods by up to 4.53 F1 points. Third, it presents a systematic study of many-shot in-context learning for named entity recognition and develops an in-context annotation framework for low-resource settings. The results show that using around 100 human-labeled examples, LLMs can generate high-quality labeled corpora for training smaller models, yielding improvements of approximately 10 absolute F1 points over strong baselines in low-resource domain-specific NER. Taken together, these contributions support a unified view of scalable information extraction in which realistic benchmarks define meaningful tasks, noise-aware supervision improves the quality of automatically generated labels, and LLMs act as annotation engines that amplify limited human supervision. More broadly, this dissertation shows that scalable IE is not only a modeling problem, but also a data-centric systems problem requiring the joint design of resources, supervision, and adaptation mechanisms.","abstract_has_math":false,"creators":["Zhang, Qi"],"institution":"Temple University. Libraries","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Dragut, Eduard Constantin"],"committee_chairs":[],"committee_members":["Latecki, Longin","MacNeil, Stephen, 1987-","Caragea, Cornelia"],"year":2026,"date_issued":"2026-05","date_published":"2026-05","updated_at":"2026-07-27T21:20:27Z","subjects":["Computer science","Agentic AI","In-context learning","Information extraction","Large language model","Named entity recognition","Relation extraction","Computer and information science"],"languages":["eng"],"rights":["IN COPYRIGHT- This Rights Statement can be used for an Item that is in copyright. Using this statement implies that the organization making this Item available has determined that the Item is in copyright and either is the rights-holder, has obtained permission from the rights-holder(s) to make their Work(s) available, or makes the Item available under an exception or limitation to copyright (including Fair Use) that entitles it to make the Item available."],"rights_urls":["http://rightsstatements.org/vocab/InC/1.0/"],"identifier_entries":[]},"links":{"outbound_url":"https://scholarshare.temple.edu/handle/20.500.12613/12237","outbound_label":"Repository record","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Dragut, Eduard Constantin"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Latecki, Longin","MacNeil, Stephen, 1987-","Caragea, Cornelia"]},{"key":"dc:creator","label":"Author","values":["Zhang, Qi"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2026-06-10T14:54:08Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2026-06-10T14:54:08Z"]},{"key":"dc:date.issued","label":"Date","values":["2026-05"]},{"key":"dc:publisher","label":"Institution","values":["Temple University. Libraries"]},{"key":"dc:type","label":"Dc Type","values":["Text"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer science","Agentic AI","In-context learning","Information extraction","Large language model","Named entity recognition","Relation extraction","Computer and information science"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["IN COPYRIGHT- This Rights Statement can be used for an Item that is in copyright. Using this statement implies that the organization making this Item available has determined that the Item is in copyright and either is the rights-holder, has obtained permission from the rights-holder(s) to make their Work(s) available, or makes the Item available under an exception or limitation to copyright (including Fair Use) that entitles it to make the Item available."]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://rightsstatements.org/vocab/InC/1.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://scholarshare.temple.edu/handle/20.500.12613/12237"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Information extraction (IE) transforms unstructured text into structured knowledge such as entities and relations, and is fundamental to applications including knowledge graph construction, information retrieval, question answering, and domain-specific document understanding. Although large language models (LLMs) have broadened the scope of IE through zero-shot and in-context extraction, scalable IE remains challenging in realistic settings, particularly for scientific and other specialized domains where labeled data is scarce, expensive, and difficult to curate. This dissertation studies scalable information extraction with large language models from a data-centric perspective, arguing that progress requires advances in benchmark construction, noise-aware supervision, and LLM-based annotation. This dissertation makes three main contributions. First, it introduces a full-text benchmark for scientific information extraction that supports the extraction of datasets, methods, tasks, and their relations from scientific publications. The benchmark contains 106 manually annotated full-text papers with more than 24,000 entity mentions and 12,000 relations, providing a more realistic testbed than prior resources limited to abstracts or selected paragraphs. Second, it proposes DynClean, a training dynamics-based label cleaning framework for distantly supervised named entity recognition. By locating false positive and false negative annotations in weakly supervised data, DynClean improves downstream F1 by 3.19% to 8.95% across four benchmark datasets and outperforms prior distantly supervised NER methods by up to 4.53 F1 points. Third, it presents a systematic study of many-shot in-context learning for named entity recognition and develops an in-context annotation framework for low-resource settings. The results show that using around 100 human-labeled examples, LLMs can generate high-quality labeled corpora for training smaller models, yielding improvements of approximately 10 absolute F1 points over strong baselines in low-resource domain-specific NER. Taken together, these contributions support a unified view of scalable information extraction in which realistic benchmarks define meaningful tasks, noise-aware supervision improves the quality of automatically generated labels, and LLMs act as annotation engines that amplify limited human supervision. More broadly, this dissertation shows that scalable IE is not only a modeling problem, but also a data-centric systems problem requiring the joint design of resources, supervision, and adaptation mechanisms."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["Ph.D."]},{"key":"dc:title","label":"Title","values":["Scalable information extraction with large language models"]}]}],"canonical_facts":{"dc:contributor.advisor":["Dragut, Eduard Constantin"],"dc:contributor.committeemember":["Latecki, Longin","MacNeil, Stephen, 1987-","Caragea, Cornelia"],"dc:creator":["Zhang, Qi"],"dc:date.accessioned":["2026-06-10T14:54:08Z"],"dc:date.available":["2026-06-10T14:54:08Z"],"dc:date.issued":["2026-05"],"dc:description.abstract":["Information extraction (IE) transforms unstructured text into structured knowledge such as entities and relations, and is fundamental to applications including knowledge graph construction, information retrieval, question answering, and domain-specific document understanding. Although large language models (LLMs) have broadened the scope of IE through zero-shot and in-context extraction, scalable IE remains challenging in realistic settings, particularly for scientific and other specialized domains where labeled data is scarce, expensive, and difficult to curate. This dissertation studies scalable information extraction with large language models from a data-centric perspective, arguing that progress requires advances in benchmark construction, noise-aware supervision, and LLM-based annotation. This dissertation makes three main contributions. First, it introduces a full-text benchmark for scientific information extraction that supports the extraction of datasets, methods, tasks, and their relations from scientific publications. The benchmark contains 106 manually annotated full-text papers with more than 24,000 entity mentions and 12,000 relations, providing a more realistic testbed than prior resources limited to abstracts or selected paragraphs. Second, it proposes DynClean, a training dynamics-based label cleaning framework for distantly supervised named entity recognition. By locating false positive and false negative annotations in weakly supervised data, DynClean improves downstream F1 by 3.19% to 8.95% across four benchmark datasets and outperforms prior distantly supervised NER methods by up to 4.53 F1 points. Third, it presents a systematic study of many-shot in-context learning for named entity recognition and develops an in-context annotation framework for low-resource settings. The results show that using around 100 human-labeled examples, LLMs can generate high-quality labeled corpora for training smaller models, yielding improvements of approximately 10 absolute F1 points over strong baselines in low-resource domain-specific NER. Taken together, these contributions support a unified view of scalable information extraction in which realistic benchmarks define meaningful tasks, noise-aware supervision improves the quality of automatically generated labels, and LLMs act as annotation engines that amplify limited human supervision. More broadly, this dissertation shows that scalable IE is not only a modeling problem, but also a data-centric systems problem requiring the joint design of resources, supervision, and adaptation mechanisms."],"dc:description.degree":["Ph.D."],"dc:identifier.uri":["https://scholarshare.temple.edu/handle/20.500.12613/12237"],"dc:language.iso":["eng"],"dc:publisher":["Temple University. Libraries"],"dc:rights":["IN COPYRIGHT- This Rights Statement can be used for an Item that is in copyright. Using this statement implies that the organization making this Item available has determined that the Item is in copyright and either is the rights-holder, has obtained permission from the rights-holder(s) to make their Work(s) available, or makes the Item available under an exception or limitation to copyright (including Fair Use) that entitles it to make the Item available."],"dc:rights.uri":["http://rightsstatements.org/vocab/InC/1.0/"],"dc:subject":["Computer science","Agentic AI","In-context learning","Information extraction","Large language model","Named entity recognition","Relation extraction","Computer and information science"],"dc:title":["Scalable information extraction with large language models"],"dc:type":["Text"]},"updated_at":"2026-07-27T21:20:27Z"}