{"id":{"repo_id":"vt","oai_identifier":"oai:vtechworks.lib.vt.edu:10919/141065"},"canonical_url":"https://search.dev.ndltd.org/etd/vt/oai:vtechworks.lib.vt.edu:10919/141065","repository":{"repo_id":"vt","name":"Virginia Tech","base_url":"https://vtechworks.lib.vt.edu/oai/request"},"display":{"title":"Scholarly Information System for Long Documents and their Elements: Structured Representation and Exploration using Knowledge Graphs and LLMs","abstract":"Every year, thousands of students write long and detailed theses and dissertations as part of their graduate requirement. These documents contain valuable knowledge—methods, explanations, figures, and insights—but because they are so long, it is often difficult for researchers to find exactly what they are looking for. Most search tools only point researchers to the beginning of a document, leaving them to scroll through hundreds of pages on their own. Even modern artificial intelligence (AI) tools, like large language models, that can give fluent answers, don't always know where the information came from or how the different parts of a thesis fit together. This dissertation creates a new way to explore these long academic documents. Instead of treating a thesis as one giant block of text, the system breaks it into meaningful pieces like captions, figures, and key sentences, so they can be searched directly. It also builds a structured map, called a knowledge graph, that shows how these pieces relate to each other. This allows the system to understand the document more like a human would. Using this structure, the system can answer questions in a more helpful way. Rather than giving a list of documents for the user to read, it gathers relevant information and produces clear, paragraph-level answers that are grounded in specific parts of the dissertations. To see how well this works, 47 participants from different engineering fields tested three versions of the system. The version that combined the knowledge graph aided retrieval with AI-generated summaries received far better ratings for clarity, informativeness, and overall usefulness. Users said the answers were easier to read, more complete, and more helpful than those produced by traditional search or AI alone. Overall, this dissertation shows that by combining search technology, AI, and structured representations of knowledge, we can make long academic documents far more accessible. This approach helps people find what they need faster, understand it more easily, and discover information that would otherwise be buried deep within hundreds of pages. It represents a meaningful step toward smarter, more helpful tools for navigating scholarly knowledge.","abstract_html":"Every year, thousands of students write long and detailed theses and dissertations as part of their graduate requirement. These documents contain valuable knowledge—methods, explanations, figures, and insights—but because they are so long, it is often difficult for researchers to find exactly what they are looking for. Most search tools only point researchers to the beginning of a document, leaving them to scroll through hundreds of pages on their own. Even modern artificial intelligence (AI) tools, like large language models, that can give fluent answers, don&#x27;t always know where the information came from or how the different parts of a thesis fit together. This dissertation creates a new way to explore these long academic documents. Instead of treating a thesis as one giant block of text, the system breaks it into meaningful pieces like captions, figures, and key sentences, so they can be searched directly. It also builds a structured map, called a knowledge graph, that shows how these pieces relate to each other. This allows the system to understand the document more like a human would. Using this structure, the system can answer questions in a more helpful way. Rather than giving a list of documents for the user to read, it gathers relevant information and produces clear, paragraph-level answers that are grounded in specific parts of the dissertations. To see how well this works, 47 participants from different engineering fields tested three versions of the system. The version that combined the knowledge graph aided retrieval with AI-generated summaries received far better ratings for clarity, informativeness, and overall usefulness. Users said the answers were easier to read, more complete, and more helpful than those produced by traditional search or AI alone. Overall, this dissertation shows that by combining search technology, AI, and structured representations of knowledge, we can make long academic documents far more accessible. This approach helps people find what they need faster, understand it more easily, and discover information that would otherwise be buried deep within hundreds of pages. It represents a meaningful step toward smarter, more helpful tools for navigating scholarly knowledge.","abstract_has_math":false,"creators":["Chekuri, Satvik"],"institution":"Virginia Tech","degree_name":"Doctor of Philosophy","degree_level":"doctoral","degree_discipline":"Computer Science & Applications","degree_department":"Computer Science and Applications","school":null,"contributors":[],"advisors":[],"committee_chairs":["Fox, Edward A."],"committee_members":["Wu, Jian","Rho, Ha Rim","Lourentzou, Ismini","North, Christopher L."],"year":2026,"date_issued":"2026-01-29","date_published":"2026-01-29","updated_at":"2026-07-22T22:20:17Z","subjects":["Knowledge Graph","Question-Answering","Information Retrieval","Large Language Models","Natural Language Processing","Sentence Classification"],"languages":["en"],"rights":["Creative Commons Attribution 4.0 International"],"rights_urls":["http://creativecommons.org/licenses/by/4.0/"],"identifier_entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:45605"],"render_values":[{"text":"vt_gsexam:45605","href":null,"code":true}]}]},"links":{"outbound_url":"https://hdl.handle.net/10919/141065","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.committeechair","label":"Committee Chair","values":["Fox, Edward A."]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Wu, Jian","Rho, Ha Rim","Lourentzou, Ismini","North, Christopher L."]},{"key":"dc:contributor.department","label":"Department","values":["Computer Science and Applications"]},{"key":"dc:creator","label":"Author","values":["Chekuri, Satvik"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2026-01-30T09:00:48Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2026-01-30T09:00:48Z"]},{"key":"dc:date.issued","label":"Date","values":["2026-01-29"]},{"key":"dc:publisher","label":"Institution","values":["Virginia Tech"]},{"key":"dc:type","label":"Dc Type","values":["Dissertation"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science & Applications"]},{"key":"thesis:degree_level","label":"Degree Level","values":["doctoral"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctor of Philosophy"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Virginia Polytechnic Institute and State University"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Knowledge Graph","Question-Answering","Information Retrieval","Large Language Models","Natural Language Processing","Sentence Classification"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Creative Commons Attribution 4.0 International"]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://creativecommons.org/licenses/by/4.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:45605"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10919/141065"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstractgeneral","label":"General Abstract","values":["Every year, thousands of students write long and detailed theses and dissertations as part of their graduate requirement. These documents contain valuable knowledge—methods, explanations, figures, and insights—but because they are so long, it is often difficult for researchers to find exactly what they are looking for. Most search tools only point researchers to the beginning of a document, leaving them to scroll through hundreds of pages on their own. Even modern artificial intelligence (AI) tools, like large language models, that can give fluent answers, don't always know where the information came from or how the different parts of a thesis fit together. This dissertation creates a new way to explore these long academic documents. Instead of treating a thesis as one giant block of text, the system breaks it into meaningful pieces like captions, figures, and key sentences, so they can be searched directly. It also builds a structured map, called a knowledge graph, that shows how these pieces relate to each other. This allows the system to understand the document more like a human would. Using this structure, the system can answer questions in a more helpful way. Rather than giving a list of documents for the user to read, it gathers relevant information and produces clear, paragraph-level answers that are grounded in specific parts of the dissertations. To see how well this works, 47 participants from different engineering fields tested three versions of the system. The version that combined the knowledge graph aided retrieval with AI-generated summaries received far better ratings for clarity, informativeness, and overall usefulness. Users said the answers were easier to read, more complete, and more helpful than those produced by traditional search or AI alone. Overall, this dissertation shows that by combining search technology, AI, and structured representations of knowledge, we can make long academic documents far more accessible. This approach helps people find what they need faster, understand it more easily, and discover information that would otherwise be buried deep within hundreds of pages. It represents a meaningful step toward smarter, more helpful tools for navigating scholarly knowledge."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["Doctor of Philosophy"]},{"key":"dc:format.medium","label":"Dc Format Medium","values":["ETD"]},{"key":"dc:title","label":"Title","values":["Scholarly Information System for Long Documents and their Elements: Structured Representation and Exploration using Knowledge Graphs and LLMs"]}]}],"canonical_facts":{"dc:contributor.committeechair":["Fox, Edward A."],"dc:contributor.committeemember":["Wu, Jian","Rho, Ha Rim","Lourentzou, Ismini","North, Christopher L."],"dc:contributor.department":["Computer Science and Applications"],"dc:creator":["Chekuri, Satvik"],"dc:date.accessioned":["2026-01-30T09:00:48Z"],"dc:date.available":["2026-01-30T09:00:48Z"],"dc:date.issued":["2026-01-29"],"dc:description.abstractgeneral":["Every year, thousands of students write long and detailed theses and dissertations as part of their graduate requirement. These documents contain valuable knowledge—methods, explanations, figures, and insights—but because they are so long, it is often difficult for researchers to find exactly what they are looking for. Most search tools only point researchers to the beginning of a document, leaving them to scroll through hundreds of pages on their own. Even modern artificial intelligence (AI) tools, like large language models, that can give fluent answers, don't always know where the information came from or how the different parts of a thesis fit together. This dissertation creates a new way to explore these long academic documents. Instead of treating a thesis as one giant block of text, the system breaks it into meaningful pieces like captions, figures, and key sentences, so they can be searched directly. It also builds a structured map, called a knowledge graph, that shows how these pieces relate to each other. This allows the system to understand the document more like a human would. Using this structure, the system can answer questions in a more helpful way. Rather than giving a list of documents for the user to read, it gathers relevant information and produces clear, paragraph-level answers that are grounded in specific parts of the dissertations. To see how well this works, 47 participants from different engineering fields tested three versions of the system. The version that combined the knowledge graph aided retrieval with AI-generated summaries received far better ratings for clarity, informativeness, and overall usefulness. Users said the answers were easier to read, more complete, and more helpful than those produced by traditional search or AI alone. Overall, this dissertation shows that by combining search technology, AI, and structured representations of knowledge, we can make long academic documents far more accessible. This approach helps people find what they need faster, understand it more easily, and discover information that would otherwise be buried deep within hundreds of pages. It represents a meaningful step toward smarter, more helpful tools for navigating scholarly knowledge."],"dc:description.degree":["Doctor of Philosophy"],"dc:format.medium":["ETD"],"dc:identifier.other":["vt_gsexam:45605"],"dc:identifier.uri":["https://hdl.handle.net/10919/141065"],"dc:language.iso":["en"],"dc:publisher":["Virginia Tech"],"dc:rights":["Creative Commons Attribution 4.0 International"],"dc:rights.uri":["http://creativecommons.org/licenses/by/4.0/"],"dc:subject":["Knowledge Graph","Question-Answering","Information Retrieval","Large Language Models","Natural Language Processing","Sentence Classification"],"dc:title":["Scholarly Information System for Long Documents and their Elements: Structured Representation and Exploration using Knowledge Graphs and LLMs"],"dc:type":["Dissertation"],"thesis:degree_discipline":["Computer Science & Applications"],"thesis:degree_level":["doctoral"],"thesis:degree_name":["Doctor of Philosophy"],"thesis:institution_name":["Virginia Polytechnic Institute and State University"]},"updated_at":"2026-07-22T22:20:17Z"}