{"id":{"repo_id":"texas-state","oai_identifier":"oai:digital.library.txst.edu:10877/21830"},"canonical_url":"https://search.dev.ndltd.org/etd/texas-state/oai:digital.library.txst.edu:10877/21830","repository":{"repo_id":"texas-state","name":"Texas State University","base_url":"https://digital.library.txst.edu/server/oai/request"},"display":{"title":"ResearchBuddy AI: LLM-Powered Assistant","abstract":"This paper presents ResearchBuddy AI, an LLM-powered assistant that addresses modern research information overload problems. The project has three essential features, including finetuning of large language models for structured research paper summarization, prompt engineering for flexible summarization of variable formats of GitHub READMEs, and Retrieval-Augmented Generation (RAG) for context-aware question answering over domain-specific documents. The research explains model selection through self-hosted open-source models with LoRA (Low Rank Adaptation) and other techniques to achieve cost-effectiveness and control. The research paper summarization capability was developed through an iterative process. I began with small model and dataset experiments before scaling to a larger model with expanded datasets. During optimization on powerful hardware, I identified a “token performance paradox” about long context utilization. The development process faced three main challenges, which involved dealing with different input formats and managing token limits, optimizing computing resources, and solutions are presented. The RAG chatbot architecture uses vector databases together with powerful LLMs to generate responses that remain grounded in the original content. The future research agenda includes two main directions: investigating new model architecture suitable for processing long documents and developing YouTube video summarization with timestamp capabilities. ResearchBuddy AI is designed to provide researchers with intelligent tools for more efficient information management, with early development indicating strong potential to address research information overload.","abstract_html":"This paper presents ResearchBuddy AI, an LLM-powered assistant that addresses modern research information overload problems. The project has three essential features, including finetuning of large language models for structured research paper summarization, prompt engineering for flexible summarization of variable formats of GitHub READMEs, and Retrieval-Augmented Generation (RAG) for context-aware question answering over domain-specific documents. The research explains model selection through self-hosted open-source models with LoRA (Low Rank Adaptation) and other techniques to achieve cost-effectiveness and control. The research paper summarization capability was developed through an iterative process. I began with small model and dataset experiments before scaling to a larger model with expanded datasets. During optimization on powerful hardware, I identified a “token performance paradox” about long context utilization. The development process faced three main challenges, which involved dealing with different input formats and managing token limits, optimizing computing resources, and solutions are presented. The RAG chatbot architecture uses vector databases together with powerful LLMs to generate responses that remain grounded in the original content. The future research agenda includes two main directions: investigating new model architecture suitable for processing long documents and developing YouTube video summarization with timestamp capabilities. ResearchBuddy AI is designed to provide researchers with intelligent tools for more efficient information management, with early development indicating strong potential to address research information overload.","abstract_has_math":false,"creators":["Qassim, Muhammad"],"institution":"Texas State University","degree_name":null,"degree_level":null,"degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":[],"advisors":["Lehr, Ted"],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-05","date_published":"2025-05","updated_at":"2026-07-27T21:22:32Z","subjects":["large language models","research assistants","summarization","retrieval-augmented generation","LoRA","prompt engineering","fine-tuning","RAG chatbot","vector databases","GitHub summarization"],"languages":["en"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/10877/21830","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Lehr, Ted"]},{"key":"dc:creator","label":"Author","values":["Qassim, Muhammad"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2025-08-12T15:05:42Z"]},{"key":"dc:date.issued","label":"Date","values":["2025-05"]},{"key":"dc:type","label":"Dc Type","values":["Capstone"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Texas State University"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["large language models","research assistants","summarization","retrieval-augmented generation","LoRA","prompt engineering","fine-tuning","RAG chatbot","vector databases","GitHub summarization"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10877/21830"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["This paper presents ResearchBuddy AI, an LLM-powered assistant that addresses modern research information overload problems. The project has three essential features, including finetuning of large language models for structured research paper summarization, prompt engineering for flexible summarization of variable formats of GitHub READMEs, and Retrieval-Augmented Generation (RAG) for context-aware question answering over domain-specific documents. The research explains model selection through self-hosted open-source models with LoRA (Low Rank Adaptation) and other techniques to achieve cost-effectiveness and control. The research paper summarization capability was developed through an iterative process. I began with small model and dataset experiments before scaling to a larger model with expanded datasets. During optimization on powerful hardware, I identified a “token performance paradox” about long context utilization. The development process faced three main challenges, which involved dealing with different input formats and managing token limits, optimizing computing resources, and solutions are presented. The RAG chatbot architecture uses vector databases together with powerful LLMs to generate responses that remain grounded in the original content. The future research agenda includes two main directions: investigating new model architecture suitable for processing long documents and developing YouTube video summarization with timestamp capabilities. ResearchBuddy AI is designed to provide researchers with intelligent tools for more efficient information management, with early development indicating strong potential to address research information overload."]},{"key":"dc:format","label":"Dc Format","values":["Text"]},{"key":"dc:format.medium","label":"Dc Format Medium","values":["1 file (.pdf)"]},{"key":"dc:title","label":"Title","values":["ResearchBuddy AI: LLM-Powered Assistant"]}]}],"canonical_facts":{"dc:contributor.advisor":["Lehr, Ted"],"dc:creator":["Qassim, Muhammad"],"dc:date.accessioned":["2025-08-12T15:05:42Z"],"dc:date.issued":["2025-05"],"dc:description.abstract":["This paper presents ResearchBuddy AI, an LLM-powered assistant that addresses modern research information overload problems. The project has three essential features, including finetuning of large language models for structured research paper summarization, prompt engineering for flexible summarization of variable formats of GitHub READMEs, and Retrieval-Augmented Generation (RAG) for context-aware question answering over domain-specific documents. The research explains model selection through self-hosted open-source models with LoRA (Low Rank Adaptation) and other techniques to achieve cost-effectiveness and control. The research paper summarization capability was developed through an iterative process. I began with small model and dataset experiments before scaling to a larger model with expanded datasets. During optimization on powerful hardware, I identified a “token performance paradox” about long context utilization. The development process faced three main challenges, which involved dealing with different input formats and managing token limits, optimizing computing resources, and solutions are presented. The RAG chatbot architecture uses vector databases together with powerful LLMs to generate responses that remain grounded in the original content. The future research agenda includes two main directions: investigating new model architecture suitable for processing long documents and developing YouTube video summarization with timestamp capabilities. ResearchBuddy AI is designed to provide researchers with intelligent tools for more efficient information management, with early development indicating strong potential to address research information overload."],"dc:format":["Text"],"dc:format.medium":["1 file (.pdf)"],"dc:identifier.uri":["https://hdl.handle.net/10877/21830"],"dc:language.iso":["en"],"dc:subject":["large language models","research assistants","summarization","retrieval-augmented generation","LoRA","prompt engineering","fine-tuning","RAG chatbot","vector databases","GitHub summarization"],"dc:title":["ResearchBuddy AI: LLM-Powered Assistant"],"dc:type":["Capstone"],"thesis:degree_discipline":["Computer Science"],"thesis:institution_name":["Texas State University"]},"updated_at":"2026-07-27T21:22:32Z"}