{"id":{"repo_id":"queens","oai_identifier":"oai:queensu.scholaris.ca:1974/35343"},"canonical_url":"https://search.dev.ndltd.org/etd/queens/oai:queensu.scholaris.ca:1974/35343","repository":{"repo_id":"queens","name":"Queens University","base_url":"https://qspace.library.queensu.ca/server/oai/request"},"display":{"title":"Detect, Explain, Ground: A Sustainable Pipeline for Robust Large Language Model–Powered Agents","abstract":"Large language models (LLMs) enable fluent dialogue but still suffer from critical breakdowns, such as incoherence, irrelevance, or factual inaccuracies (hallucinations). This thesis develops a sustainable, three-stage pipeline to detect, manage, and ground LLM-powered agents, enhancing their robustness and reliability. First, to establish a baseline for self-monitoring, we evaluate general-purpose LLMs on DBDC5, demonstrating that models like GPT-4 can surpass specialized detectors. While effective, relying on such large models for continuous monitoring is computationally unsustainable. Building on this finding, we propose a novel \"monitor-explain-escalate\" architecture. In this system, a lightweight detector provides real-time analysis and escalates high-risk conversational turns to a stronger model, preserving quality while significantly reducing costs. Finally, to address the core issue of factual accuracy that underlies many breakdowns, we introduce a Hierarchical Lexical Graph to improve evidence retrieval for multi-hop question answering. Across five datasets, this graph-augmented retrieval method achieves a 23.1% relative improvement in recall and correctness over baseline RAG systems. Together, these contributions advance the development of reliable, interpretable, and resource-efficient LLM agents by establishing a comprehensive framework that manages conversational flow, operationalizes low-carbon management, and grounds multi-document reasoning in verifiable evidence.","abstract_html":"Large language models (LLMs) enable fluent dialogue but still suffer from critical breakdowns, such as incoherence, irrelevance, or factual inaccuracies (hallucinations). This thesis develops a sustainable, three-stage pipeline to detect, manage, and ground LLM-powered agents, enhancing their robustness and reliability. First, to establish a baseline for self-monitoring, we evaluate general-purpose LLMs on DBDC5, demonstrating that models like GPT-4 can surpass specialized detectors. While effective, relying on such large models for continuous monitoring is computationally unsustainable. Building on this finding, we propose a novel &quot;monitor-explain-escalate&quot; architecture. In this system, a lightweight detector provides real-time analysis and escalates high-risk conversational turns to a stronger model, preserving quality while significantly reducing costs. Finally, to address the core issue of factual accuracy that underlies many breakdowns, we introduce a Hierarchical Lexical Graph to improve evidence retrieval for multi-hop question answering. Across five datasets, this graph-augmented retrieval method achieves a 23.1% relative improvement in recall and correctness over baseline RAG systems. Together, these contributions advance the development of reliable, interpretable, and resource-efficient LLM agents by establishing a comprehensive framework that manages conversational flow, operationalizes low-carbon management, and grounds multi-document reasoning in verifiable evidence.","abstract_has_math":false,"creators":["Ghassel, Abdellah"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":"Electrical and Computer Engineering","school":null,"contributors":[],"advisors":["Zhu, Xiaodan"],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-10-03","date_published":"2025-10-03","updated_at":"2026-07-27T20:35:29Z","subjects":["conversational artificial intelligence","graph machine learning","large language models"],"languages":["eng"],"rights":["Attribution-ShareAlike 4.0 International"],"rights_urls":["http://creativecommons.org/licenses/by-sa/4.0/"],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/1974/35343","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.department","label":"Department","values":["Electrical and Computer Engineering"]},{"key":"dc:contributor.supervisor","label":"Supervisor","values":["Zhu, Xiaodan"]},{"key":"dc:creator","label":"Author","values":["Ghassel, Abdellah"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2025-10-03T15:02:14Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2025-10-03T15:02:14Z"]},{"key":"dc:date.issued","label":"Date","values":["2025-10-03"]},{"key":"dc:type","label":"Dc Type","values":["thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["conversational artificial intelligence","graph machine learning","large language models"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Attribution-ShareAlike 4.0 International"]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://creativecommons.org/licenses/by-sa/4.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/1974/35343"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Large language models (LLMs) enable fluent dialogue but still suffer from critical breakdowns, such as incoherence, irrelevance, or factual inaccuracies (hallucinations). This thesis develops a sustainable, three-stage pipeline to detect, manage, and ground LLM-powered agents, enhancing their robustness and reliability. First, to establish a baseline for self-monitoring, we evaluate general-purpose LLMs on DBDC5, demonstrating that models like GPT-4 can surpass specialized detectors. While effective, relying on such large models for continuous monitoring is computationally unsustainable. Building on this finding, we propose a novel \"monitor-explain-escalate\" architecture. In this system, a lightweight detector provides real-time analysis and escalates high-risk conversational turns to a stronger model, preserving quality while significantly reducing costs. Finally, to address the core issue of factual accuracy that underlies many breakdowns, we introduce a Hierarchical Lexical Graph to improve evidence retrieval for multi-hop question answering. Across five datasets, this graph-augmented retrieval method achieves a 23.1% relative improvement in recall and correctness over baseline RAG systems. Together, these contributions advance the development of reliable, interpretable, and resource-efficient LLM agents by establishing a comprehensive framework that manages conversational flow, operationalizes low-carbon management, and grounds multi-document reasoning in verifiable evidence."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["M.A.Sc."]},{"key":"dc:title","label":"Title","values":["Detect, Explain, Ground: A Sustainable Pipeline for Robust Large Language Model–Powered Agents"]}]}],"canonical_facts":{"dc:contributor.department":["Electrical and Computer Engineering"],"dc:contributor.supervisor":["Zhu, Xiaodan"],"dc:creator":["Ghassel, Abdellah"],"dc:date.accessioned":["2025-10-03T15:02:14Z"],"dc:date.available":["2025-10-03T15:02:14Z"],"dc:date.issued":["2025-10-03"],"dc:description.abstract":["Large language models (LLMs) enable fluent dialogue but still suffer from critical breakdowns, such as incoherence, irrelevance, or factual inaccuracies (hallucinations). This thesis develops a sustainable, three-stage pipeline to detect, manage, and ground LLM-powered agents, enhancing their robustness and reliability. First, to establish a baseline for self-monitoring, we evaluate general-purpose LLMs on DBDC5, demonstrating that models like GPT-4 can surpass specialized detectors. While effective, relying on such large models for continuous monitoring is computationally unsustainable. Building on this finding, we propose a novel \"monitor-explain-escalate\" architecture. In this system, a lightweight detector provides real-time analysis and escalates high-risk conversational turns to a stronger model, preserving quality while significantly reducing costs. Finally, to address the core issue of factual accuracy that underlies many breakdowns, we introduce a Hierarchical Lexical Graph to improve evidence retrieval for multi-hop question answering. Across five datasets, this graph-augmented retrieval method achieves a 23.1% relative improvement in recall and correctness over baseline RAG systems. Together, these contributions advance the development of reliable, interpretable, and resource-efficient LLM agents by establishing a comprehensive framework that manages conversational flow, operationalizes low-carbon management, and grounds multi-document reasoning in verifiable evidence."],"dc:description.degree":["M.A.Sc."],"dc:identifier.uri":["https://hdl.handle.net/1974/35343"],"dc:language.iso":["eng"],"dc:rights":["Attribution-ShareAlike 4.0 International"],"dc:rights.uri":["http://creativecommons.org/licenses/by-sa/4.0/"],"dc:subject":["conversational artificial intelligence","graph machine learning","large language models"],"dc:title":["Detect, Explain, Ground: A Sustainable Pipeline for Robust Large Language Model–Powered Agents"],"dc:type":["thesis"]},"updated_at":"2026-07-27T20:35:29Z"}