{"id":{"repo_id":"embry-riddle","oai_identifier":"oai:commons.erau.edu:edt-1925"},"canonical_url":"https://search.dev.ndltd.org/etd/embry-riddle/oai:commons.erau.edu:edt-1925","repository":{"repo_id":"embry-riddle","name":"Embry Riddle Aeronautical University","base_url":"https://commons.erau.edu/do/oai/"},"display":{"title":"On the Provenance of Software Systems: Automating Software Traceability with Knowledge Graph and Large Language Model Synergy","abstract":"<p>The present dissertation delineates a system that enables those engaged in software development to automatically generate and maintain project life cycle provenance. All projects are implemented and made manifest with the development of artifacts, e.g., papers, code files, etc. Tools exist to accelerate artifact creation, but little focus is paid to the processes that produce them. In terms of Ontology, or, from Ancient Greek, the <em>study of being</em>, the two most basic entities in reality are Continuant and Occurrent, or, roughly, “Artifact” and “Process”. This dissertation posits that for any created artifact, its process of creation, i.e., its life cycle<em> provenance</em>, must also be captured and maintained. Artifacts are often delivered without an explicit trace of their evolution. This is particularly unacceptable for critical systems, where requirements documents, codebases, and other meta-artifacts are revisited without a corresponding history of how or why they came to be, leading to confusion and rework.</p> <p>While the software development life cycle (SDLC) incorporates meta-artifacts like traceability matrices to improve artifact provenance, these are typically informal, heavy with natural language, and lack structured explainability. This work proposes that each artifact should be attended by a machine-readable, human-interpretable, extensible provenance record, implemented in the form of a knowledge graph, backed by well-established ontologies. The developed system, ProvTracer, leverages structured knowledge via PROV-O and the Basic Formal Ontology alongside generative natural language capabilities via the Generative Pretrained Transformer (GPT) series of Multimodal Large Language Models (MLLMs), to create real-time, traceable, and explainable links between development activities and their resulting artifacts. By capturing these provenance trace links automatically through multimodal signals, e.g., screenshots, peripheral device input, etc., ProvTracer aims to bridge the gap between implicit processes and explicit traces, enabling developers to understand, query, maintain, integrate, and trust the evolution of their projects and systems.</p> <p>The synergy between knowledge graphs and MLLMs enables a novel form of interactive, explainable software development. Natural language queries of provenance trace link knowledge graphs can reduce information overload, extract developer rationale and decision histories, support task assignment, and a range of project management activities. This aligns with a burgeoning trend in research demonstrating that structured knowledge improves machine learning trust, transparency and reproducibility. The present dissertation addresses the challenges of traceability and explainability in the SDLC by presenting a system that automatically captures artifact provenance and operationalizes it for practical use in real-world software development.</p>","abstract_html":"&lt;p&gt;The present dissertation delineates a system that enables those engaged in software development to automatically generate and maintain project life cycle provenance. All projects are implemented and made manifest with the development of artifacts, e.g., papers, code files, etc. Tools exist to accelerate artifact creation, but little focus is paid to the processes that produce them. In terms of Ontology, or, from Ancient Greek, the &lt;em&gt;study of being&lt;/em&gt;, the two most basic entities in reality are Continuant and Occurrent, or, roughly, “Artifact” and “Process”. This dissertation posits that for any created artifact, its process of creation, i.e., its life cycle&lt;em&gt; provenance&lt;/em&gt;, must also be captured and maintained. Artifacts are often delivered without an explicit trace of their evolution. This is particularly unacceptable for critical systems, where requirements documents, codebases, and other meta-artifacts are revisited without a corresponding history of how or why they came to be, leading to confusion and rework.&lt;/p&gt; &lt;p&gt;While the software development life cycle (SDLC) incorporates meta-artifacts like traceability matrices to improve artifact provenance, these are typically informal, heavy with natural language, and lack structured explainability. This work proposes that each artifact should be attended by a machine-readable, human-interpretable, extensible provenance record, implemented in the form of a knowledge graph, backed by well-established ontologies. The developed system, ProvTracer, leverages structured knowledge via PROV-O and the Basic Formal Ontology alongside generative natural language capabilities via the Generative Pretrained Transformer (GPT) series of Multimodal Large Language Models (MLLMs), to create real-time, traceable, and explainable links between development activities and their resulting artifacts. By capturing these provenance trace links automatically through multimodal signals, e.g., screenshots, peripheral device input, etc., ProvTracer aims to bridge the gap between implicit processes and explicit traces, enabling developers to understand, query, maintain, integrate, and trust the evolution of their projects and systems.&lt;/p&gt; &lt;p&gt;The synergy between knowledge graphs and MLLMs enables a novel form of interactive, explainable software development. Natural language queries of provenance trace link knowledge graphs can reduce information overload, extract developer rationale and decision histories, support task assignment, and a range of project management activities. This aligns with a burgeoning trend in research demonstrating that structured knowledge improves machine learning trust, transparency and reproducibility. The present dissertation addresses the challenges of traceability and explainability in the SDLC by presenting a system that automatically captures artifact provenance and operationalizes it for practical use in real-world software development.&lt;/p&gt;","abstract_has_math":false,"creators":["Procko, Tyler"],"institution":null,"degree_name":"Doctor of Philosophy in Electrical Engineering & Computer Science","degree_level":"Dissertation - Open Access","degree_discipline":"Electrical Engineering and Computer Science","degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-04-01T07:00:00Z","date_published":"2025-04-01T07:00:00Z","updated_at":"2026-07-27T19:26:16Z","subjects":["knowledge graphs","ontology","bfo","provenance","llm","gpt","traceability","software engineering","sdlc","metascience","Applied Behavior Analysis","Archival Science","Cataloging and Metadata","Cognition and Perception","Cognitive Science","Communication Technology and New Media","Computational Linguistics","Computer and Systems Architecture","Computer Engineering","Data Storage Systems","Experimental Analysis of Behavior","Graphic Communications","Library and Information Science","Linguistics","Operational Research","Organizational Communication","Signal Processing","Systems Engineering"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://commons.erau.edu/edt/898","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["Procko, Tyler"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical Engineering and Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation - Open Access"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctor of Philosophy in Electrical Engineering & Computer Science"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["knowledge graphs","ontology","bfo","provenance","llm","gpt","traceability","software engineering","sdlc","metascience","Applied Behavior Analysis","Archival Science","Cataloging and Metadata","Cognition and Perception","Cognitive Science","Communication Technology and New Media","Computational Linguistics","Computer and Systems Architecture","Computer Engineering","Data Storage Systems","Experimental Analysis of Behavior","Graphic Communications","Library and Information Science","Linguistics","Operational Research","Organizational Communication","Signal Processing","Systems Engineering"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://commons.erau.edu/edt/898"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>The present dissertation delineates a system that enables those engaged in software development to automatically generate and maintain project life cycle provenance. All projects are implemented and made manifest with the development of artifacts, e.g., papers, code files, etc. Tools exist to accelerate artifact creation, but little focus is paid to the processes that produce them. In terms of Ontology, or, from Ancient Greek, the <em>study of being</em>, the two most basic entities in reality are Continuant and Occurrent, or, roughly, “Artifact” and “Process”. This dissertation posits that for any created artifact, its process of creation, i.e., its life cycle<em> provenance</em>, must also be captured and maintained. Artifacts are often delivered without an explicit trace of their evolution. This is particularly unacceptable for critical systems, where requirements documents, codebases, and other meta-artifacts are revisited without a corresponding history of how or why they came to be, leading to confusion and rework.</p> <p>While the software development life cycle (SDLC) incorporates meta-artifacts like traceability matrices to improve artifact provenance, these are typically informal, heavy with natural language, and lack structured explainability. This work proposes that each artifact should be attended by a machine-readable, human-interpretable, extensible provenance record, implemented in the form of a knowledge graph, backed by well-established ontologies. The developed system, ProvTracer, leverages structured knowledge via PROV-O and the Basic Formal Ontology alongside generative natural language capabilities via the Generative Pretrained Transformer (GPT) series of Multimodal Large Language Models (MLLMs), to create real-time, traceable, and explainable links between development activities and their resulting artifacts. By capturing these provenance trace links automatically through multimodal signals, e.g., screenshots, peripheral device input, etc., ProvTracer aims to bridge the gap between implicit processes and explicit traces, enabling developers to understand, query, maintain, integrate, and trust the evolution of their projects and systems.</p> <p>The synergy between knowledge graphs and MLLMs enables a novel form of interactive, explainable software development. Natural language queries of provenance trace link knowledge graphs can reduce information overload, extract developer rationale and decision histories, support task assignment, and a range of project management activities. This aligns with a burgeoning trend in research demonstrating that structured knowledge improves machine learning trust, transparency and reproducibility. The present dissertation addresses the challenges of traceability and explainability in the SDLC by presenting a system that automatically captures artifact provenance and operationalizes it for practical use in real-world software development.</p>"]},{"key":"dc:title","label":"Title","values":["On the Provenance of Software Systems: Automating Software Traceability with Knowledge Graph and Large Language Model Synergy"]}]}],"canonical_facts":{"dc:creator":["Procko, Tyler"],"dc:description.abstract":["<p>The present dissertation delineates a system that enables those engaged in software development to automatically generate and maintain project life cycle provenance. All projects are implemented and made manifest with the development of artifacts, e.g., papers, code files, etc. Tools exist to accelerate artifact creation, but little focus is paid to the processes that produce them. In terms of Ontology, or, from Ancient Greek, the <em>study of being</em>, the two most basic entities in reality are Continuant and Occurrent, or, roughly, “Artifact” and “Process”. This dissertation posits that for any created artifact, its process of creation, i.e., its life cycle<em> provenance</em>, must also be captured and maintained. Artifacts are often delivered without an explicit trace of their evolution. This is particularly unacceptable for critical systems, where requirements documents, codebases, and other meta-artifacts are revisited without a corresponding history of how or why they came to be, leading to confusion and rework.</p> <p>While the software development life cycle (SDLC) incorporates meta-artifacts like traceability matrices to improve artifact provenance, these are typically informal, heavy with natural language, and lack structured explainability. This work proposes that each artifact should be attended by a machine-readable, human-interpretable, extensible provenance record, implemented in the form of a knowledge graph, backed by well-established ontologies. The developed system, ProvTracer, leverages structured knowledge via PROV-O and the Basic Formal Ontology alongside generative natural language capabilities via the Generative Pretrained Transformer (GPT) series of Multimodal Large Language Models (MLLMs), to create real-time, traceable, and explainable links between development activities and their resulting artifacts. By capturing these provenance trace links automatically through multimodal signals, e.g., screenshots, peripheral device input, etc., ProvTracer aims to bridge the gap between implicit processes and explicit traces, enabling developers to understand, query, maintain, integrate, and trust the evolution of their projects and systems.</p> <p>The synergy between knowledge graphs and MLLMs enables a novel form of interactive, explainable software development. Natural language queries of provenance trace link knowledge graphs can reduce information overload, extract developer rationale and decision histories, support task assignment, and a range of project management activities. This aligns with a burgeoning trend in research demonstrating that structured knowledge improves machine learning trust, transparency and reproducibility. The present dissertation addresses the challenges of traceability and explainability in the SDLC by presenting a system that automatically captures artifact provenance and operationalizes it for practical use in real-world software development.</p>"],"dc:identifier":["https://commons.erau.edu/edt/898"],"dc:subject":["knowledge graphs","ontology","bfo","provenance","llm","gpt","traceability","software engineering","sdlc","metascience","Applied Behavior Analysis","Archival Science","Cataloging and Metadata","Cognition and Perception","Cognitive Science","Communication Technology and New Media","Computational Linguistics","Computer and Systems Architecture","Computer Engineering","Data Storage Systems","Experimental Analysis of Behavior","Graphic Communications","Library and Information Science","Linguistics","Operational Research","Organizational Communication","Signal Processing","Systems Engineering"],"dc:title":["On the Provenance of Software Systems: Automating Software Traceability with Knowledge Graph and Large Language Model Synergy"],"thesis:degree_discipline":["Electrical Engineering and Computer Science"],"thesis:degree_level":["Dissertation - Open Access"],"thesis:degree_name":["Doctor of Philosophy in Electrical Engineering & Computer Science"]},"updated_at":"2026-07-27T19:26:16Z"}