University of Ontario Institute of Technology
Addressing data challenges in LLM-enhanced software engineering
Abstract
dc:description.abstractThe rapid adoption of large language models (LLMs) is reshaping software engineering practice, yet it reveals a critical dichotomy: while resource-intensive tasks like code generation benefit from large datasets and standardized benchmarks, they also face significant risks from data contamination and inflated model evaluations. Conversely, resource-constrained methods, such as flaky test detection, often operate under strict constraints, including limited labeled data and restricted computational resources. This thesis investigates this dichotomy, focusing on key tensions related to dataset integrity, computational efficiency, and benchmark reliability. Through comprehensive empirical analysis and targeted methodological innovations, we propose practical solutions enabling robust, transparent, fair, and sustainable integration of LLMs into software engineering workflows. By addressing both resource-rich and resource-constrained contexts, this work seeks to bridge fundamental gaps between the theoretical capabilities of LLMs and their real-world applicability, guiding researchers and practitioners towards more responsible and trustworthy automation.
Degree
thesis:*- Name thesis:degree_name
- Master of Science (MSc)
- Discipline thesis:degree_discipline
- Computer Science
- Grantor
- University of Ontario Institute of Technology
- Year dc:date.issued
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- More, Riddhi
- Advisor dc:contributor.advisor
-
- Bradbury, Jeremy
Rights
- Language dc:language.iso
- en
Identifiers
dc:identifier.*- Handle dc:identifier.uri
- https://hdl.handle.net/10155/2014
- OAI identifier oai:identifier
- oai:ontariotechu.scholaris.ca:10155/2014