University of Illinois at Urbana-Champaign
Exploring efficiency improvements to distillation of language models for code
Abstract
dc:descriptionLarge Language Models (LLMs) trained on code have demonstrated great success in the code domain across various tasks including code generation, repair, test generation, and more. These models are increasingly being adopted in both industry for practical software development and in academia to address challenges in software engineering that previously relied on heuristic-based methods or conventional machine learning techniques. While prompting can be an easy way to deploy these models for certain tasks, fine-tuning them for specific tasks often yields more reliable results. Therefore fine-tuning is widely prevalent in the industry and is still used for several tasks in academia. However, with the increasing size of these models, a major drawback of fine-tuning is the high cost, arising primarily from the extensive GPU compute time required. This compute time is especially large when we consider the full development and debugging phase of fine-tuning these models. Therefore, in this study, we aim to reduce these costs by reducing the time needed for fine-tuning while striving to minimally impact predictive performance. We primarily explore data subsampling techniques that can help reduce the size of the fine-tuning dataset thereby reducing associated costs. We specifically focus on fine-tuning within the context of model distillation. Previous research has utilized functional correctness-based methods for data sub-sampling, which can be effective for standalone programming tasks. However, many real-world coding tasks involve complexities such as setting up environments or integrating external resources, making evaluating functional correctness infeasible. Our work explores alternative sampling-based techniques that rely only on the probabilistic outputs of models instead of code execution. Through extensive empirical evaluation on an MBPP-based dataset, we show the feasibility of reducing fine-tuning time while maintaining performance. More specifically, we demonstrate that our methodology can achieve a 37% reduction in fine-tuning time with negligible impact on predictive performance, and in some instances, up to a 75% reduction without any loss in performance.
Degree
thesis:*- Name thesis:degree_name
- M.S.
- Level thesis:degree_level
- Thesis
- Discipline thesis:degree_discipline
- Computer Science
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2024
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Karthikeyan, Ajaykrishna
- Contributors dc:contributor
-
- Zhang, Lingming
Subjects
dc:subject × 6Rights
dc:rights- Statement dc:rights
-
- Copyright 2024 Ajaykrishna Karthikeyan
- Language dc:language
- en, eng
Identifiers
dc:identifier.*- Handle dc:identifier
- https://hdl.handle.net/2142/124446