University of Illinois at Urbana-Champaign
Efficient code-specific speculative decoding: Enhancing predictive accuracy and execution speed
Abstract
dc:descriptionLarge Language Models (LLMs) have gained prominence due to their remarkable performance across diverse domains, providing high-quality and contextually relevant outputs that can be easily adapted to various applications. The success of these models has captured researchers’ attention, particularly regarding to the trade-offs between computational cost and performance: while larger LLMs yield superior outputs compared to their smaller counterparts, they also require significantly more computational and memory resources. This resource-intensive nature poses challenges in real-world applications, as scaling these models requires extensive infrastructure that may not be accessible or sustainable for all users and applications. To address these limitations, recent research has proposed Speculative Decoding—an approach that takes advantage of the efficiency of smaller models while preserving the performance benefits of larger ones. Speculative Decoding is based on the observation that many complex tasks can be decomposed into simpler subtasks, which both large and small models can handle with similar accuracy. By using a smaller model to perform preliminary inferences and selectively involving the larger model for more complex instances, this method significantly reduces resource consumption without sacrificing performance quality. In the context of programming language generation, however, unique structural and syntactic characteristics can offer additional opportunities for optimization. Unlike natural language, programming languages adhere to rigid syntactic rules and patterns, such as function and variable declarations, that provide context early in the generation process. This thesis investigates the potential for leveraging these syntactic structures within Speculative Decoding, aiming to further streamline the generation process for code. By integrating structured syntax with speculative execution techniques, this research seeks to advance computational efficiency in code generation, presenting a path toward resource-efficient LLMs that retain high performance in programming-specific tasks.
Degree
thesis:*- Name thesis:degree_name
- M.S.
- Level thesis:degree_level
- Thesis
- Discipline thesis:degree_discipline
- Computer Science
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2024
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Wang, Yuhan
- Contributors dc:contributor
-
- Zhang, Lingming
Subjects
dc:subject × 2Rights
dc:rights- Statement dc:rights
-
- Copyright 2024 Yuhan Wang
- Language dc:language
- en, eng
Identifiers
dc:identifier.*- Handle dc:identifier
- https://hdl.handle.net/2142/127359