Back to results

University of Illinois at Urbana-Champaign

Efficient code-specific speculative decoding: Enhancing predictive accuracy and execution speed

Abstract

dc:description

Large Language Models (LLMs) have gained prominence due to their remarkable performance across diverse domains, providing high-quality and contextually relevant outputs that can be easily adapted to various applications. The success of these models has captured researchers’ attention, particularly regarding to the trade-offs between computational cost and performance: while larger LLMs yield superior outputs compared to their smaller counterparts, they also require significantly more computational and memory resources. This resource-intensive nature poses challenges in real-world applications, as scaling these models requires extensive infrastructure that may not be accessible or sustainable for all users and applications. To address these limitations, recent research has proposed Speculative Decoding—an approach that takes advantage of the efficiency of smaller models while preserving the performance benefits of larger ones. Speculative Decoding is based on the observation that many complex tasks can be decomposed into simpler subtasks, which both large and small models can handle with similar accuracy. By using a smaller model to perform preliminary inferences and selectively involving the larger model for more complex instances, this method significantly reduces resource consumption without sacrificing performance quality. In the context of programming language generation, however, unique structural and syntactic characteristics can offer additional opportunities for optimization. Unlike natural language, programming languages adhere to rigid syntactic rules and patterns, such as function and variable declarations, that provide context early in the generation process. This thesis investigates the potential for leveraging these syntactic structures within Speculative Decoding, aiming to further streamline the generation process for code. By integrating structured syntax with speculative execution techniques, this research seeks to advance computational efficiency in code generation, presenting a path toward resource-efficient LLMs that retain high performance in programming-specific tasks.

Degree

thesis:*
Name thesis:degree_name
M.S.
Level thesis:degree_level
Thesis
Discipline thesis:degree_discipline
Computer Science
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2024

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Wang, Yuhan
Contributors dc:contributor
  • Zhang, Lingming

Subjects

dc:subject × 2

Rights

dc:rights
Statement dc:rights
  • Copyright 2024 Yuhan Wang
Language dc:language
en, eng

Identifiers

dc:identifier.*
Handle dc:identifier
https://hdl.handle.net/2142/127359

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Wang, Yuhan. Efficient code-specific speculative decoding: Enhancing predictive accuracy and execution speed. Thesis thesis, University of Illinois at Urbana-Champaign, 2024. https://hdl.handle.net/2142/127359