University of Illinois Urbana-Champaign
Large language model for programming by example
Abstract
dc:descriptionProgramming by Example (PBE) is a technique in which the system generates a program to automate complex transformation, with the user simply providing input and output example pairs. While traditional PBE systems like FlashFill and Foofah have demonstrated strong ability in program synthesis, they are often restrained from domain-specific tasks and lack adaptability across varied data types. With the emergence of Large Language Models (LLMs) and their reasoning and code generation ability, there is the opportunity to enhance PBE tasks with LLMs. This thesis investigates how LLM, specifically GPT 4.0, is able to perform PBE-style data wrangling tasks. Three methods are proposed to improve LLM performance on the tasks: (1) zero-shot prompt engineering, (2) loop-based verification prompting, and (3) a hybrid framework combining LLMs with traditional PBE models. Experiments and evaluation on benchmark dataset Foofah and Prose show that LLM not only has the ability to generalize across different types of data but also achieves higher accuracy than the traditional systems, with the hybrid method reaching the highest 85% overall accuracy on the Foofah dataset. This work highlights both the promise and current limitations of LLM on the PBE style tasks and offers potential future exploration of more robust LLM-driven reasoning and program synthesis.
Degree
thesis:*- Name thesis:degree_name
- M.S.
- Level thesis:degree_level
- Thesis
- Discipline thesis:degree_discipline
- Computer Science
- Grantor
- University of Illinois Urbana-Champaign
- Year dc:date
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Zhang, Shuning
- Contributors dc:contributor
-
- Park, Yongjoo
Subjects
dc:subject × 3Rights
dc:rights- Statement dc:rights
-
- Copyright 2025 Shuning Zhang
- Language dc:language
- en, eng
Identifiers
dc:identifier.*- Handle dc:identifier
- https://hdl.handle.net/2142/129203