{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/129203"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/129203","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Large language model for programming by example","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2025-10-19 without embargo terms","abstract_has_math":false,"creators":["Zhang, Shuning"],"institution":"University of Illinois Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Park, Yongjoo"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-04-16","date_published":"2025-04-16","updated_at":"2026-07-22T22:25:04Z","subjects":["Program By Example (PBE)","LLM","Prompt Engineering"],"languages":["en","eng"],"rights":["Copyright 2025 Shuning Zhang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/129203","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Park, Yongjoo"]},{"key":"dc:creator","label":"Author","values":["Zhang, Shuning"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-04-16","2025-05"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Program By Example (PBE)","LLM","Prompt Engineering"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Shuning Zhang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/129203"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","The student, Shuning Zhang, accepted the attached license on 2025-04-15 at 17:33.","The student, Shuning Zhang, submitted this Thesis for approval on 2025-04-15 at 17:37.","This Thesis was approved for publication on 2025-04-16 at 09:26.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21783 on 2025-10-19 at 18:09:24","Programming by Example (PBE) is a technique in which the system generates a program to automate complex transformation, with the user simply providing input and output example pairs. While traditional PBE systems like FlashFill and Foofah have demonstrated strong ability in program synthesis, they are often restrained from domain-specific tasks and lack adaptability across varied data types. With the emergence of Large Language Models (LLMs) and their reasoning and code generation ability, there is the opportunity to enhance PBE tasks with LLMs. This thesis investigates how LLM, specifically GPT 4.0, is able to perform PBE-style data wrangling tasks. Three methods are proposed to improve LLM performance on the tasks: (1) zero-shot prompt engineering, (2) loop-based verification prompting, and (3) a hybrid framework combining LLMs with traditional PBE models. Experiments and evaluation on benchmark dataset Foofah and Prose show that LLM not only has the ability to generalize across different types of data but also achieves higher accuracy than the traditional systems, with the hybrid method reaching the highest 85% overall accuracy on the Foofah dataset. This work highlights both the promise and current limitations of LLM on the PBE style tasks and offers potential future exploration of more robust LLM-driven reasoning and program synthesis."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Large language model for programming by example"]}]}],"canonical_facts":{"dc:contributor":["Park, Yongjoo"],"dc:creator":["Zhang, Shuning"],"dc:date":["2025-04-16","2025-05"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","The student, Shuning Zhang, accepted the attached license on 2025-04-15 at 17:33.","The student, Shuning Zhang, submitted this Thesis for approval on 2025-04-15 at 17:37.","This Thesis was approved for publication on 2025-04-16 at 09:26.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21783 on 2025-10-19 at 18:09:24","Programming by Example (PBE) is a technique in which the system generates a program to automate complex transformation, with the user simply providing input and output example pairs. While traditional PBE systems like FlashFill and Foofah have demonstrated strong ability in program synthesis, they are often restrained from domain-specific tasks and lack adaptability across varied data types. With the emergence of Large Language Models (LLMs) and their reasoning and code generation ability, there is the opportunity to enhance PBE tasks with LLMs. This thesis investigates how LLM, specifically GPT 4.0, is able to perform PBE-style data wrangling tasks. Three methods are proposed to improve LLM performance on the tasks: (1) zero-shot prompt engineering, (2) loop-based verification prompting, and (3) a hybrid framework combining LLMs with traditional PBE models. Experiments and evaluation on benchmark dataset Foofah and Prose show that LLM not only has the ability to generalize across different types of data but also achieves higher accuracy than the traditional systems, with the hybrid method reaching the highest 85% overall accuracy on the Foofah dataset. This work highlights both the promise and current limitations of LLM on the PBE style tasks and offers potential future exploration of more robust LLM-driven reasoning and program synthesis."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/129203"],"dc:language":["en","eng"],"dc:rights":["Copyright 2025 Shuning Zhang"],"dc:subject":["Program By Example (PBE)","LLM","Prompt Engineering"],"dc:title":["Large language model for programming by example"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:04Z"}