Back to results

Virginia Tech

Reinforcement Learning–Based Discrete Prompt Optimization for Neuro-Symbolic Structured Simplification of Complex Game Descriptions with Large Language Models

Abstract

dc:description.abstract

This thesis investigates how large language models can be trained to perform structured simplification of complex, free-form game descriptions for the GameChangineer platform. The work formalizes simplification as a discrete prompt optimization problem and introduces a neuro-symbolic pipeline that maps raw natural language into controlled GameChangineer sentences via scenario normalization, retrieval-augmented code generation, and AST-based FACTS extraction. A reinforcement learning framework based on Proximal Policy Optimization optimizes discrete prompt edits using task-specific rewards that combine grammar compliance, semantic agreement with the FACTS contract, and compiler validity of the resulting games. Experiments on diverse arcade-style game descriptions show that the proposed GC-Repair and sentence correction agents significantly improve grammar-constrained generation, robustness to noisy user input, and end-to-end code correctness compared to direct LLM rewriting baselines.

Degree

thesis:*
Name thesis:degree_name
Master of Science
Level thesis:degree_level
masters
Discipline thesis:degree_discipline
Computer Engineering
Department dc:contributor.department
Electrical and Computer Engineering
Grantor dc:publisher
Virginia Tech
Year dc:date.issued
2026

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Bhatt, Shubham Satyaprakash
Chairs dc:contributor.committeechair
  • Hsiao, Michael S.
  • Abbott, Amos L.
Committee member dc:contributor.committeemember
  • Wang, Yue J.

Subjects

dc:subject × 5

Rights

dc:rights
Statement dc:rights
  • In Copyright
Language dc:language.iso
en

Identifiers

dc:identifier.*
Dc Identifier Other
vt_gsexam:45464
OAI identifier oai:identifier
oai:vtechworks.lib.vt.edu:10919/140897

Chain of custody

source
Harvested from
Virginia Tech
Base URL
vtechworks.lib.vt.edu/oai/request
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Bhatt, Shubham Satyaprakash. Reinforcement Learning–Based Discrete Prompt Optimization for Neuro-Symbolic Structured Simplification of Complex Game Descriptions with Large Language Models. masters thesis, Virginia Tech, 2026. https://hdl.handle.net/10919/140897