{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/125625"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/125625","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Glass onion: Compositional text-to-image generation using diffusion models and LLMs","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-02-04 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2025-02-04 without embargo terms","abstract_has_math":false,"creators":["Sarswat, Shrey"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Lazebnik, Svetlana"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-07-15","date_published":"2024-07-15","updated_at":"2026-07-22T22:25:02Z","subjects":["Text-to-image Generation","Diffusion Models","Large Language Models (llms)","Computer Vision","Language And Vision"],"languages":["en","eng"],"rights":["Copyright 2024 Shrey Sarswat"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/125625","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Lazebnik, Svetlana"]},{"key":"dc:creator","label":"Author","values":["Sarswat, Shrey"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2024-07-15","2024-08"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Text-to-image Generation","Diffusion Models","Large Language Models (llms)","Computer Vision","Language And Vision"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2024 Shrey Sarswat"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/125625"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-02-04 without embargo terms","The student, Shrey Sarswat, accepted the attached license on 2024-07-12 at 11:01.","The student, Shrey Sarswat, submitted this Thesis for approval on 2024-07-12 at 17:09.","This Thesis was approved for publication on 2024-07-15 at 14:12.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21089 on 2025-02-04 at 21:05:13","Text-to-image generation has seen substantial advancements in recent years, particularly with the advent of diffusion models, which have transformed how images are generated from text prompts. Despite these advancements, current methods often struggle with complex challenges such as accurately interpreting spatial and numerical relationships, unusual attributes, and logically intricate prompts. This work introduces a pioneering experimental approach specifically designed to tackle these complexities. Inspired by the layered editing techniques found in Adobe Photoshop, we propose an iterative framework that leverages Large Language Models (LLMs) along with state-of-the-art image generation models. This approach aims to create highly accurate and controllable visual representations from detailed textual descriptions. Our methodology involves decomposing text prompts into structured sub-prompts via LLMs, which are then sequentially rendered into images. To refine this process further, we integrate a dynamic feedback mechanism using a suite of foundational models. This system meticulously evaluates each generated image to ensure it aligns closely with the original prompt and maintains high visual fidelity. Our findings demonstrate the effectiveness of this approach, showing notable advancements for specific types of prompts while also revealing areas for improvement in others. This research underscores the considerable potential of our method and sets a foundation for future explorations in enhancing automated visual content generation."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Glass onion: Compositional text-to-image generation using diffusion models and LLMs"]}]}],"canonical_facts":{"dc:contributor":["Lazebnik, Svetlana"],"dc:creator":["Sarswat, Shrey"],"dc:date":["2024-07-15","2024-08"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-02-04 without embargo terms","The student, Shrey Sarswat, accepted the attached license on 2024-07-12 at 11:01.","The student, Shrey Sarswat, submitted this Thesis for approval on 2024-07-12 at 17:09.","This Thesis was approved for publication on 2024-07-15 at 14:12.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21089 on 2025-02-04 at 21:05:13","Text-to-image generation has seen substantial advancements in recent years, particularly with the advent of diffusion models, which have transformed how images are generated from text prompts. Despite these advancements, current methods often struggle with complex challenges such as accurately interpreting spatial and numerical relationships, unusual attributes, and logically intricate prompts. This work introduces a pioneering experimental approach specifically designed to tackle these complexities. Inspired by the layered editing techniques found in Adobe Photoshop, we propose an iterative framework that leverages Large Language Models (LLMs) along with state-of-the-art image generation models. This approach aims to create highly accurate and controllable visual representations from detailed textual descriptions. Our methodology involves decomposing text prompts into structured sub-prompts via LLMs, which are then sequentially rendered into images. To refine this process further, we integrate a dynamic feedback mechanism using a suite of foundational models. This system meticulously evaluates each generated image to ensure it aligns closely with the original prompt and maintains high visual fidelity. Our findings demonstrate the effectiveness of this approach, showing notable advancements for specific types of prompts while also revealing areas for improvement in others. This research underscores the considerable potential of our method and sets a foundation for future explorations in enhancing automated visual content generation."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/125625"],"dc:language":["en","eng"],"dc:rights":["Copyright 2024 Shrey Sarswat"],"dc:subject":["Text-to-image Generation","Diffusion Models","Large Language Models (llms)","Computer Vision","Language And Vision"],"dc:title":["Glass onion: Compositional text-to-image generation using diffusion models and LLMs"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:02Z"}