Back to results

University of Illinois at Urbana-Champaign

Glass onion: Compositional text-to-image generation using diffusion models and LLMs

Abstract

dc:description

Text-to-image generation has seen substantial advancements in recent years, particularly with the advent of diffusion models, which have transformed how images are generated from text prompts. Despite these advancements, current methods often struggle with complex challenges such as accurately interpreting spatial and numerical relationships, unusual attributes, and logically intricate prompts. This work introduces a pioneering experimental approach specifically designed to tackle these complexities. Inspired by the layered editing techniques found in Adobe Photoshop, we propose an iterative framework that leverages Large Language Models (LLMs) along with state-of-the-art image generation models. This approach aims to create highly accurate and controllable visual representations from detailed textual descriptions. Our methodology involves decomposing text prompts into structured sub-prompts via LLMs, which are then sequentially rendered into images. To refine this process further, we integrate a dynamic feedback mechanism using a suite of foundational models. This system meticulously evaluates each generated image to ensure it aligns closely with the original prompt and maintains high visual fidelity. Our findings demonstrate the effectiveness of this approach, showing notable advancements for specific types of prompts while also revealing areas for improvement in others. This research underscores the considerable potential of our method and sets a foundation for future explorations in enhancing automated visual content generation.

Degree

thesis:*
Name thesis:degree_name
M.S.
Level thesis:degree_level
Thesis
Discipline thesis:degree_discipline
Computer Science
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2024

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Sarswat, Shrey
Contributors dc:contributor
  • Lazebnik, Svetlana

Subjects

dc:subject × 5

Rights

dc:rights
Statement dc:rights
  • Copyright 2024 Shrey Sarswat
Language dc:language
en, eng

Identifiers

dc:identifier.*
Handle dc:identifier
https://hdl.handle.net/2142/125625

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Sarswat, Shrey. Glass onion: Compositional text-to-image generation using diffusion models and LLMs. Thesis thesis, University of Illinois at Urbana-Champaign, 2024. https://hdl.handle.net/2142/125625