Back to results

University of Illinois at Urbana-Champaign

Concepts from unclear textual embeddings for text-to-image synthesis

Abstract

dc:description

Automatically generating images based on a natural language description is a challenging problem with several key applications in the fields of retail, marketing, education and entertainment. In the last few years, some progress has been made in this direction specifically by using Generative Adversarial Networks(GANs). Although current state of the art models can generate images that roughly adhere to the textual description, there still remains a long way to go, in terms of producing high quality images that adhere to the nuances of the sentence. To this end we propose CuteGAN, our simple text-to-image generation approach that encourages the model to leverage the attribute information while also attending to more relevant words in a sentence while generating images. We perform experiments on the competitive CUB-200 and MS-COCO datasets and achieve state-of-the-art performance on standard metrics of inception score and R-precision, indicating that our method produces more photo-realistic images that are better correlated with the text.

Degree

thesis:*
Name thesis:degree_name
M.S.
Level thesis:degree_level
Thesis
Discipline thesis:degree_discipline
Computer Science
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2020

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Kumar, Maghav
Contributors dc:contributor
  • Schwing, Alexander

Subjects

dc:subject × 5

Rights

dc:rights
Statement dc:rights
  • Copyright 2019 Maghav Kumar
Language dc:language
en

Identifiers

dc:identifier.*
Handle dc:identifier
http://hdl.handle.net/2142/106392
OAI identifier oai:identifier
oai:www.ideals.illinois.edu:2142/106392

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Kumar, Maghav. Concepts from unclear textual embeddings for text-to-image synthesis. Thesis thesis, University of Illinois at Urbana-Champaign, 2020. http://hdl.handle.net/2142/106392