{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/106392"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/106392","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Concepts from unclear textual embeddings for text-to-image synthesis","abstract":"Automatically generating images based on a natural language description is a challenging problem with several key applications in the fields of retail, marketing, education and entertainment. In the last few years, some progress has been made in this direction specifically by using Generative Adversarial Networks(GANs). Although current state of the art models can generate images that roughly adhere to the textual description, there still remains a long way to go, in terms of producing high quality images that adhere to the nuances of the sentence. To this end we propose CuteGAN, our simple text-to-image generation approach that encourages the model to leverage the attribute information while also attending to more relevant words in a sentence while generating images. We perform experiments on the competitive CUB-200 and MS-COCO datasets and achieve state-of-the-art performance on standard metrics of inception score and R-precision, indicating that our method produces more photo-realistic images that are better correlated with the text.","abstract_html":"Automatically generating images based on a natural language description is a challenging problem with several key applications in the fields of retail, marketing, education and entertainment. In the last few years, some progress has been made in this direction specifically by using Generative Adversarial Networks(GANs). Although current state of the art models can generate images that roughly adhere to the textual description, there still remains a long way to go, in terms of producing high quality images that adhere to the nuances of the sentence. To this end we propose CuteGAN, our simple text-to-image generation approach that encourages the model to leverage the attribute information while also attending to more relevant words in a sentence while generating images. We perform experiments on the competitive CUB-200 and MS-COCO datasets and achieve state-of-the-art performance on standard metrics of inception score and R-precision, indicating that our method produces more photo-realistic images that are better correlated with the text.","abstract_has_math":false,"creators":["Kumar, Maghav"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Schwing, Alexander"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2020,"date_issued":"2020-03-02T22:15:17Z","date_published":"2020-03-02T22:15:17Z","updated_at":"2026-07-22T22:24:47Z","subjects":["Computer Vision","deep learning","GANs","text-to-image","machine learning"],"languages":["en"],"rights":["Copyright 2019 Maghav Kumar"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/106392","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Schwing, Alexander"]},{"key":"dc:creator","label":"Author","values":["Kumar, Maghav"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2020-03-02T22:15:17Z","2022-03-03T10:15:16Z","2019-12-10","2019-12"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer Vision","deep learning","GANs","text-to-image","machine learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2019 Maghav Kumar"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/106392"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Automatically generating images based on a natural language description is a challenging problem with several key applications in the fields of retail, marketing, education and entertainment. In the last few years, some progress has been made in this direction specifically by using Generative Adversarial Networks(GANs). Although current state of the art models can generate images that roughly adhere to the textual description, there still remains a long way to go, in terms of producing high quality images that adhere to the nuances of the sentence. To this end we propose CuteGAN, our simple text-to-image generation approach that encourages the model to leverage the attribute information while also attending to more relevant words in a sentence while generating images. We perform experiments on the competitive CUB-200 and MS-COCO datasets and achieve state-of-the-art performance on standard metrics of inception score and R-precision, indicating that our method produces more photo-realistic images that are better correlated with the text.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2021-12-01","The student, Maghav Kumar, accepted the attached license on 2019-12-10 at 14:10.","The student, Maghav Kumar, submitted this Thesis for approval on 2019-12-10 at 14:14.","This Thesis was approved for publication on 2019-12-10 at 14:40.","DSpace SAF Submission Ingestion Package generated from Vireo submission #14773 on 2020-02-28 at 17:24:16","Made available in DSpace on 2020-03-02T22:15:17Z (GMT). No. of bitstreams: 2 KUMAR-THESIS-2019.pdf: 25611126 bytes, checksum: b3f3ce2eca99d13a4341b144c25fc4ae (MD5) LICENSE.txt: 4209 bytes, checksum: 4307707901b94b5ad7d4ec99a5857325 (MD5) Previous issue date: 2019-12-10","Embargo set by: Seth Robbins for item 113934 Lift date: 2022-03-02T22:15:21Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 113934 Lift date: 2022-03-02T22:18:25Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 113934 on 2022-03-03T10:15:16Z."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Concepts from unclear textual embeddings for text-to-image synthesis"]}]}],"canonical_facts":{"dc:contributor":["Schwing, Alexander"],"dc:creator":["Kumar, Maghav"],"dc:date":["2020-03-02T22:15:17Z","2022-03-03T10:15:16Z","2019-12-10","2019-12"],"dc:description":["Automatically generating images based on a natural language description is a challenging problem with several key applications in the fields of retail, marketing, education and entertainment. In the last few years, some progress has been made in this direction specifically by using Generative Adversarial Networks(GANs). Although current state of the art models can generate images that roughly adhere to the textual description, there still remains a long way to go, in terms of producing high quality images that adhere to the nuances of the sentence. To this end we propose CuteGAN, our simple text-to-image generation approach that encourages the model to leverage the attribute information while also attending to more relevant words in a sentence while generating images. We perform experiments on the competitive CUB-200 and MS-COCO datasets and achieve state-of-the-art performance on standard metrics of inception score and R-precision, indicating that our method produces more photo-realistic images that are better correlated with the text.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2021-12-01","The student, Maghav Kumar, accepted the attached license on 2019-12-10 at 14:10.","The student, Maghav Kumar, submitted this Thesis for approval on 2019-12-10 at 14:14.","This Thesis was approved for publication on 2019-12-10 at 14:40.","DSpace SAF Submission Ingestion Package generated from Vireo submission #14773 on 2020-02-28 at 17:24:16","Made available in DSpace on 2020-03-02T22:15:17Z (GMT). No. of bitstreams: 2 KUMAR-THESIS-2019.pdf: 25611126 bytes, checksum: b3f3ce2eca99d13a4341b144c25fc4ae (MD5) LICENSE.txt: 4209 bytes, checksum: 4307707901b94b5ad7d4ec99a5857325 (MD5) Previous issue date: 2019-12-10","Embargo set by: Seth Robbins for item 113934 Lift date: 2022-03-02T22:15:21Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 113934 Lift date: 2022-03-02T22:18:25Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 113934 on 2022-03-03T10:15:16Z."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/106392"],"dc:language":["en"],"dc:rights":["Copyright 2019 Maghav Kumar"],"dc:subject":["Computer Vision","deep learning","GANs","text-to-image","machine learning"],"dc:title":["Concepts from unclear textual embeddings for text-to-image synthesis"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:47Z"}