{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/121419"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/121419","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Compositional visual generation with energy-based modeling","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2023-12-04 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2023-12-04 without embargo terms","abstract_has_math":false,"creators":["Liu, Nan"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Lazebnik, Svetlana"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2023,"date_issued":"2023-08","date_published":"2023-08","updated_at":"2026-07-22T22:24:57Z","subjects":["Energy-based Models","Diffusion Models"],"languages":["en","eng"],"rights":["Copyright 2023 Nan Liu"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/121419","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Lazebnik, Svetlana"]},{"key":"dc:creator","label":"Author","values":["Liu, Nan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2023-08","2023-06-21"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Energy-based Models","Diffusion Models"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2023 Nan Liu"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/121419"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2023-12-04 without embargo terms","The student, Nan Liu, accepted the attached license on 2023-06-14 at 20:07.","The student, Nan Liu, submitted this Thesis for approval on 2023-06-14 at 20:12.","This Thesis was approved for publication on 2023-06-21 at 14:25.","DSpace SAF Submission Ingestion Package generated from Vireo submission #19429 on 2023-12-04 at 16:59:59","Our understanding of the visual world around us is highly compositional in nature, since humans can rapidly understand individual concepts in a scene, and even compose them to describe the world states we encounter. However, machines struggle to understand complex composition of challenging concepts, such as confusing attributes of different objects or relations between objects. While a larger body of work has explored inferring and understanding objects in a scene, less work has been done on building a composable system that can enable “infinite use of finite means”, i.e., repeatedly reuse and recombine acquired concepts. This thesis endeavors to construct machine learning systems to have such com- positional capabilities, particularly in the context of generative modeling. First, existing works primarily compose relations by utilizing a holistic encoder that encodes inputs into fixed-size vectors, in the form of text or graphs. We instead propose to represent each relation as an unnormalized density (an energy-based model), enabling us to compose separate relations in a factorized manner. We show that such a factorized decomposition allows the model to both generate and edit scenes that have multiple sets of relations more faithfully. Second, we further extend our previous work to understand the composition of various concepts, including objects, relations and text descriptions. Our alternative structured approach for compositional generation involves interpreting diffusion models as energy-based mod- els, which allow us to explicitly combine data distributions defined by energy functions. This proposed method can generate scenes at test time that are substantially more complex than those seen in training, composing sentence descriptions, object relations, human facial attributes, and even generalizing to new combinations that are rarely seen. Third, we consider the inverse problem – given a collection of different images, can we discover the underlying generative concepts that represent each image? We present an approach to decompose and represent images into a set of different concepts, disentangling different art styles in paintings, objects, and lighting from kitchen scenes, and discovering image classes given ImageNet images. We illustrate how such discovered concepts accurately represent the underlying content of images and illustrate how they may further be composed with other concepts to construct new artistic and hybrid images. In summary, the proposed methods in this thesis showcase the potential for compositional modeling to enhance machine learning systems’ ability to generate complex and realistic scenes by intelligently combining learned generative concepts."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Compositional visual generation with energy-based modeling"]}]}],"canonical_facts":{"dc:contributor":["Lazebnik, Svetlana"],"dc:creator":["Liu, Nan"],"dc:date":["2023-08","2023-06-21"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2023-12-04 without embargo terms","The student, Nan Liu, accepted the attached license on 2023-06-14 at 20:07.","The student, Nan Liu, submitted this Thesis for approval on 2023-06-14 at 20:12.","This Thesis was approved for publication on 2023-06-21 at 14:25.","DSpace SAF Submission Ingestion Package generated from Vireo submission #19429 on 2023-12-04 at 16:59:59","Our understanding of the visual world around us is highly compositional in nature, since humans can rapidly understand individual concepts in a scene, and even compose them to describe the world states we encounter. However, machines struggle to understand complex composition of challenging concepts, such as confusing attributes of different objects or relations between objects. While a larger body of work has explored inferring and understanding objects in a scene, less work has been done on building a composable system that can enable “infinite use of finite means”, i.e., repeatedly reuse and recombine acquired concepts. This thesis endeavors to construct machine learning systems to have such com- positional capabilities, particularly in the context of generative modeling. First, existing works primarily compose relations by utilizing a holistic encoder that encodes inputs into fixed-size vectors, in the form of text or graphs. We instead propose to represent each relation as an unnormalized density (an energy-based model), enabling us to compose separate relations in a factorized manner. We show that such a factorized decomposition allows the model to both generate and edit scenes that have multiple sets of relations more faithfully. Second, we further extend our previous work to understand the composition of various concepts, including objects, relations and text descriptions. Our alternative structured approach for compositional generation involves interpreting diffusion models as energy-based mod- els, which allow us to explicitly combine data distributions defined by energy functions. This proposed method can generate scenes at test time that are substantially more complex than those seen in training, composing sentence descriptions, object relations, human facial attributes, and even generalizing to new combinations that are rarely seen. Third, we consider the inverse problem – given a collection of different images, can we discover the underlying generative concepts that represent each image? We present an approach to decompose and represent images into a set of different concepts, disentangling different art styles in paintings, objects, and lighting from kitchen scenes, and discovering image classes given ImageNet images. We illustrate how such discovered concepts accurately represent the underlying content of images and illustrate how they may further be composed with other concepts to construct new artistic and hybrid images. In summary, the proposed methods in this thesis showcase the potential for compositional modeling to enhance machine learning systems’ ability to generate complex and realistic scenes by intelligently combining learned generative concepts."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/121419"],"dc:language":["en","eng"],"dc:rights":["Copyright 2023 Nan Liu"],"dc:subject":["Energy-based Models","Diffusion Models"],"dc:title":["Compositional visual generation with energy-based modeling"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:57Z"}