University of Illinois at Urbana-Champaign
Compositional visual generation with energy-based modeling
Abstract
dc:descriptionOur understanding of the visual world around us is highly compositional in nature, since humans can rapidly understand individual concepts in a scene, and even compose them to describe the world states we encounter. However, machines struggle to understand complex composition of challenging concepts, such as confusing attributes of different objects or relations between objects. While a larger body of work has explored inferring and understanding objects in a scene, less work has been done on building a composable system that can enable “infinite use of finite means”, i.e., repeatedly reuse and recombine acquired concepts. This thesis endeavors to construct machine learning systems to have such com- positional capabilities, particularly in the context of generative modeling. First, existing works primarily compose relations by utilizing a holistic encoder that encodes inputs into fixed-size vectors, in the form of text or graphs. We instead propose to represent each relation as an unnormalized density (an energy-based model), enabling us to compose separate relations in a factorized manner. We show that such a factorized decomposition allows the model to both generate and edit scenes that have multiple sets of relations more faithfully. Second, we further extend our previous work to understand the composition of various concepts, including objects, relations and text descriptions. Our alternative structured approach for compositional generation involves interpreting diffusion models as energy-based mod- els, which allow us to explicitly combine data distributions defined by energy functions. This proposed method can generate scenes at test time that are substantially more complex than those seen in training, composing sentence descriptions, object relations, human facial attributes, and even generalizing to new combinations that are rarely seen. Third, we consider the inverse problem – given a collection of different images, can we discover the underlying generative concepts that represent each image? We present an approach to decompose and represent images into a set of different concepts, disentangling different art styles in paintings, objects, and lighting from kitchen scenes, and discovering image classes given ImageNet images. We illustrate how such discovered concepts accurately represent the underlying content of images and illustrate how they may further be composed with other concepts to construct new artistic and hybrid images. In summary, the proposed methods in this thesis showcase the potential for compositional modeling to enhance machine learning systems’ ability to generate complex and realistic scenes by intelligently combining learned generative concepts.
Degree
thesis:*- Name thesis:degree_name
- M.S.
- Level thesis:degree_level
- Thesis
- Discipline thesis:degree_discipline
- Computer Science
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2023
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Liu, Nan
- Contributors dc:contributor
-
- Lazebnik, Svetlana
Subjects
dc:subject × 2Rights
dc:rights- Statement dc:rights
-
- Copyright 2023 Nan Liu
- Language dc:language
- en, eng
Identifiers
dc:identifier.*- Handle dc:identifier
- https://hdl.handle.net/2142/121419