Back to search

University of Illinois at Urbana-Champaign

Compositional visual generation with energy-based modeling

Abstract

dc:description

Our understanding of the visual world around us is highly compositional in nature, since humans can rapidly understand individual concepts in a scene, and even compose them to describe the world states we encounter. However, machines struggle to understand complex composition of challenging concepts, such as confusing attributes of different objects or relations between objects. While a larger body of work has explored inferring and understanding objects in a scene, less work has been done on building a composable system that can enable “infinite use of finite means”, i.e., repeatedly reuse and recombine acquired concepts. This thesis endeavors to construct machine learning systems to have such com- positional capabilities, particularly in the context of generative modeling. First, existing works primarily compose relations by utilizing a holistic encoder that encodes inputs into fixed-size vectors, in the form of text or graphs. We instead propose to represent each relation as an unnormalized density (an energy-based model), enabling us to compose separate relations in a factorized manner. We show that such a factorized decomposition allows the model to both generate and edit scenes that have multiple sets of relations more faithfully. Second, we further extend our previous work to understand the composition of various concepts, including objects, relations and text descriptions. Our alternative structured approach for compositional generation involves interpreting diffusion models as energy-based mod- els, which allow us to explicitly combine data distributions defined by energy functions. This proposed method can generate scenes at test time that are substantially more complex than those seen in training, composing sentence descriptions, object relations, human facial attributes, and even generalizing to new combinations that are rarely seen. Third, we consider the inverse problem – given a collection of different images, can we discover the underlying generative concepts that represent each image? We present an approach to decompose and represent images into a set of different concepts, disentangling different art styles in paintings, objects, and lighting from kitchen scenes, and discovering image classes given ImageNet images. We illustrate how such discovered concepts accurately represent the underlying content of images and illustrate how they may further be composed with other concepts to construct new artistic and hybrid images. In summary, the proposed methods in this thesis showcase the potential for compositional modeling to enhance machine learning systems’ ability to generate complex and realistic scenes by intelligently combining learned generative concepts.

Degree

thesis:*
Name thesis:degree_name
M.S.
Level thesis:degree_level
Thesis
Discipline thesis:degree_discipline
Computer Science
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2023

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Liu, Nan
Contributors dc:contributor
  • Lazebnik, Svetlana

Subjects

dc:subject × 2

Rights

dc:rights
Statement dc:rights
  • Copyright 2023 Nan Liu
Language dc:language
en, eng

Identifiers

dc:identifier.*
Handle dc:identifier
https://hdl.handle.net/2142/121419

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Liu, Nan. Compositional visual generation with energy-based modeling. Thesis thesis, University of Illinois at Urbana-Champaign, 2023. https://hdl.handle.net/2142/121419