Back to search

University of Illinois at Urbana-Champaign

Towards high-quality, accessible, and universal video matting

Abstract

dc:description

Video matting seeks to predict alpha mattes for each frame in a given video sequence. While current approaches leverage deep convolutional neural networks (CNNs) to achieve accurate alpha matte predictions, these methods are typically trained and evaluated on private or inaccessible matting datasets. Furthermore, existing solutions generate only a single alpha matte per frame, without distinguishing between individual human instances present in the video. In this dissertation, we aim to address these limitations in video matting techniques and extend the capabilities of video matting to accommodate more complex scenarios. we first propose VideoMatt, a simple and strong real-time video matting baseline model VideoMatt-S/T, with a fully composited accessible video matting benchmark. Then, we propose VMFormer, which utilizes a transformer-based end-to-end method for video matting that outperforms previous CNN-based state-of-the-art solutions, which sets a new track for video matting solutions and demonstrates the potential of our proposed approach. Considering that both VideoMatt and VMFormer can only handle semantic-aware video matting, we further propose Video Instance Matting (VIM), a new task estimating alpha mattes of each instance at each frame of a video sequence, with corresponding VIM50 benchmark and the baseline model MSG-VIM. VIM extends video matting to an instance-aware problem and improves its applicability. Finally, we introduce Matting Anything, a method to estimate the instance-aware and class-aware alpha matte with one universal model. This approach broadens the matting task to accommodate more general use cases, addressing the evolving needs of modern image and video editing. We also explore future directions, such as developing an end-to-end video editing system and integrating the video matting capacity into the multimodal large language models (LLMs).

Degree

thesis:*
Name thesis:degree_name
Ph.D.
Level thesis:degree_level
Dissertation
Discipline thesis:degree_discipline
Electrical & Computer Engr
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2024

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Li, Jiachen
Contributors dc:contributor
  • Shi, Humphrey
  • Hwu, Wen-Mei
  • Hasegawa-Johnson, Mark
  • Xiong, Jinjun

Subjects

dc:subject × 1

Rights

dc:rights
Statement dc:rights
  • Copyright 2024 Jiachen Li
Language dc:language
en, eng

Identifiers

dc:identifier.*
Handle dc:identifier
https://hdl.handle.net/2142/127197

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Li, Jiachen. Towards high-quality, accessible, and universal video matting. Dissertation thesis, University of Illinois at Urbana-Champaign, 2024. https://hdl.handle.net/2142/127197