{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/127197"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/127197","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Towards high-quality, accessible, and universal video matting","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-03-28 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2025-03-28 without embargo terms","abstract_has_math":false,"creators":["Li, Jiachen"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Shi, Humphrey","Hwu, Wen-Mei","Hasegawa-Johnson, Mark","Xiong, Jinjun"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-12-03","date_published":"2024-12-03","updated_at":"2026-07-22T22:25:03Z","subjects":["Video Matting"],"languages":["en","eng"],"rights":["Copyright 2024 Jiachen Li"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/127197","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Shi, Humphrey","Hwu, Wen-Mei","Hasegawa-Johnson, Mark","Xiong, Jinjun"]},{"key":"dc:creator","label":"Author","values":["Li, Jiachen"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2024-12-03","2024-12"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Video Matting"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2024 Jiachen Li"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/127197"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-03-28 without embargo terms","The student, Jiachen Li, accepted the attached license on 2024-11-18 at 01:38.","The student, Jiachen Li, submitted this Dissertation for approval on 2024-11-18 at 01:45.","This Dissertation was approved for publication on 2024-12-03 at 15:40.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21332 on 2025-03-28 at 14:25:40","Video matting seeks to predict alpha mattes for each frame in a given video sequence. While current approaches leverage deep convolutional neural networks (CNNs) to achieve accurate alpha matte predictions, these methods are typically trained and evaluated on private or inaccessible matting datasets. Furthermore, existing solutions generate only a single alpha matte per frame, without distinguishing between individual human instances present in the video. In this dissertation, we aim to address these limitations in video matting techniques and extend the capabilities of video matting to accommodate more complex scenarios. we first propose VideoMatt, a simple and strong real-time video matting baseline model VideoMatt-S/T, with a fully composited accessible video matting benchmark. Then, we propose VMFormer, which utilizes a transformer-based end-to-end method for video matting that outperforms previous CNN-based state-of-the-art solutions, which sets a new track for video matting solutions and demonstrates the potential of our proposed approach. Considering that both VideoMatt and VMFormer can only handle semantic-aware video matting, we further propose Video Instance Matting (VIM), a new task estimating alpha mattes of each instance at each frame of a video sequence, with corresponding VIM50 benchmark and the baseline model MSG-VIM. VIM extends video matting to an instance-aware problem and improves its applicability. Finally, we introduce Matting Anything, a method to estimate the instance-aware and class-aware alpha matte with one universal model. This approach broadens the matting task to accommodate more general use cases, addressing the evolving needs of modern image and video editing. We also explore future directions, such as developing an end-to-end video editing system and integrating the video matting capacity into the multimodal large language models (LLMs)."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Towards high-quality, accessible, and universal video matting"]}]}],"canonical_facts":{"dc:contributor":["Shi, Humphrey","Hwu, Wen-Mei","Hasegawa-Johnson, Mark","Xiong, Jinjun"],"dc:creator":["Li, Jiachen"],"dc:date":["2024-12-03","2024-12"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-03-28 without embargo terms","The student, Jiachen Li, accepted the attached license on 2024-11-18 at 01:38.","The student, Jiachen Li, submitted this Dissertation for approval on 2024-11-18 at 01:45.","This Dissertation was approved for publication on 2024-12-03 at 15:40.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21332 on 2025-03-28 at 14:25:40","Video matting seeks to predict alpha mattes for each frame in a given video sequence. While current approaches leverage deep convolutional neural networks (CNNs) to achieve accurate alpha matte predictions, these methods are typically trained and evaluated on private or inaccessible matting datasets. Furthermore, existing solutions generate only a single alpha matte per frame, without distinguishing between individual human instances present in the video. In this dissertation, we aim to address these limitations in video matting techniques and extend the capabilities of video matting to accommodate more complex scenarios. we first propose VideoMatt, a simple and strong real-time video matting baseline model VideoMatt-S/T, with a fully composited accessible video matting benchmark. Then, we propose VMFormer, which utilizes a transformer-based end-to-end method for video matting that outperforms previous CNN-based state-of-the-art solutions, which sets a new track for video matting solutions and demonstrates the potential of our proposed approach. Considering that both VideoMatt and VMFormer can only handle semantic-aware video matting, we further propose Video Instance Matting (VIM), a new task estimating alpha mattes of each instance at each frame of a video sequence, with corresponding VIM50 benchmark and the baseline model MSG-VIM. VIM extends video matting to an instance-aware problem and improves its applicability. Finally, we introduce Matting Anything, a method to estimate the instance-aware and class-aware alpha matte with one universal model. This approach broadens the matting task to accommodate more general use cases, addressing the evolving needs of modern image and video editing. We also explore future directions, such as developing an end-to-end video editing system and integrating the video matting capacity into the multimodal large language models (LLMs)."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/127197"],"dc:language":["en","eng"],"dc:rights":["Copyright 2024 Jiachen Li"],"dc:subject":["Video Matting"],"dc:title":["Towards high-quality, accessible, and universal video matting"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:03Z"}