Back to search

University of Illinois at Urbana-Champaign

Learning-based saliency-aware compression framework

Abstract

dc:description

Networked vision systems are prevalent nowadays due to the dominance of video traffic over the Internet. Videos in the networked vision system are transmitted over the Internet to a vision application, which processes videos by deep neural networks (DNN), i.e., computation offloading, or presents videos to human viewers, i.e., video streaming. The video codec is an essential component in the networked vision system, which compresses videos to a small size with minimal impact on the video quality. Traditional video codecs like MPEG, AVC, and HEVC have been widely adopted in the networked vision system in the past few decades. Recently, learned video codecs built on DNNs are gaining more attention due to their superior coding efficiency than traditional ones. Nonetheless, DNNs in networked vision systems introduce rate and quality heterogeneity, which lead to some issues in the networked vision system, e.g., low processing rate and poor coding efficiency. This dissertation aims at addressing the heterogeneity in networked vision systems. The thesis statement is that the rate and quality heterogeneity in networked vision systems should be addressed by leveraging the spatiotemporal saliency with learning-based adaptation in video compression. The spatiotemporal saliency characterizes how important pixels in the video are, which by our observation varies across different networked vision systems. Specifically, certain regions of interest (ROI) or keyframes are more significant in a given networked vision system. The key intuition is to adapt how information in a video is encoded spatially and temporally based on spatiotemporal saliency, i.e., saliency-aware adaptation to improve processing rate or coding efficiency with minimal impact on other metrics. The challenge lies in properly designing and training the adaptation module in a given networked vision system. To prove the thesis, this dissertation provides a suite of saliency-aware adaptation modules in the video codec, i.e., temporal adaptation (Learned Frame Sampling), spatial adaptation (Context-aware Compression and Spatial-adaptive Filter), and spatiotemporal adaptation (Space-time-aware Entropy Module), addressing the aforementioned challenge. Our evaluation covers various video categories, diverse vision applications, and state-of-the-art compression algorithms. The result demonstrates our approaches effectively improve the processing rate or the coding efficiency in networked vision systems compared to existing solutions without harming other performance metrics.

Degree

thesis:*
Name thesis:degree_name
Ph.D.
Level thesis:degree_level
Dissertation
Discipline thesis:degree_discipline
Computer Science
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2022

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Chen, Bo
Contributors dc:contributor
  • Nahrstedt, Klara
  • Abdelzaher, Tarek
  • Sundaram, Hari
  • Yan, Zhisheng

Subjects

dc:subject × 4

Rights

dc:rights
Statement dc:rights
  • Copyright 2022 Bo Chen
Language dc:language
en, eng

Identifiers

dc:identifier.*
Handle dc:identifier
https://hdl.handle.net/2142/115495

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Chen, Bo. Learning-based saliency-aware compression framework. Dissertation thesis, University of Illinois at Urbana-Champaign, 2022. https://hdl.handle.net/2142/115495