{"id":{"repo_id":"umkc","oai_identifier":"oai:mospace.umsystem.edu:10355/75785"},"canonical_url":"https://search.dev.ndltd.org/etd/umkc/oai:mospace.umsystem.edu:10355/75785","repository":{"repo_id":"umkc","name":"University of Missouri - Kansas City","base_url":"https://mospace.umsystem.edu/oai/request"},"display":{"title":"Compression for Machine Vision and Beyond","abstract":"Compression has been one of the most fundamental and elusive challenges in both academia and industry. With the sheer increase of high-definition video content over the internet, developing improved compression algorithms becomes an urgent necessity. This thesis tackles the problem of visual content compression: how to reduce the transmitted data volume under specific application scenarios. One of the core steps is how to remove the redundancy to achieve a compact latent representation. We approach the problem from two directions: prediction and transform. While a typical prediction process targets at removing the statistical redundancy between the reference and current image blocks, and transform further removes the inter-pixel redundancy between residual samples. We will address compression for machine vision and related topics. In compression for machine vision, machines will communicate amongst themselves to perform tasks without a human in the mix, which requires a separate pipeline to achieve optimal coding performance. We aim to investigate how to efficiently transmit image features in low latency scenario and focus on developing a multiple-transform solution to achieve a more compact data representation for image retrieval task. Multiple-transform solution is proven to be more efficient to preserve more distinguishable properties for a large-scale dataset. However, over-sized transform candidate list burdens the bit-rate constraint. We develop a merge scheme to search for the optimal transforms from available transform candidates. We will also present our efforts at contributing the development of next-generation video coding standard: Versatile Video Coding (VVC), and exploring improved intra prediction schemes beyond the High Efficiency Video Coding (HEVC) standard. 1) Based on observations on the properties of DST-7 and DCT-8, a dual-implementation support solution is developed to reduce the software run-time complexity. The (anti-)symmetric features are leveraged to reduce the number of arithmetic operations involved in deriving the transformed coefficients from the residual block. The scheme has been adopted by MPEG VVC standardization development group and was integrated into VVC reference software. 2) In prediction-relevant attempts, we explore both traditional and Convolutional Neural Network (CNN)-based schemes. Multiple Linear Regression is utilized to further explore spatial correlation with reference pixels and existing intra prediction. A CNN-based scheme is developed by combining local and non-local information for intra prediction. We demonstrate the effectiveness of these approaches.","abstract_html":"Compression has been one of the most fundamental and elusive challenges in both academia and industry. With the sheer increase of high-definition video content over the internet, developing improved compression algorithms becomes an urgent necessity. This thesis tackles the problem of visual content compression: how to reduce the transmitted data volume under specific application scenarios. One of the core steps is how to remove the redundancy to achieve a compact latent representation. We approach the problem from two directions: prediction and transform. While a typical prediction process targets at removing the statistical redundancy between the reference and current image blocks, and transform further removes the inter-pixel redundancy between residual samples. We will address compression for machine vision and related topics. In compression for machine vision, machines will communicate amongst themselves to perform tasks without a human in the mix, which requires a separate pipeline to achieve optimal coding performance. We aim to investigate how to efficiently transmit image features in low latency scenario and focus on developing a multiple-transform solution to achieve a more compact data representation for image retrieval task. Multiple-transform solution is proven to be more efficient to preserve more distinguishable properties for a large-scale dataset. However, over-sized transform candidate list burdens the bit-rate constraint. We develop a merge scheme to search for the optimal transforms from available transform candidates. We will also present our efforts at contributing the development of next-generation video coding standard: Versatile Video Coding (VVC), and exploring improved intra prediction schemes beyond the High Efficiency Video Coding (HEVC) standard. 1) Based on observations on the properties of DST-7 and DCT-8, a dual-implementation support solution is developed to reduce the software run-time complexity. The (anti-)symmetric features are leveraged to reduce the number of arithmetic operations involved in deriving the transformed coefficients from the residual block. The scheme has been adopted by MPEG VVC standardization development group and was integrated into VVC reference software. 2) In prediction-relevant attempts, we explore both traditional and Convolutional Neural Network (CNN)-based schemes. Multiple Linear Regression is utilized to further explore spatial correlation with reference pixels and existing intra prediction. A CNN-based scheme is developed by combining local and non-local information for intra prediction. We demonstrate the effectiveness of these approaches.","abstract_has_math":false,"creators":["Zhang, Zhaobin"],"institution":"University of Missouri--Kansas City","degree_name":"Ph.D. (Doctor of Philosophy)","degree_level":"Doctoral","degree_discipline":"Electrical and Electronics Engineering (UMKC)","degree_department":null,"school":null,"contributors":[],"advisors":["Li, Zhu"],"committee_chairs":[],"committee_members":[],"year":2020,"date_issued":"2020","date_published":"2020","updated_at":"2026-07-24T05:16:20Z","subjects":[],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/10355/75785","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Li, Zhu"]},{"key":"dc:creator","label":"Author","values":["Zhang, Zhaobin"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2020-08-13T16:03:27Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2020-08-13T16:03:27Z"]},{"key":"dc:date.issued","label":"Date","values":["2020"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical and Electronics Engineering (UMKC)","Computer Science (UMKC)"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Doctoral"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D. (Doctor of Philosophy)"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Missouri--Kansas City"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10355/75785"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Title from PDF of title page viewed August 31, 2020","Dissertation advisor: Zhu Li","Vita","Includes bibliographical references (pages 123-140)","Thesis (Ph.D.)--School of Computing and Engineering. University of Missouri--Kansas City, 2020"]},{"key":"dc:description.abstract","label":"Abstract","values":["Compression has been one of the most fundamental and elusive challenges in both academia and industry. With the sheer increase of high-definition video content over the internet, developing improved compression algorithms becomes an urgent necessity. This thesis tackles the problem of visual content compression: how to reduce the transmitted data volume under specific application scenarios. One of the core steps is how to remove the redundancy to achieve a compact latent representation. We approach the problem from two directions: prediction and transform. While a typical prediction process targets at removing the statistical redundancy between the reference and current image blocks, and transform further removes the inter-pixel redundancy between residual samples. We will address compression for machine vision and related topics. In compression for machine vision, machines will communicate amongst themselves to perform tasks without a human in the mix, which requires a separate pipeline to achieve optimal coding performance. We aim to investigate how to efficiently transmit image features in low latency scenario and focus on developing a multiple-transform solution to achieve a more compact data representation for image retrieval task. Multiple-transform solution is proven to be more efficient to preserve more distinguishable properties for a large-scale dataset. However, over-sized transform candidate list burdens the bit-rate constraint. We develop a merge scheme to search for the optimal transforms from available transform candidates. We will also present our efforts at contributing the development of next-generation video coding standard: Versatile Video Coding (VVC), and exploring improved intra prediction schemes beyond the High Efficiency Video Coding (HEVC) standard. 1) Based on observations on the properties of DST-7 and DCT-8, a dual-implementation support solution is developed to reduce the software run-time complexity. The (anti-)symmetric features are leveraged to reduce the number of arithmetic operations involved in deriving the transformed coefficients from the residual block. The scheme has been adopted by MPEG VVC standardization development group and was integrated into VVC reference software. 2) In prediction-relevant attempts, we explore both traditional and Convolutional Neural Network (CNN)-based schemes. Multiple Linear Regression is utilized to further explore spatial correlation with reference pixels and existing intra prediction. A CNN-based scheme is developed by combining local and non-local information for intra prediction. We demonstrate the effectiveness of these approaches."]},{"key":"dc:title","label":"Title","values":["Compression for Machine Vision and Beyond"]}]}],"canonical_facts":{"dc:contributor.advisor":["Li, Zhu"],"dc:creator":["Zhang, Zhaobin"],"dc:date.accessioned":["2020-08-13T16:03:27Z"],"dc:date.available":["2020-08-13T16:03:27Z"],"dc:date.issued":["2020"],"dc:description":["Title from PDF of title page viewed August 31, 2020","Dissertation advisor: Zhu Li","Vita","Includes bibliographical references (pages 123-140)","Thesis (Ph.D.)--School of Computing and Engineering. University of Missouri--Kansas City, 2020"],"dc:description.abstract":["Compression has been one of the most fundamental and elusive challenges in both academia and industry. With the sheer increase of high-definition video content over the internet, developing improved compression algorithms becomes an urgent necessity. This thesis tackles the problem of visual content compression: how to reduce the transmitted data volume under specific application scenarios. One of the core steps is how to remove the redundancy to achieve a compact latent representation. We approach the problem from two directions: prediction and transform. While a typical prediction process targets at removing the statistical redundancy between the reference and current image blocks, and transform further removes the inter-pixel redundancy between residual samples. We will address compression for machine vision and related topics. In compression for machine vision, machines will communicate amongst themselves to perform tasks without a human in the mix, which requires a separate pipeline to achieve optimal coding performance. We aim to investigate how to efficiently transmit image features in low latency scenario and focus on developing a multiple-transform solution to achieve a more compact data representation for image retrieval task. Multiple-transform solution is proven to be more efficient to preserve more distinguishable properties for a large-scale dataset. However, over-sized transform candidate list burdens the bit-rate constraint. We develop a merge scheme to search for the optimal transforms from available transform candidates. We will also present our efforts at contributing the development of next-generation video coding standard: Versatile Video Coding (VVC), and exploring improved intra prediction schemes beyond the High Efficiency Video Coding (HEVC) standard. 1) Based on observations on the properties of DST-7 and DCT-8, a dual-implementation support solution is developed to reduce the software run-time complexity. The (anti-)symmetric features are leveraged to reduce the number of arithmetic operations involved in deriving the transformed coefficients from the residual block. The scheme has been adopted by MPEG VVC standardization development group and was integrated into VVC reference software. 2) In prediction-relevant attempts, we explore both traditional and Convolutional Neural Network (CNN)-based schemes. Multiple Linear Regression is utilized to further explore spatial correlation with reference pixels and existing intra prediction. A CNN-based scheme is developed by combining local and non-local information for intra prediction. We demonstrate the effectiveness of these approaches."],"dc:identifier.uri":["https://hdl.handle.net/10355/75785"],"dc:title":["Compression for Machine Vision and Beyond"],"thesis:degree_discipline":["Electrical and Electronics Engineering (UMKC)","Computer Science (UMKC)"],"thesis:degree_level":["Doctoral"],"thesis:degree_name":["Ph.D. (Doctor of Philosophy)"],"thesis:institution_name":["University of Missouri--Kansas City"]},"updated_at":"2026-07-24T05:16:20Z"}