{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/105709"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/105709","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Audio compression via nonlinear transform coding and stochastic binary activation","abstract":"Engineers have pushed the boundaries of audio compression and designed numerous lossy audio compression codecs, such as ACC, WNA, and others, that have surpassed the longstanding MP3 coding format. However most of the methods are laboriously engineered using psychoacoustic modeling, and some of them are proprietary and only see limited use. This thesis, inspired by recent major breakthroughs in lossy image compression via machine learning methods, explores the possibilities of a neural network trained for lossy audio compression. Currently there are few if any audio compression methods that utilize machine learning. This thesis presents a brief introduction to lossy transform compression and compares it to similar machine learning concepts, then systematically presents a convolutional autoencoder network with a stochastic binary activation for a sparse representation of the code space to achieve compression. A similar network is employed for encoding the residual of the main network. Our network achieves average compression rates of roughly 5 to 2 and introduces few if any audible artifacts, presenting a promising opening to audio compression using machine learning.","abstract_html":"Engineers have pushed the boundaries of audio compression and designed numerous lossy audio compression codecs, such as ACC, WNA, and others, that have surpassed the longstanding MP3 coding format. However most of the methods are laboriously engineered using psychoacoustic modeling, and some of them are proprietary and only see limited use. This thesis, inspired by recent major breakthroughs in lossy image compression via machine learning methods, explores the possibilities of a neural network trained for lossy audio compression. Currently there are few if any audio compression methods that utilize machine learning. This thesis presents a brief introduction to lossy transform compression and compares it to similar machine learning concepts, then systematically presents a convolutional autoencoder network with a stochastic binary activation for a sparse representation of the code space to achieve compression. A similar network is employed for encoding the residual of the main network. Our network achieves average compression rates of roughly 5 to 2 and introduces few if any audible artifacts, presenting a promising opening to audio compression using machine learning.","abstract_has_math":false,"creators":["Yan, Yuanheng"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Smaragdis, Paris"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019-11-26T20:35:14Z","date_published":"2019-11-26T20:35:14Z","updated_at":"2026-07-22T22:24:44Z","subjects":["Audio compression","Neural network","Convolutional neural network (CNN)","Stochastic binary activation"],"languages":["en"],"rights":["Copyright 2019 Yuanheng Yan"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/105709","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Smaragdis, Paris"]},{"key":"dc:creator","label":"Author","values":["Yan, Yuanheng"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2019-11-26T20:35:14Z","2019-07-18","2019-08"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Audio compression","Neural network","Convolutional neural network (CNN)","Stochastic binary activation"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2019 Yuanheng Yan"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/105709"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Engineers have pushed the boundaries of audio compression and designed numerous lossy audio compression codecs, such as ACC, WNA, and others, that have surpassed the longstanding MP3 coding format. However most of the methods are laboriously engineered using psychoacoustic modeling, and some of them are proprietary and only see limited use. This thesis, inspired by recent major breakthroughs in lossy image compression via machine learning methods, explores the possibilities of a neural network trained for lossy audio compression. Currently there are few if any audio compression methods that utilize machine learning. This thesis presents a brief introduction to lossy transform compression and compares it to similar machine learning concepts, then systematically presents a convolutional autoencoder network with a stochastic binary activation for a sparse representation of the code space to achieve compression. A similar network is employed for encoding the residual of the main network. Our network achieves average compression rates of roughly 5 to 2 and introduces few if any audible artifacts, presenting a promising opening to audio compression using machine learning.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2019-11-26 without embargo terms","The student, Yuanheng Yan, accepted the attached license on 2019-07-17 at 11:20.","The student, Yuanheng Yan, submitted this Thesis for approval on 2019-07-17 at 11:27.","This Thesis was approved for publication on 2019-07-18 at 10:26.","DSpace SAF Submission Ingestion Package generated from Vireo submission #14353 on 2019-11-26 at 12:53:54","Made available in DSpace on 2019-11-26T20:35:14Z (GMT). No. of bitstreams: 2 YAN-THESIS-2019.pdf: 10479858 bytes, checksum: 80dd087c0d334d621aa87dc899291522 (MD5) LICENSE.txt: 4209 bytes, checksum: 2daabd56eded156b34f21f5945ae41a9 (MD5) Previous issue date: 2019-07-18"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Audio compression via nonlinear transform coding and stochastic binary activation"]}]}],"canonical_facts":{"dc:contributor":["Smaragdis, Paris"],"dc:creator":["Yan, Yuanheng"],"dc:date":["2019-11-26T20:35:14Z","2019-07-18","2019-08"],"dc:description":["Engineers have pushed the boundaries of audio compression and designed numerous lossy audio compression codecs, such as ACC, WNA, and others, that have surpassed the longstanding MP3 coding format. However most of the methods are laboriously engineered using psychoacoustic modeling, and some of them are proprietary and only see limited use. This thesis, inspired by recent major breakthroughs in lossy image compression via machine learning methods, explores the possibilities of a neural network trained for lossy audio compression. Currently there are few if any audio compression methods that utilize machine learning. This thesis presents a brief introduction to lossy transform compression and compares it to similar machine learning concepts, then systematically presents a convolutional autoencoder network with a stochastic binary activation for a sparse representation of the code space to achieve compression. A similar network is employed for encoding the residual of the main network. Our network achieves average compression rates of roughly 5 to 2 and introduces few if any audible artifacts, presenting a promising opening to audio compression using machine learning.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2019-11-26 without embargo terms","The student, Yuanheng Yan, accepted the attached license on 2019-07-17 at 11:20.","The student, Yuanheng Yan, submitted this Thesis for approval on 2019-07-17 at 11:27.","This Thesis was approved for publication on 2019-07-18 at 10:26.","DSpace SAF Submission Ingestion Package generated from Vireo submission #14353 on 2019-11-26 at 12:53:54","Made available in DSpace on 2019-11-26T20:35:14Z (GMT). No. of bitstreams: 2 YAN-THESIS-2019.pdf: 10479858 bytes, checksum: 80dd087c0d334d621aa87dc899291522 (MD5) LICENSE.txt: 4209 bytes, checksum: 2daabd56eded156b34f21f5945ae41a9 (MD5) Previous issue date: 2019-07-18"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/105709"],"dc:language":["en"],"dc:rights":["Copyright 2019 Yuanheng Yan"],"dc:subject":["Audio compression","Neural network","Convolutional neural network (CNN)","Stochastic binary activation"],"dc:title":["Audio compression via nonlinear transform coding and stochastic binary activation"],"dc:type":["text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:44Z"}