{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/120294"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/120294","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Unsupervised sound separation","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2023-09-01 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2023-09-01 without embargo terms","abstract_has_math":false,"creators":["Tzinis, Efthymios"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Smaragdis, Paris","Hasegawa-Johnson, Mark","Misailovic, Sasa","Hershey, John R"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2023,"date_issued":"2023-05","date_published":"2023-05","updated_at":"2026-07-22T22:24:57Z","subjects":["Sound Separation","Audio-visual Perception","Unsupervised Learning","Self-supervised Learning","Speech Enhancement","Efficient Neural Networks","Federated Learning"],"languages":["en","eng"],"rights":["Copyright 2023 Efthymios Tzinis"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/120294","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Smaragdis, Paris","Hasegawa-Johnson, Mark","Misailovic, Sasa","Hershey, John R"]},{"key":"dc:creator","label":"Author","values":["Tzinis, Efthymios"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2023-05","2023-04-24"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Sound Separation","Audio-visual Perception","Unsupervised Learning","Self-supervised Learning","Speech Enhancement","Efficient Neural Networks","Federated Learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2023 Efthymios Tzinis"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/120294"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2023-09-01 without embargo terms","The student, Efthymios Tzinis, accepted the attached license on 2023-04-20 at 07:58.","The student, Efthymios Tzinis, submitted this Dissertation for approval on 2023-04-20 at 08:07.","This Dissertation was approved for publication on 2023-04-24 at 09:35.","DSpace SAF Submission Ingestion Package generated from Vireo submission #19058 on 2023-09-01 at 17:09:05","In this thesis, we tackle the problem of training a machine to perceive, disentangle and reconstruct independent source waveforms from a given mixture audio recording without the need of explicit supervision. The problem becomes even more apparent with current state-of-the-art approaches which are largely dependent on the existence of vast amounts of carefully curated data. In essence, this thesis presents a holistic approach on how people can develop sound separation algorithms based on neural networks which are able to scale up to multiple users, modalities and datasets without the need of annotated data. To that end, the contributions of this thesis is threefold. The first part of the thesis describes novel unsupervised and self-supervised algorithms for sound source separation problems under a wide spectrum of environmental setups. The second part aims to expand the applicability of sound separation systems using external condition information (e.g. video, text and other semantic discriminative concepts) which consists of the multi-modal aspect of this work. Finally, the last chapter presents potential obstacles towards the deployment of the aforementioned algorithms (e.g. scarcity of labels, lack of data on the same device during training, limited computational resources, reluctance of the users to share their private data, erroneous predictions, etc.) as well as proposes novel solutions which can be seamlessly integrated into their real-world implementations."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Unsupervised sound separation"]}]}],"canonical_facts":{"dc:contributor":["Smaragdis, Paris","Hasegawa-Johnson, Mark","Misailovic, Sasa","Hershey, John R"],"dc:creator":["Tzinis, Efthymios"],"dc:date":["2023-05","2023-04-24"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2023-09-01 without embargo terms","The student, Efthymios Tzinis, accepted the attached license on 2023-04-20 at 07:58.","The student, Efthymios Tzinis, submitted this Dissertation for approval on 2023-04-20 at 08:07.","This Dissertation was approved for publication on 2023-04-24 at 09:35.","DSpace SAF Submission Ingestion Package generated from Vireo submission #19058 on 2023-09-01 at 17:09:05","In this thesis, we tackle the problem of training a machine to perceive, disentangle and reconstruct independent source waveforms from a given mixture audio recording without the need of explicit supervision. The problem becomes even more apparent with current state-of-the-art approaches which are largely dependent on the existence of vast amounts of carefully curated data. In essence, this thesis presents a holistic approach on how people can develop sound separation algorithms based on neural networks which are able to scale up to multiple users, modalities and datasets without the need of annotated data. To that end, the contributions of this thesis is threefold. The first part of the thesis describes novel unsupervised and self-supervised algorithms for sound source separation problems under a wide spectrum of environmental setups. The second part aims to expand the applicability of sound separation systems using external condition information (e.g. video, text and other semantic discriminative concepts) which consists of the multi-modal aspect of this work. Finally, the last chapter presents potential obstacles towards the deployment of the aforementioned algorithms (e.g. scarcity of labels, lack of data on the same device during training, limited computational resources, reluctance of the users to share their private data, erroneous predictions, etc.) as well as proposes novel solutions which can be seamlessly integrated into their real-world implementations."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/120294"],"dc:language":["en","eng"],"dc:rights":["Copyright 2023 Efthymios Tzinis"],"dc:subject":["Sound Separation","Audio-visual Perception","Unsupervised Learning","Self-supervised Learning","Speech Enhancement","Efficient Neural Networks","Federated Learning"],"dc:title":["Unsupervised sound separation"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:57Z"}