Back to results

The University of Western Ontario

Contrastive Learning of Auditory Representations

Abstract

dc:description.abstract

Learning rich visual representations using contrastive self-supervised learning has been extremely successful. However, it is still a major question whether we could use a similar approach to learn more efficient auditory and audio-visual representations. In this thesis, we expand on prior self-supervised methods to learn better auditory and audio-visual representations. We introduce various data augmentations suitable for auditory and audio-visual data and evaluate their impact on predictive performance, and demonstrate that training with both supervised and contrastive losses simultaneously improves the learned representations compared to self-supervised pre-training followed by supervised fine-tuning. We illustrate that by combining all these methods and with substantially less labeled data, our framework achieves significant improvement on prediction performance compared to the supervised approach. Moreover, compared to the self-supervised approach, our framework converges faster with significantly better representations.

Degree

thesis:*
Name thesis:degree_name
M Sc
Discipline thesis:degree_discipline
Computer Science
Grantor dc:publisher
The University of Western Ontario
Year dc:date.issued
2021

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Al-Tahan, Haider
Advisor dc:contributor.advisor
  • Mohsenzadeh, Yalda

Subjects

dc:subject × 3

Rights

Language dc:language.iso
en_ca

Identifiers

dc:identifier.*
OAI identifier oai:identifier
oai:uwo.scholaris.ca:20.500.14721/31151

Chain of custody

source
Harvested from
Western University
Base URL
uwo.scholaris.ca/server/oai/request
Last updated
2026-07-27
Source record
OAI-PMH GetRecord
citation

Al-Tahan, Haider. Contrastive Learning of Auditory Representations. The University of Western Ontario, 2021. https://hdl.handle.net/20.500.14721/31151