University of Toronto
Bad Data, Good Representations: Transferable Embeddings of Heterogeneous, Unordered, and Incomplete Datasets
Abstract
dc:description.abstractThe challenge of enabling neural systems to effectively process complex, irregular data structures remains largely unresolved. While recent advances in deep learning have delivered remarkable success in domains with well-structured inputs--such as computer vision and natural language processing--many real-world datasets continue to pose persistent challenges. Data originating from domains like healthcare, the Internet of Things, or the financial sector often appear in messy forms: tabular, multimodal, variably structured, riddled with missing values, and often lacking any natural ordering. This thesis asks whether we can still learn good representations from such ``bad data''. To answer this, the thesis makes three core contributions, each addressing a distinct category of ``bad data.'' First, it introduces a novel method for contrastive representation learning in tabular data, treating columns as modalities, to derive column-wise embeddings that scale efficiently and capture inter-variable relationships. Second, it proposes an expressive framework for learning from set-structured data via newly introduced quasi-arithmetic neural networks, which replace classical algebraic set-pooling operations by learnable neural functions. Third, it extends technique known as arbitrary conditioning to the realm of multimodal generative modeling through a Wasserstein Autoencoder, resulting in a model capable of handling missing modalities, allowing flexible inference, and modality translation. Extensive experiments demonstrate that the proposed models consistently match or surpass state-of-the-art baselines.They produce embeddings with broad multi-task transferability, supporting supervised, unsupervised, and generative tasks alike. Moreover, the models are data-agnostic, allows efficient adaptation to a wide range of data types and structures. By embracing unorderedness, incompleteness, and heterogeneity not as flaws but as inherent attributes to be modeled, this work paves the way for more robust, flexible, and general-purpose neural systems.
Degree
thesis:*- Department dc:contributor.department
- Mechanical and Industrial Engineering
- Year dc:date.issued
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Tokar, Tomas
- Advisor dc:contributor.advisor
-
- Sanner, Scott
Subjects
dc:subject × 2Rights
dc:rights- Statement dc:rights
-
- Attribution 4.0 International
- Licence dc:rights.uri
Identifiers
dc:identifier.*- Handle dc:identifier.uri
- https://hdl.handle.net/1807/150476
- OAI identifier oai:identifier
- oai:utoronto.scholaris.ca:1807/150476