Back to results

University of Toronto

Bad Data, Good Representations: Transferable Embeddings of Heterogeneous, Unordered, and Incomplete Datasets

Abstract

dc:description.abstract

The challenge of enabling neural systems to effectively process complex, irregular data structures remains largely unresolved. While recent advances in deep learning have delivered remarkable success in domains with well-structured inputs--such as computer vision and natural language processing--many real-world datasets continue to pose persistent challenges. Data originating from domains like healthcare, the Internet of Things, or the financial sector often appear in messy forms: tabular, multimodal, variably structured, riddled with missing values, and often lacking any natural ordering. This thesis asks whether we can still learn good representations from such ``bad data''. To answer this, the thesis makes three core contributions, each addressing a distinct category of ``bad data.'' First, it introduces a novel method for contrastive representation learning in tabular data, treating columns as modalities, to derive column-wise embeddings that scale efficiently and capture inter-variable relationships. Second, it proposes an expressive framework for learning from set-structured data via newly introduced quasi-arithmetic neural networks, which replace classical algebraic set-pooling operations by learnable neural functions. Third, it extends technique known as arbitrary conditioning to the realm of multimodal generative modeling through a Wasserstein Autoencoder, resulting in a model capable of handling missing modalities, allowing flexible inference, and modality translation. Extensive experiments demonstrate that the proposed models consistently match or surpass state-of-the-art baselines.They produce embeddings with broad multi-task transferability, supporting supervised, unsupervised, and generative tasks alike. Moreover, the models are data-agnostic, allows efficient adaptation to a wide range of data types and structures. By embracing unorderedness, incompleteness, and heterogeneity not as flaws but as inherent attributes to be modeled, this work paves the way for more robust, flexible, and general-purpose neural systems.

Degree

thesis:*
Department dc:contributor.department
Mechanical and Industrial Engineering
Year dc:date.issued
2025

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Tokar, Tomas
Advisor dc:contributor.advisor
  • Sanner, Scott

Subjects

dc:subject × 2

Rights

dc:rights
Statement dc:rights
  • Attribution 4.0 International

Identifiers

dc:identifier.*
Handle dc:identifier.uri
https://hdl.handle.net/1807/150476
OAI identifier oai:identifier
oai:utoronto.scholaris.ca:1807/150476

Chain of custody

source
Harvested from
University of Toronto
Base URL
utoronto.scholaris.ca/server/oai/request
Last updated
2026-07-27
Source record
OAI-PMH GetRecord
citation

Tokar, Tomas. Bad Data, Good Representations: Transferable Embeddings of Heterogeneous, Unordered, and Incomplete Datasets. 2025. https://hdl.handle.net/1807/150476