{"id":{"repo_id":"toronto-retro","oai_identifier":"oai:utoronto.scholaris.ca:1807/150476"},"canonical_url":"https://search.dev.ndltd.org/etd/toronto-retro/oai:utoronto.scholaris.ca:1807/150476","repository":{"repo_id":"toronto-retro","name":"University of Toronto","base_url":"https://utoronto.scholaris.ca/server/oai/request"},"display":{"title":"Bad Data, Good Representations: Transferable Embeddings of Heterogeneous, Unordered, and Incomplete Datasets","abstract":"The challenge of enabling neural systems to effectively process complex, irregular data structures remains largely unresolved. While recent advances in deep learning have delivered remarkable success in domains with well-structured inputs--such as computer vision and natural language processing--many real-world datasets continue to pose persistent challenges. Data originating from domains like healthcare, the Internet of Things, or the financial sector often appear in messy forms: tabular, multimodal, variably structured, riddled with missing values, and often lacking any natural ordering. This thesis asks whether we can still learn good representations from such ``bad data''. To answer this, the thesis makes three core contributions, each addressing a distinct category of ``bad data.'' First, it introduces a novel method for contrastive representation learning in tabular data, treating columns as modalities, to derive column-wise embeddings that scale efficiently and capture inter-variable relationships. Second, it proposes an expressive framework for learning from set-structured data via newly introduced quasi-arithmetic neural networks, which replace classical algebraic set-pooling operations by learnable neural functions. Third, it extends technique known as arbitrary conditioning to the realm of multimodal generative modeling through a Wasserstein Autoencoder, resulting in a model capable of handling missing modalities, allowing flexible inference, and modality translation. Extensive experiments demonstrate that the proposed models consistently match or surpass state-of-the-art baselines.They produce embeddings with broad multi-task transferability, supporting supervised, unsupervised, and generative tasks alike. Moreover, the models are data-agnostic, allows efficient adaptation to a wide range of data types and structures. By embracing unorderedness, incompleteness, and heterogeneity not as flaws but as inherent attributes to be modeled, this work paves the way for more robust, flexible, and general-purpose neural systems.","abstract_html":"The challenge of enabling neural systems to effectively process complex, irregular data structures remains largely unresolved. While recent advances in deep learning have delivered remarkable success in domains with well-structured inputs--such as computer vision and natural language processing--many real-world datasets continue to pose persistent challenges. Data originating from domains like healthcare, the Internet of Things, or the financial sector often appear in messy forms: tabular, multimodal, variably structured, riddled with missing values, and often lacking any natural ordering. This thesis asks whether we can still learn good representations from such ``bad data&#x27;&#x27;. To answer this, the thesis makes three core contributions, each addressing a distinct category of ``bad data.&#x27;&#x27; First, it introduces a novel method for contrastive representation learning in tabular data, treating columns as modalities, to derive column-wise embeddings that scale efficiently and capture inter-variable relationships. Second, it proposes an expressive framework for learning from set-structured data via newly introduced quasi-arithmetic neural networks, which replace classical algebraic set-pooling operations by learnable neural functions. Third, it extends technique known as arbitrary conditioning to the realm of multimodal generative modeling through a Wasserstein Autoencoder, resulting in a model capable of handling missing modalities, allowing flexible inference, and modality translation. Extensive experiments demonstrate that the proposed models consistently match or surpass state-of-the-art baselines.They produce embeddings with broad multi-task transferability, supporting supervised, unsupervised, and generative tasks alike. Moreover, the models are data-agnostic, allows efficient adaptation to a wide range of data types and structures. By embracing unorderedness, incompleteness, and heterogeneity not as flaws but as inherent attributes to be modeled, this work paves the way for more robust, flexible, and general-purpose neural systems.","abstract_has_math":false,"creators":["Tokar, Tomas"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":"Mechanical and Industrial Engineering","school":null,"contributors":[],"advisors":["Sanner, Scott"],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-10","date_published":"2025-10","updated_at":"2026-07-27T21:28:16Z","subjects":["generative modeling","representation learning"],"languages":[],"rights":["Attribution 4.0 International"],"rights_urls":["http://creativecommons.org/licenses/by/4.0/"],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/1807/150476","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Sanner, Scott"]},{"key":"dc:contributor.department","label":"Department","values":["Mechanical and Industrial Engineering"]},{"key":"dc:creator","label":"Author","values":["Tokar, Tomas"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-10"]},{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2025-12-01T16:47:05Z"]},{"key":"dc:date.issued","label":"Date","values":["2025-10"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["generative modeling","representation learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["Attribution 4.0 International"]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://creativecommons.org/licenses/by/4.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/1807/150476"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["The challenge of enabling neural systems to effectively process complex, irregular data structures remains largely unresolved. While recent advances in deep learning have delivered remarkable success in domains with well-structured inputs--such as computer vision and natural language processing--many real-world datasets continue to pose persistent challenges. Data originating from domains like healthcare, the Internet of Things, or the financial sector often appear in messy forms: tabular, multimodal, variably structured, riddled with missing values, and often lacking any natural ordering. This thesis asks whether we can still learn good representations from such ``bad data''. To answer this, the thesis makes three core contributions, each addressing a distinct category of ``bad data.'' First, it introduces a novel method for contrastive representation learning in tabular data, treating columns as modalities, to derive column-wise embeddings that scale efficiently and capture inter-variable relationships. Second, it proposes an expressive framework for learning from set-structured data via newly introduced quasi-arithmetic neural networks, which replace classical algebraic set-pooling operations by learnable neural functions. Third, it extends technique known as arbitrary conditioning to the realm of multimodal generative modeling through a Wasserstein Autoencoder, resulting in a model capable of handling missing modalities, allowing flexible inference, and modality translation. Extensive experiments demonstrate that the proposed models consistently match or surpass state-of-the-art baselines.They produce embeddings with broad multi-task transferability, supporting supervised, unsupervised, and generative tasks alike. Moreover, the models are data-agnostic, allows efficient adaptation to a wide range of data types and structures. By embracing unorderedness, incompleteness, and heterogeneity not as flaws but as inherent attributes to be modeled, this work paves the way for more robust, flexible, and general-purpose neural systems."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["Ph.D."]},{"key":"dc:title","label":"Title","values":["Bad Data, Good Representations: Transferable Embeddings of Heterogeneous, Unordered, and Incomplete Datasets"]}]}],"canonical_facts":{"dc:contributor.advisor":["Sanner, Scott"],"dc:contributor.department":["Mechanical and Industrial Engineering"],"dc:creator":["Tokar, Tomas"],"dc:date":["2025-10"],"dc:date.accessioned":["2025-12-01T16:47:05Z"],"dc:date.issued":["2025-10"],"dc:description.abstract":["The challenge of enabling neural systems to effectively process complex, irregular data structures remains largely unresolved. While recent advances in deep learning have delivered remarkable success in domains with well-structured inputs--such as computer vision and natural language processing--many real-world datasets continue to pose persistent challenges. Data originating from domains like healthcare, the Internet of Things, or the financial sector often appear in messy forms: tabular, multimodal, variably structured, riddled with missing values, and often lacking any natural ordering. This thesis asks whether we can still learn good representations from such ``bad data''. To answer this, the thesis makes three core contributions, each addressing a distinct category of ``bad data.'' First, it introduces a novel method for contrastive representation learning in tabular data, treating columns as modalities, to derive column-wise embeddings that scale efficiently and capture inter-variable relationships. Second, it proposes an expressive framework for learning from set-structured data via newly introduced quasi-arithmetic neural networks, which replace classical algebraic set-pooling operations by learnable neural functions. Third, it extends technique known as arbitrary conditioning to the realm of multimodal generative modeling through a Wasserstein Autoencoder, resulting in a model capable of handling missing modalities, allowing flexible inference, and modality translation. Extensive experiments demonstrate that the proposed models consistently match or surpass state-of-the-art baselines.They produce embeddings with broad multi-task transferability, supporting supervised, unsupervised, and generative tasks alike. Moreover, the models are data-agnostic, allows efficient adaptation to a wide range of data types and structures. By embracing unorderedness, incompleteness, and heterogeneity not as flaws but as inherent attributes to be modeled, this work paves the way for more robust, flexible, and general-purpose neural systems."],"dc:description.degree":["Ph.D."],"dc:identifier.uri":["https://hdl.handle.net/1807/150476"],"dc:rights":["Attribution 4.0 International"],"dc:rights.uri":["http://creativecommons.org/licenses/by/4.0/"],"dc:subject":["generative modeling","representation learning"],"dc:title":["Bad Data, Good Representations: Transferable Embeddings of Heterogeneous, Unordered, and Incomplete Datasets"],"dc:type":["Thesis"]},"updated_at":"2026-07-27T21:28:16Z"}