{"id":{"repo_id":"nus","oai_identifier":"oai:scholarbank.nus.edu.sg:10635/310424"},"canonical_url":"https://search.dev.ndltd.org/etd/nus/oai:scholarbank.nus.edu.sg:10635/310424","repository":{"repo_id":"nus","name":"National University of Singapore","base_url":"https://scholarbank.nus.edu.sg/oai/request"},"display":{"title":"TOWARDS DATA-EFFICIENT DEEP LEARNING","abstract":"This thesis advances data-efficient machine learning by tackling the limitations of current dataset distillation (DD) methods, which aim to compress large datasets into compact synthetic ones for faster training and enhanced privacy. First, it introduces Dataset Factorization, a novel framework that decomposes datasets into learnable bases and a hallucination network, significantly improving performance and establishing a new research paradigm. Second, it proposes the first slimmable dataset condensation method, enabling flexible dataset sizing without re-accessing the original data, thus supporting varying storage budgets. Third, to reduce the high latency of DD, the thesis pioneers Few-Shot Dataset Distillation, using a translator network to refine distilled datasets efficiently. Finally, it introduces pre-training for DD with a meta-learning scheme, accelerating DD by up to tenfold while maintaining quality. Collectively, these innovations push the boundaries of data efficiency, making model training faster, more flexible, and scalable across storage and computational constraints.","abstract_html":"This thesis advances data-efficient machine learning by tackling the limitations of current dataset distillation (DD) methods, which aim to compress large datasets into compact synthetic ones for faster training and enhanced privacy. First, it introduces Dataset Factorization, a novel framework that decomposes datasets into learnable bases and a hallucination network, significantly improving performance and establishing a new research paradigm. Second, it proposes the first slimmable dataset condensation method, enabling flexible dataset sizing without re-accessing the original data, thus supporting varying storage budgets. Third, to reduce the high latency of DD, the thesis pioneers Few-Shot Dataset Distillation, using a translator network to refine distilled datasets efficiently. Finally, it introduces pre-training for DD with a meta-learning scheme, accelerating DD by up to tenfold while maintaining quality. Collectively, these innovations push the boundaries of data efficiency, making model training faster, more flexible, and scalable across storage and computational constraints.","abstract_has_math":false,"creators":["LIU SONGHUA"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-03-18","date_published":"2025-03-18","updated_at":"2026-07-24T03:31:26Z","subjects":["Data Compression","Generative Modeling","Synthetic Data","Data Synthesis","Dataset Distillation","Data-Efficient Machine Learning"],"languages":[],"rights":[],"rights_urls":["https://scholarbank.nus.edu.sg/bitstreams/aa18bf99-f7ad-4e2e-b789-37dde56129f3/download"],"identifier_entries":[]},"links":{"outbound_url":null,"outbound_label":null,"outbound_source":null},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["LIU SONGHUA"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2025-03-18"]},{"key":"dc:relation.isreferencedby","label":"Dc Relation Isreferencedby","values":["https://scholarbank.nus.edu.sg/handle/10635/310424"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Data Compression","Generative Modeling","Synthetic Data","Data Synthesis","Dataset Distillation","Data-Efficient Machine Learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["https://scholarbank.nus.edu.sg/bitstreams/aa18bf99-f7ad-4e2e-b789-37dde56129f3/download"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://scholarbank.nus.edu.sg/bitstreams/9353c2f6-1b10-46bc-a2a9-5f78c091be88/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["This thesis advances data-efficient machine learning by tackling the limitations of current dataset distillation (DD) methods, which aim to compress large datasets into compact synthetic ones for faster training and enhanced privacy. First, it introduces Dataset Factorization, a novel framework that decomposes datasets into learnable bases and a hallucination network, significantly improving performance and establishing a new research paradigm. Second, it proposes the first slimmable dataset condensation method, enabling flexible dataset sizing without re-accessing the original data, thus supporting varying storage budgets. Third, to reduce the high latency of DD, the thesis pioneers Few-Shot Dataset Distillation, using a translator network to refine distilled datasets efficiently. Finally, it introduces pre-training for DD with a meta-learning scheme, accelerating DD by up to tenfold while maintaining quality. Collectively, these innovations push the boundaries of data efficiency, making model training faster, more flexible, and scalable across storage and computational constraints."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["9f3c6f2aa8f96291ef77b0a18484521d","cad1c8e18c8148174cd01e9280ba70c4","898d745fc1b150ec4ab7f69bf603b270"]},{"key":"dc:title","label":"Title","values":["TOWARDS DATA-EFFICIENT DEEP LEARNING"]}]}],"canonical_facts":{"dc:creator":["LIU SONGHUA"],"dc:date.issued":["2025-03-18"],"dc:description.abstract":["This thesis advances data-efficient machine learning by tackling the limitations of current dataset distillation (DD) methods, which aim to compress large datasets into compact synthetic ones for faster training and enhanced privacy. First, it introduces Dataset Factorization, a novel framework that decomposes datasets into learnable bases and a hallucination network, significantly improving performance and establishing a new research paradigm. Second, it proposes the first slimmable dataset condensation method, enabling flexible dataset sizing without re-accessing the original data, thus supporting varying storage budgets. Third, to reduce the high latency of DD, the thesis pioneers Few-Shot Dataset Distillation, using a translator network to refine distilled datasets efficiently. Finally, it introduces pre-training for DD with a meta-learning scheme, accelerating DD by up to tenfold while maintaining quality. Collectively, these innovations push the boundaries of data efficiency, making model training faster, more flexible, and scalable across storage and computational constraints."],"dc:format.checksum.md5":["9f3c6f2aa8f96291ef77b0a18484521d","cad1c8e18c8148174cd01e9280ba70c4","898d745fc1b150ec4ab7f69bf603b270"],"dc:identifier.uri":["https://scholarbank.nus.edu.sg/bitstreams/9353c2f6-1b10-46bc-a2a9-5f78c091be88/download"],"dc:relation.isreferencedby":["https://scholarbank.nus.edu.sg/handle/10635/310424"],"dc:rights":["https://scholarbank.nus.edu.sg/bitstreams/aa18bf99-f7ad-4e2e-b789-37dde56129f3/download"],"dc:subject":["Data Compression","Generative Modeling","Synthetic Data","Data Synthesis","Dataset Distillation","Data-Efficient Machine Learning"],"dc:title":["TOWARDS DATA-EFFICIENT DEEP LEARNING"],"dc:type":["Thesis"]},"updated_at":"2026-07-24T03:31:26Z"}