Abstract
dc:description.abstractThis thesis advances data-efficient machine learning by tackling the limitations of current dataset distillation (DD) methods, which aim to compress large datasets into compact synthetic ones for faster training and enhanced privacy. First, it introduces Dataset Factorization, a novel framework that decomposes datasets into learnable bases and a hallucination network, significantly improving performance and establishing a new research paradigm. Second, it proposes the first slimmable dataset condensation method, enabling flexible dataset sizing without re-accessing the original data, thus supporting varying storage budgets. Third, to reduce the high latency of DD, the thesis pioneers Few-Shot Dataset Distillation, using a translator network to refine distilled datasets efficiently. Finally, it introduces pre-training for DD with a meta-learning scheme, accelerating DD by up to tenfold while maintaining quality. Collectively, these innovations push the boundaries of data efficiency, making model training faster, more flexible, and scalable across storage and computational constraints.
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- LIU SONGHUA