National University of Singapore
EFFICIENT DATA CURATION AND UTILIZATION FOR DEEP LEARNING
Abstract
dc:description.abstractThis thesis studies practical methods to improve the training and construction efficiency of large-scale vision datasets, aiming to reduce computational and annotation costs. We propose InfoBatch, an unbiased dynamic data pruning framework that losslessly accelerates training and saves 20–40% of computation across diverse vision tasks. To address dataset quality in large-scale and multi-modal settings, we introduce InfoGrowth, an efficient online algorithm for data cleaning and selection that maintains cleanliness and diversity as data grow. Finally, we present Info-Coevolution, a bias-free framework for online selective annotation that enables models and data to coevolve, reducing annotation and training costs by 32–50% on ImageNet-1K without performance degradation. Together, these approaches significantly lower the cost of training and building large-scale datasets in real-world applications.
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- QIN ZIHENG