National University of Singapore
EFFECTIVE TRAINING OF NEURAL NETWORKS FOR BETTER GENERALIZATION
Abstract
dc:description.abstractDeep learning has achieved remarkable success, yet training deep neural networks remains costly, unstable, and poorly understood in terms of generalization. This thesis aims to make training more efficient and generalization-aware, addressing from optimization and data perspectives. From the optimization side, we develop practical algorithms that improve convergence speed and stability. We propose DRAG, a dimension-reduced adaptive gradient method that unifies the benefits of SGD and Adam, and a memory-efficient Shampoo using 4-bit Cholesky quantization with error feedback, enabling scalable second-order training. From the data side, we investigate why data-centric strategies enhance generalization. We analyze semi-supervised learning and data augmentation through the lens of feature learning, uncovering how they promote semantic diversity and robustness, and propose improved variants such as SA-FixMatch.
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- LI JINGYANG