University of Illinois at Urbana-Champaign
Efficient and robust algorithms for training machine learning models
Abstract
dc:descriptionDeep Learning (DL) models have been widely successful at solving many large-scale and challenging tasks. However, to achieve state-of-the-art performance, these models need to be extremely large, and they need to be trained with a massive amount of data. Therefore, the best DL models are computing-resource and data-hungry. This makes training such models very expensive and sometimes prohibitively so. Detrimentally, this can also lead to a high energy usage. With this as a motivation, we study whether we can improve the computational and sample complexities of Machine Learning (ML) training algorithms for a few specific problems. The computational complexity of an algorithm characterizes the number of computational operations required to run that algorithm in terms of the problem size and parameters. In the first part of this dissertation, we investigate algorithms for solving minimax optimization and constrained minimization problems. These problems have several applications in modern ML. We propose improved algorithms for solving these and prove that they achieve better computational complexity than baseline algorithms. Sample complexity of a learning algorithm and model characterizes its achievable testing loss for a given number of potentially noisy samples, i.e.~its statistical efficiency. In the second part of the dissertation, we investigate the statistical efficiency of solving some modern DL tasks. First, we propose an architecture and loss to learn unbiased conditional generative adversarial networks using noisy labeled samples and then characterize their sample complexity in terms of the noise level in the labels. Next, we study the problem of meta-representation learning of many related tasks under a few-shot learning regime, where only very few samples are available per task. We prove that a recently popular DL algorithm can faithfully learn a linear meta-representation for regression tasks with very few samples each.
Degree
thesis:*- Name thesis:degree_name
- Ph.D.
- Level thesis:degree_level
- Dissertation
- Discipline thesis:degree_discipline
- Electrical & Computer Engr
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2022
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Thekumparampil, Kiran Koshy
- Contributors dc:contributor
-
- Oh, Sewoong
- Hajek, Bruce
- Srikant, Rayadurgam
- Sun, Ruoyu
Subjects
dc:subject × 4Rights
dc:rights- Statement dc:rights
-
- Copyright 2022 Kiran Koshy Thekumparampil
- Language dc:language
- en, eng
Identifiers
dc:identifier.*- Handle dc:identifier
- https://hdl.handle.net/2142/117818