University of Illinois at Urbana-Champaign
Toward communication-efficient and secure distributed machine learning
Abstract
dc:description"In recent years, there is an increasing interest in distributed machine learning. On one hand, distributed machine learning is motivated by assigning the training workload to multiple devices for acceleration and better throughput. On the other hand, there are machine-learning tasks requiring distributed training locally on remote devices due to privacy concerns. Stochastic Gradient Descent (SGD) and its variants are commonly used for training large-scale deep neural networks, as well as the distributed training. Unlike machine learning on a single device, distributed machine learning requires collaboration and communication among the devices, which incur the new challenges: 1) the heavy communication overhead can be the bottleneck that slows down the training; 2) the unreliable communication and weaker control over the remote entities makes the distributed system vulnerable to systematic failures and malicious attacks. In this dissertation, we aim to find new approaches to make distributed SGD faster and more secure. We present four main parts of research. We first study approaches for reducing the communication overhead, including message compression and infrequent synchronization. Then, we investigate the possibility of combining asynchrony with infrequent synchronization. To address security in distributed SGD, we study the tolerance to Byzantine failures. Finally, we explore the possibility of combining both communication efficiency and security techniques into one distributed learning system. Specifically, we present the following techniques to improve the communication efficiency and security of distributed SGD: 1) a technique called ""error reset"" to adapt both infrequent synchronization and message compression to distributed SGD, to reduce the communication overhead; 2) federated optimization in asynchronous mode; 3) a framework of score-based approaches for Byzantine tolerance in distributed SGD; 4) a distributed learning system integrating all these three techniques. The proposed system provides communication reduction, both synchronous and asynchronous training, and Byzantine tolerance, with both theoretical guarantees and empirical evaluations."
Degree
thesis:*- Name thesis:degree_name
- Ph.D.
- Level thesis:degree_level
- Dissertation
- Discipline thesis:degree_discipline
- Computer Science
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2021
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Xie, Cong
- Contributors dc:contributor
-
- Koyejo, Oluwasanmi
- Gupta, Indranil
- Raginsky, Maxim
- McMahan, Hugh Brendan
Subjects
dc:subject × 2Rights
dc:rights- Statement dc:rights
-
- Copyright 2021 Cong Xie
- Language dc:language
- en
Identifiers
dc:identifier.*- Handle dc:identifier
- http://hdl.handle.net/2142/110466
- OAI identifier oai:identifier
- oai:www.ideals.illinois.edu:2142/110466