Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 20 of 24 for “"Distributed Training"”.
-
Distributed Training with Heterogeneous Data: Bridging Median- and Mean-Based Algorithms
… in the study of median-based algorithms for distributed non-convex optimization. Two prominent examples include signSGD with majority vote, an effective approach for communication reduction via 1-bit compression on the local gradients, and medianSGD, an algorithm recently proposed to ensure …
-
Comparison of distributed training architecture for convolutional neural network in cloud
… natural language processing and data mining. The training process of the DNN is a computationally intensive application that can be accelerated by parallel computing devices such as graphic processing units (GPUs) and field programmable gate arrays (FP- GAs). However, sometimes the amount of the …
-
Trustworthy Federated Learning Systems: From Secure Distributed Training to Reliable Fine-tuning
Training and fine-tuning Deep Learning (DL) models require vast amounts of domain-specific data. However, this data is often distributed across different organizations and devices, restricted from direct sharing by privacy, ownership, and regulatory constraints. While Federated Learning (FL) …
-
Robust Machine Learning Against Faults in Micro-Controllers and Stragglers in Distributed Training on the Cloud
… ML models onto smaller devices and efficiently training more powerful ML models in parallel in different distributed system topologies have drawn interests. This thesis studies the robustness of ML models in the two scenarios when deployed on portable micro-controller units and while being …
-
GraphDHT: Scaling Graph Neural Networks' Distributed Training on Edge Devices on a Peer-to-Peer Distributed Hash Table Network
This thesis presents an innovative strategy for distributed Graph Neural Network (GNN) training, leveraging a peer-to-peer network of heterogeneous edge devices interconnected through a Distributed Hash Table (DHT). As GNNs become increasingly vital in analyzing graph-structured data across various …
-
Distributed On-line Training for Object Detection on Embedded Devices.
In this thesis, we develop a scalable distributed approach for object detection model training and inference, using low cost embedded devices. Examples of the usage of such an approach is automatic beach litter collection using mobile robots, IoT based intelligent video surveillance system for …
-
Adaptive parameter servers
… and recommender systems. For large ML tasks, distributed training has become a necessity for keeping up with increasing dataset sizes and model complexity. A key challenge in distributed training is to synchronize model parameters among cluster nodes and to do so efficiently. Parameter servers …
-
Accelerating distributed neural network training with network-centric approach
Distributed training of Deep Neural Networks (DNN) is an important technique to reduce the training time of large DNNs for a wide range of applications. In existing distributed training approaches, however, the communication time to periodically exchange parameters (i.e., weights) and gradients …
-
Network Requirements for Distributed Machine Learning Training in the Cloud
… characterize the impact of network bandwidth on distributed machine learning training. I test four popular machine learning models (ResNet, DenseNet, VGG, and BERT) on an Nvidia A-100 cluster to determine the impact of bursty and non-bursty cross traffic (such as web-search traffic and long-lived …
-
TopoOpt: Co-optimizing Network Topology and Parallelization Strategy for Distributed Machine Learning Training Jobs
… a novel approach for building direct-connect DNN training clusters. The proposed system, called TopoOpt, co-optimizes the distributed training process across three dimensions: computation, communication, and network topology. TopoOpt uses a novel alternating optimization technique and a group …
-
Toward communication-efficient and secure distributed machine learning
… recent years, there is an increasing interest in distributed machine learning. On one hand, distributed machine learning is motivated by assigning the training workload to multiple devices for acceleration and better throughput. On the other hand, there are machine-learning tasks requiring …
-
ACCELERATING DNN INFERENCE AND TRAINING IN DISTRIBUTED SYSTEMS
… complex, which introduces extremely long DNN training time. Although DNN inference typically runs a single round of forward propagation on the DNN model, the inference time for large DNN models is still intolerable for time-sensitive applications in mobile devices. We propose to reduce the DNN …
-
LARGE-SCALE NEURAL 3D SCENE RECONSTRUCTION, RENDERING, AND BEYOND
… imperfect camera poses, and consistency across distributed captures. To address these, we introduce a series of innovations. For scalability, AdaSfM enables distributed camera pose estimation by adapting to scene complexity. Building on this, DOGS, a distributed training framework for 3DGS, …
-
Olfactory conditioned ejaculatory preference in the male rat : implications for the role of learning in sexual partner preferences
… receiving explicitly-unpaired or randomly-paired training failed to display this preference, implicating classical conditioning mechanisms in the development of this behavior. Examination of the time course of the development of CEP found that it develops rapidly, demonstrating the importance of …
-
Table manners at scale: introducing Lora Lens for efficient analysis & detoxification of language models
… toxicity without requiring complete model retraining. Despite its effectiveness, Attention Lens faces severe limitations in scalability and computational efficiency. To overcome these limitations, we propose and implement the Lora Lens, an innovative adaptation of Attention Lens employing …
-
The own-group bias in face processing: the effect of training on recognition performance
… a way to decrease or eliminate it, through training. Reporting five studies involving memory or matching tasks, the aim of the present thesis was to develop and to explore to what extent training can decrease or remove the OGB. French White participants, and South African White, Black and …
-
Interpretability and Debugging for Distributed Privacy Preserving Machine Learning
… systems increasingly rely on privacy-preserving distributed training to leverage sensitive data across multiple organizations without centralization. Federated Learning (FL), a distributed privacy-preserving machine learning paradigm, enables hospitals, devices, and enterprises to collaboratively …
-
Distributed practice for cardiopulmonary resuscitation (CPR) training: improving educational efficiency and cost-effectiveness in clinical settings
… outcomes of the victims. Despite annual training, healthcare providers struggle to conduct guideline compliant CPR during the management of cardiac arrests. Increased likelihood of survival from cardiac arrest depends upon the integration of medical science, educational efficiency and …
-
Decentralized Machine Learning over Fragmented Data
… Federated Learning (FL), which enables model training over distributed data. However, FL’s dependence on central coordination and inflexibility with heterogeneous systems limit its applicability in healthcare settings. Motivated by these challenges, I explore the following three core themes: …
-
Dora: QoE-aware hybrid parallelism for distributed edge AI
… capacity of individual devices, necessitating distributed execution across heterogeneous devices over variable and contention-prone networks. Existing planners for hybrid (e.g., data and pipeline) parallelism largely optimize for throughput or device utilization, overlooking QoE, leading to …
Page 1 of 2