University of Missouri--Kansas City
Resource efficient distributed inference of deep neural networks for Edge AI
Abstract
dc:description.abstractDeep Neural Networks (DNNs) have become the cornerstone of modern artificial intelligence, powering applications such as image recognition, language understanding, and multimodal reasoning. However, executing these models efficiently on edge devices remains a major challenge due to their high computational, energy, and bandwidth demands. This dissertation presents a unified framework for resource-efficient distributed inference of DNNs in Edge AI, built upon three tightly coupled research objectives: asynchronous split inference, adaptive batching for resource efficiency, and tensor compression at the client–server boundary. First, the study introduces an asynchronous split inference paradigm, where model computation is partitioned between clients and servers, and communication is overlapped with computation. The proposed framework, ASAP (Asynchronous Split Inference for Accelerated DNN Execution), significantly reduces idle time and improves overall throughput. Second, the work explores input batching and slice batching strategies that maximize hardware utilization and minimize inference latency across heterogeneous clients. The batching framework adapts dynamically to network and device variations, enabling balanced resource sharing in distributed environments. Third, the dissertation proposes a tensor compression mechanism that reduces data transfer volume between clients and servers by encoding and decoding intermediate activations efficiently, thereby mitigating communication bottlenecks without degrading model accuracy. Comprehensive experiments are conducted using advanced vision architectures such as Vision Transformer (ViT), Swin Transformer, DenseNet, and ResNet under diverse deployment scenarios. Results demonstrate up to 67% reduction in total inference time, alongside improvements in GPU utilization and energy efficiency. The integration of asynchronous scheduling, batching, and compression enables scalable, low-latency, and adaptive inference for real-world edge applications. Collectively, this research advances the frontier of distributed deep learning by bridging computation and communication in heterogeneous edge--server systems. It provides a foundation for future work on adaptive slicing, compression-aware model partitioning, and real-time multimodal inference in large-scale Edge AI systems.
Degree
thesis:*- Name thesis:degree_name
- Ph.D. (Doctor of Philosophy)
- Level thesis:degree_level
- Doctoral
- Discipline thesis:degree_discipline
- Computer Science (UMKC)
- Grantor
- University of Missouri--Kansas City
- Year dc:date.issued
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Mubark, Waleed Hassan
- Advisor dc:contributor.advisor
-
- Uddin, Md Yusuf Sarwar (Mohammad Yusuf Sarwar)
Rights
- Language dc:language.iso
- en_US
Identifiers
dc:identifier.*- Handle dc:identifier.uri
- https://hdl.handle.net/10355/110343
- OAI identifier oai:identifier
- oai:mospace.umsystem.edu:10355/110343