Back to results

Technische Universität Berlin

Data-driven transfer optimizations for big data in the industrial internet of things

Abstract

dc:description.abstract

In the last two decades, the Internet of Things (IoT) has grown from a mere vision to everyday reality. Its fundamental idea is that devices become interconnected with each other and digital services. The consumer side of the IoT, the Consumer Internet of Things (CIoTs), has become omnipresent in the form of wearables, virtual assistants, and smart home solutions. The industrial side of the IoT, the Industrial Internet of Things (IIoT), has received less attention from the general public. The IIoT takes the shape of industrial-grade devices, from trucks to industrial robots, that are equipped with sensors and networking chipsets. It promises to reduce waste, increase machine lifespans, improve energy efficiency, and enable mass customization. The CIoT predominantly creates big data sparsely across wide areas, e.g., distributed over many households. CIoT applications collect and process this data in the cloud. In contrast, the IIoT predominantly creates big data at industrial facilities that are densely populated with devices. Because these industrial facilities are often connected to the cloud by low-bandwidth access networks, IIoT big data cannot be entirely transferred to the cloud. Simultaneously, industrial facilities are often equipped with limited computing resources. This creates a data-compute asymmetry where most data stays at resource-constrained industrial facilities, and only a fraction is transferred to the resource rich cloud. Unmitigated, the network bottleneck delays the installation of IIoT applications. This thesis introduces software solutions that reduce the impact of the network bottleneck. Systems processing IIoT big data face complexity from both the data sources and application requirements. On the one side, the data is generated by inherently hierarchical and distributed industrial processes and retains these qualities. On the other side, IIoT applications have diverse requirements on data access and processing (e.g., requiring database-like access to historic IIoT big data or processing recent IIoT big data as data streams). This work proposes a high-level architecture that connects both sides using novel computing primitives. Our novel computing primitives flexibly aggregate and combine data across hierarchies and locations. As part of our architecture, we introduce data-driven transfer optimizations to reduce the impact of the network bottleneck. The remainder of the thesis presents three case studies that implement data-driven transfer optimizations for different data processing frameworks. In our first case study, IIoT applications in the cloud access a data store at an industrial facility. They face a trade-off between processing individual queries at the industrial facility and transferring raw data to the cloud. We introduce online replication strategies that make fine-granular choices based on data access patterns. In our second case study, an IIoT application identifies the top-k most relevant objects (e.g., machine failures) across multiple industrial facilities. We introduce a new fixed-phase distributed top-k algorithm. This algorithm uses fewer phases than related work while simultaneously reducing the data transfer volume compared to the state-of-the-art. In our final case study, IIoT applications process data streams using dataflow programs. Dataflow programs process data by moving it through an operator graph. A sudden rise in the data input rate or a software or hardware failure risks to increase the dataflow program’s latency and decrease its throughput. We introduce a load shedding solution that mitigates this risk and simultaneously balances the data loss with the loss of previously done work. Our work enables IIoT applications for resource and bandwidth-constrained industrial facilities.

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Semmler, Niklas Bernhard
Advisor dc:contributor.advisor
  • Feldmann, Anja

Rights

Language dc:language.iso
en

Identifiers

dc:identifier.*
OAI identifier oai:identifier
oai:depositonce.tu-berlin.de:11303/16751

Chain of custody

source
Harvested from
Technische Universität Berlin
Base URL
api-depositonce.tu-berlin.de/server/oai/request
Last updated
2026-07-27
Source record
OAI-PMH GetRecord
related terms
citation

Semmler, Niklas Bernhard. Data-driven transfer optimizations for big data in the industrial internet of things. 2022. https://depositonce.tu-berlin.de/handle/11303/16751