Back to results

Boston University

Automating telemetry- and trace-based analytics on large-scale distributed systems

Abstract

dc:description.abstract

Large-scale distributed systems---such as supercomputers, cloud computing platforms, and distributed applications---routinely suffer from slowdowns and crashes due to software and hardware problems, resulting in reduced efficiency and wasted resources. These large-scale systems typically deploy monitoring or tracing systems that gather a variety of statistics about the state of the hardware and the software. State-of-the-art methods either analyze this data manually, or design unique automated methods for each specific problem. This thesis builds on the vision that generalized automated analytics methods on the data sets collected from these complex computing systems provide critical information about the causes of the problems, and this analysis can then enable proactive management to improve performance, resilience, efficiency, or security significantly beyond current limits. This thesis seeks to design scalable, automated analytics methods and frameworks for large-scale distributed systems that minimize dependency on expert knowledge, automate parts of the solution process, and help make systems more resilient. In addition to analyzing data that is already collected from systems, our frameworks also identify what to collect from where in the system, such that the collected data would be concise and useful for manual analytics. We focus on two data sources for conducting analytics: numeric telemetry data, which is typically collected from operating system or hardware counters, and end-to-end traces collected from distributed applications. This thesis makes the following contributions in large-scale distributed systems: (1) Designing a framework for accurately diagnosing previously encountered performance variations, (2) designing a technique for detecting (unwanted) applications running on the systems, (3) developing a suite for reproducing performance variations that can be used to systematically develop analytics methods, (4) designing a method to explain predictions of black-box machine learning frameworks, and (5) constructing an end-to-end tracing framework that can dynamically adjust instrumentation for effective diagnosis of performance problems.

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Ateş, Emre
Advisor dc:contributor.advisor
  • Coskun, Ayse K.

Subjects

dc:subject × 7

Rights

dc:rights
Statement dc:rights
  • Attribution 4.0 International
Language dc:language.iso
en_US

Identifiers

dc:identifier.*
Handle dc:identifier.uri
https://hdl.handle.net/2144/41472
OAI identifier oai:identifier
oai:open.bu.edu:2144/41472

Chain of custody

source
Harvested from
Boston University
Base URL
open.bu.edu/oai/request
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Ateş, Emre. Automating telemetry- and trace-based analytics on large-scale distributed systems. 2020. https://hdl.handle.net/2144/41472