Back to search

University of Exeter

Log-Driven Robust Anomaly Detection and Root Cause Localisation for Distributed Computing Systems

Abstract

dc:description

Modern networked services produce large volumes of operational logs that record what components did at the moment events occurred. These records are central during incident response, yet they are difficult to use at scale. Message wording changes across releases, timestamps are not always comparable across hosts, and identifiers for forming sessions are sometimes missing or inconsistently propagated. Event-level labels are also scarce. This thesis concentrates on two capabilities that work together in such conditions: detecting abnormal executions at the session level, and ranking a short list of diagnostic lines within each abnormal session so that engineers can act quickly. The first contribution is a session level detector that combines a deterministic task splitting policy with a two layer LSTM over event sequences. The detector is trained on representative successful runs and assigns each session an anomaly score that supports threshold based triage. Experiments on an OpenStack corpus collected in our laboratory show clear separation between normal and abnormal activities and reveal gradual degradation trends that appear before failure, which is useful for early warning. The second contribution, LogAR, addresses robustness to logging noise and wording variation. LogAR couples semantic vectorisation of messages with an encoder–decoder that includes a self expressive layer which encourages each event to be reconstructed from a small related set. On the public HDFS dataset, with controlled noise introduced into the logs, LogAR remains effective and outperforms strong baselines including Decision Tree, PCA, SVM, LogCluster and an LSTM sequence model. The results indicate improved tolerance to duplication, omission and incidental rephrasing. The third contribution formulates localisation as a within-session learning-to-rank problem over template-indexed events. Each event receives a score that blends (i) sequence-context deviation (how strongly the event violates the learned routine of its surrounding template stream), with (ii) content cues derived from template weighting, and (iii) simple suppressors for repetitive progress lines. The ranker operates on the same template-and-fields interface as detection, so outputs map cleanly back to concrete log lines. On the public BGL dataset, the approach attains high Recall@ k (for small k) and strong mean reciprocal rank in our experiments, with a diagnostic line typically appearing within the first few suggestions for abnormal sessions, and performance remains stable across threshold settings. The result is a concise, actionable shortlist that integrates with routine triage and requires only modest computational overhead. Taken together, the thesis shows that stable session formation, sequence modelling, and semantics aware representation can be combined to deliver accurate detection and actionable localisation on real operational logs. The integrated and systematic designs favour reproducibility, modest computational cost, and explanations that map back to the lines practitioners use to explain and fix faults.<p></p>

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Yujia Zhu (21037775)

Subjects

dc:subject × 7

Rights

dc:rights
Statement dc:rights
  • All rights reserved
  • Open Access after 2027-04-15

Identifiers

dc:identifier.*
Identifier
10779/exe.32025231.v1
OAI identifier oai:identifier
oai:figshare.com:article/32025231

Chain of custody

source
Harvested from
University of Exeter
Base URL
api.figshare.com/v2/oai
Last updated
2026-07-27
Source record
OAI-PMH GetRecord
citation

Yujia Zhu (21037775). Log-Driven Robust Anomaly Detection and Root Cause Localisation for Distributed Computing Systems. 2026.