Faculty of Graduate Studies and Research, University of Regina
Clustering and dimensionality reduction for time-series service monitoring data
Abstract
dc:description.abstractService monitoring applications allow customers to measure the performance, availability, and resolve application issues before they affect users. Since service monitoring applications continuously produce data to monitor their availability, therefore, high dimensionality, unlabeled data and changing data distribution are all prevalent. In this thesis, we efficiently address these three issues using the constructed service monitoring dataset. Higher dimensionality means higher computational cost to perform training and often leads to over-fitting while l earning a model. Furthermore, in the presence of high dimensionality, data are highly correlated resulting in insignificant and irrelevant features. These features have less impact on the prediction. To this end, the first part of the thesis conceptually and empirically explores the most representative dimensionality reduction (DR) methods from different categories. Next, we construct a new and challenging End-to-End (E2E) service monitoring dataset by extracting heterogeneous sub-datasets from multiple subservers, tackling data incompleteness in each sub-dataset using several imputation techniques, and fusing all the optimally imputed sub-datasets. This target dataset is highly dimensional, temporal, unlabeled, and non-linear. As the dataset is new, based on robust clustering approaches, we thoroughly assess the quality of the initial dataset and the reconstructed datasets (same dimensionality as the initial dataset) produced with Deep and Convolutional AutoEncoders. The experiments disclose that the reconstructed dataset with Deep AutoEncoder is the most performing. Later, we propose an ensemble-based DR approach to effectively handle the high-dimensionality of the E2E dataset. The approach combines Deep AutoEncoder with Kernel Principal Component Analysis to produce better data, and then reduce the feature space respectively. Due to the massive size of the dataset, we divide it into six weekly sub-datasets. We show that no vital information is lost for the reduced sub-datasets using the reconstruction error and total explained variance ratio. Based on time-series data clustering methods, and metrics, we thoroughly evaluate the efficacy of the ensemble approach. As the initial dataset is unlabeled (so are the reduced sub-datasets), we improve the previously developed ensemble-based DR approach by further combining it with incremental DR to improve clusters’ performance and increase the cluster labels’ confidence. We consider the weekly datasets as chunks for the experimental purpose. The experiments reveal that clustering performances increase significantly after utilizing the improved ensemble-based DR. Hence, the clusters’ labels are considered as the target class labels. Finally, to process the incoming data for any service monitoring application, it is critical to classify data accurately in real-time. Hence, we consider the labeled data chunks as incoming data, and propose an adaptive classification framework using Learn++ that also handles evolving data distributions. This approach sequentially predicts and updates the monitoring model with new data, and gradually forgets past knowledge. We employ consecutive data chunks to evaluate the performance of the predictors incrementally. The experimental results demonstrate that the proposed method provides high detection rates and low misclassification rate for most of the adaptive chunks.
Degree
thesis:*- Name thesis:degree_name
- Doctor of Philosophy (PhD)
- Level thesis:degree_level
- Doctoral -- first
- Discipline thesis:degree_discipline
- Computer Science
- Grantor dc:publisher
- Faculty of Graduate Studies and Research, University of Regina
- Year dc:date.issued
- 2022
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Anowar, Farzana
- Advisor dc:contributor.advisor
-
- Sadaoui, Samira
- Committee members dc:contributor.committeemember
-
- Fan, Lisa
- Sharhiar, Nashid
Rights
- Language dc:language.iso
- en
Identifiers
dc:identifier.*- OAI identifier oai:identifier
- oai:uregina.scholaris.ca:10294/16025