{"id":{"repo_id":"tu-berlin","oai_identifier":"oai:depositonce.tu-berlin.de:11303/23198"},"canonical_url":"https://search.dev.ndltd.org/etd/tu-berlin/oai:depositonce.tu-berlin.de:11303/23198","repository":{"repo_id":"tu-berlin","name":"Technische Universität Berlin","base_url":"https://api-depositonce.tu-berlin.de/server/oai/request"},"display":{"title":"A layered architecture for log analysis in complex IT systems","abstract":"In the rapidly evolving landscape of Information Technology (IT), the stability and reliability of IT systems and services are important because they underpin numerous aspects of modern life. However, their increasing complexity poses significant challenges for DevOps teams, who are responsible for their implementation and maintenance. Log analysis, a core component of Artificial Intelligence for IT Operations (AIOps), plays an essential role by serving as a major source for investigating the complex behaviors and failures of IT systems. Therefore, this dissertation addresses the critical need for effective log analysis in complex IT systems by introducing a three-layered architecture designed to enhance the capabilities of DevOps teams in failure resolution. The first layer, Log Investigation, focuses on autonomous labeling and anomaly classification to provide the groundwork for the next layers. We developed a method that accurately labels log data autonomously, facilitating supervised model training and the precise evaluation of anomaly detection methods. In addition, we created a taxonomy to classify anomalies into three different categories, guaranteeing the selection of a suitable anomaly detection method. Within the second layer, Anomaly Detection, we identify behaviors of IT systems that deviate from the norm. Therefore, we propose a flexible anomaly detection method adaptable to various training scenarios: Unsupervised, weakly supervised, or supervised. Evaluations on public and industry data sets demonstrate that our method achieves F1-Scores ranging from 0.98 to 1.0 in different training scenarios, ensuring a reliable anomaly detection. The third layer addresses Root Cause Analysis. With our developed root cause analysis method we can identify a minimal set of log lines that describe a failure along with its origin and the sequence of events that led to it. By balancing training data and identifying the primary services involved, our root cause analysis method consistently identifies 90-98 % of root cause log lines within the top 10 candidates, providing precise and actionable insights for failure mitigation. Our research answers the overarching question of how log analysis methods can be designed and optimized to help DevOps teams resolve failures efficiently. By integrating these three layers, our architecture equips DevOps teams with the necessary methods to enhance IT system reliability.","abstract_html":"In the rapidly evolving landscape of Information Technology (IT), the stability and reliability of IT systems and services are important because they underpin numerous aspects of modern life. However, their increasing complexity poses significant challenges for DevOps teams, who are responsible for their implementation and maintenance. Log analysis, a core component of Artificial Intelligence for IT Operations (AIOps), plays an essential role by serving as a major source for investigating the complex behaviors and failures of IT systems. Therefore, this dissertation addresses the critical need for effective log analysis in complex IT systems by introducing a three-layered architecture designed to enhance the capabilities of DevOps teams in failure resolution. The first layer, Log Investigation, focuses on autonomous labeling and anomaly classification to provide the groundwork for the next layers. We developed a method that accurately labels log data autonomously, facilitating supervised model training and the precise evaluation of anomaly detection methods. In addition, we created a taxonomy to classify anomalies into three different categories, guaranteeing the selection of a suitable anomaly detection method. Within the second layer, Anomaly Detection, we identify behaviors of IT systems that deviate from the norm. Therefore, we propose a flexible anomaly detection method adaptable to various training scenarios: Unsupervised, weakly supervised, or supervised. Evaluations on public and industry data sets demonstrate that our method achieves F1-Scores ranging from 0.98 to 1.0 in different training scenarios, ensuring a reliable anomaly detection. The third layer addresses Root Cause Analysis. With our developed root cause analysis method we can identify a minimal set of log lines that describe a failure along with its origin and the sequence of events that led to it. By balancing training data and identifying the primary services involved, our root cause analysis method consistently identifies 90-98 % of root cause log lines within the top 10 candidates, providing precise and actionable insights for failure mitigation. Our research answers the overarching question of how log analysis methods can be designed and optimized to help DevOps teams resolve failures efficiently. By integrating these three layers, our architecture equips DevOps teams with the necessary methods to enhance IT system reliability.","abstract_has_math":false,"creators":["Wittkopp, Thorsten"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Kao, Odej"],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024","date_published":"2024","updated_at":"2026-07-27T21:28:44Z","subjects":[],"languages":["en"],"rights":[],"rights_urls":["https://creativecommons.org/licenses/by/4.0/"],"identifier_entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://doi.org/10.14279/depositonce-22012"],"render_values":[{"text":"https://doi.org/10.14279/depositonce-22012","href":"https://doi.org/10.14279/depositonce-22012","code":true}]}]},"links":{"outbound_url":"https://depositonce.tu-berlin.de/handle/11303/23198","outbound_label":"Repository record","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Kao, Odej"]},{"key":"dc:creator","label":"Author","values":["Wittkopp, Thorsten"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2024-11-21T11:16:29Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2024-11-21T11:16:29Z"]},{"key":"dc:date.issued","label":"Date","values":["2024"]},{"key":"dc:type","label":"Dc Type","values":["Doctoral Thesis"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]},{"key":"dc:rights.uri","label":"Rights URI","values":["https://creativecommons.org/licenses/by/4.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://depositonce.tu-berlin.de/handle/11303/23198","https://doi.org/10.14279/depositonce-22012"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["In the rapidly evolving landscape of Information Technology (IT), the stability and reliability of IT systems and services are important because they underpin numerous aspects of modern life. However, their increasing complexity poses significant challenges for DevOps teams, who are responsible for their implementation and maintenance. Log analysis, a core component of Artificial Intelligence for IT Operations (AIOps), plays an essential role by serving as a major source for investigating the complex behaviors and failures of IT systems. Therefore, this dissertation addresses the critical need for effective log analysis in complex IT systems by introducing a three-layered architecture designed to enhance the capabilities of DevOps teams in failure resolution. The first layer, Log Investigation, focuses on autonomous labeling and anomaly classification to provide the groundwork for the next layers. We developed a method that accurately labels log data autonomously, facilitating supervised model training and the precise evaluation of anomaly detection methods. In addition, we created a taxonomy to classify anomalies into three different categories, guaranteeing the selection of a suitable anomaly detection method. Within the second layer, Anomaly Detection, we identify behaviors of IT systems that deviate from the norm. Therefore, we propose a flexible anomaly detection method adaptable to various training scenarios: Unsupervised, weakly supervised, or supervised. Evaluations on public and industry data sets demonstrate that our method achieves F1-Scores ranging from 0.98 to 1.0 in different training scenarios, ensuring a reliable anomaly detection. The third layer addresses Root Cause Analysis. With our developed root cause analysis method we can identify a minimal set of log lines that describe a failure along with its origin and the sequence of events that led to it. By balancing training data and identifying the primary services involved, our root cause analysis method consistently identifies 90-98 % of root cause log lines within the top 10 candidates, providing precise and actionable insights for failure mitigation. Our research answers the overarching question of how log analysis methods can be designed and optimized to help DevOps teams resolve failures efficiently. By integrating these three layers, our architecture equips DevOps teams with the necessary methods to enhance IT system reliability.","In der sich schnell entwickelnden Landschaft der Informationstechnologie (IT) sind die Stabilität und Zuverlässigkeit von IT-Systemen und -Diensten von großer Bedeutung. Diese Systeme unterstützen zahlreiche Aspekte des modernen Lebens, aber ihre zunehmende Komplexität stellt DevOps-Teams, die für ihre Wartung verantwortlich sind, vor große Herausforderungen. Die Log-Analyse, eine Kernkomponente der Artificial Intelligence for IT-Operations (AIOps), spielt eine wesentliche Rolle, da sie als eine der Hauptquellen für die Untersuchung von IT-Systemen dient und als Grundlage für eine Fehleruntersuchung benutzt werden kann. Diese Dissertation befasst sich daher mit der Log-Analyse in komplexen IT-Systemen, indem sie eine dreischichtige Architektur vorstellt, die DevOps-Teams bei der Fehleranalyse und Fehlerbehebung unterstützt. Die erste Ebene, Log Investigation, konzentriert sich auf das automatisierte Labeln von Datensätzen und die Klassifikation von Anomalien. Dafür haben wir einerseits eine Methode entwickelt, die Anomalien selbstständig labelt und anderseits eine Taxonomie zur Klassifikation von Anomalien erstellt. Somit gewährleistet diese Ebene eine Auswertung von Anomalieerkennungsmethoden durch die Bereitstellung nahezu perfekt gelabelter Datensätze sowie eine zielgerichtete Auswahl von Anomalieerkennungsmethoden durch die Bereitstellung von verschiedenen Anomalienarten in den Log-Daten. Auf der zweiten Ebene, Anomaly Detection, beschreiben wir eine allgemeine Anomalieerkennungsmethode, die sich an verschiedene Trainingsbedingungen anpassen lässt: unbeaufsichtigt, schwach überwacht oder überwacht. Dabei zeigen unsere Auswertungen auf öffentlichen und industriellen Datensätzen, dass unsere Methode in verschiedenen Trainingsszenarien F1-Werte von 0,98 bis 1,0 erreichen kann und somit eine zuverlässige Anomalieerkennung gewährleistet. Die dritte Ebene befasst sich mit der Ursachenanalyse, wobei irrelevante Anomalien herausgefiltert werden. Dafür erstellen wir automatisch ausbalancierte Trainingsdaten, um unsere Root Cause Analyse Methode zu trainieren. Im Anschluss analysieren wir, welche Services an dem Fehler beteiligt sind und präsentieren die entsprechenden anomalen Log-Zeilen dem DevOps Team. Dabei befinden sich 90-98% der präsentierten Log-Zeilen innerhalb der Top-10-Kandidaten und liefern präzise Erkenntnisse zur Fehlerbehebung. Im Ergebnis beantwortet diese Forschungsarbeit die übergreifende Frage, wie eine Log-Analyse so gestaltet und optimiert werden kann, dass sie DevOps-Teams genügend Details liefert, damit diese Fehler im System beheben können. Durch die Integration dieser drei Ebenen: Log Investigation, Anomaly Detection und Root Cause Analyse, stattet unsere Architektur DevOps-Teams mit den notwendigen Werkzeugen aus, um die Zuverlässigkeit und Leistung von IT-Systemen zu verbessern."]},{"key":"dc:title","label":"Title","values":["A layered architecture for log analysis in complex IT systems"]}]}],"canonical_facts":{"dc:contributor.advisor":["Kao, Odej"],"dc:creator":["Wittkopp, Thorsten"],"dc:date.accessioned":["2024-11-21T11:16:29Z"],"dc:date.available":["2024-11-21T11:16:29Z"],"dc:date.issued":["2024"],"dc:description.abstract":["In the rapidly evolving landscape of Information Technology (IT), the stability and reliability of IT systems and services are important because they underpin numerous aspects of modern life. However, their increasing complexity poses significant challenges for DevOps teams, who are responsible for their implementation and maintenance. Log analysis, a core component of Artificial Intelligence for IT Operations (AIOps), plays an essential role by serving as a major source for investigating the complex behaviors and failures of IT systems. Therefore, this dissertation addresses the critical need for effective log analysis in complex IT systems by introducing a three-layered architecture designed to enhance the capabilities of DevOps teams in failure resolution. The first layer, Log Investigation, focuses on autonomous labeling and anomaly classification to provide the groundwork for the next layers. We developed a method that accurately labels log data autonomously, facilitating supervised model training and the precise evaluation of anomaly detection methods. In addition, we created a taxonomy to classify anomalies into three different categories, guaranteeing the selection of a suitable anomaly detection method. Within the second layer, Anomaly Detection, we identify behaviors of IT systems that deviate from the norm. Therefore, we propose a flexible anomaly detection method adaptable to various training scenarios: Unsupervised, weakly supervised, or supervised. Evaluations on public and industry data sets demonstrate that our method achieves F1-Scores ranging from 0.98 to 1.0 in different training scenarios, ensuring a reliable anomaly detection. The third layer addresses Root Cause Analysis. With our developed root cause analysis method we can identify a minimal set of log lines that describe a failure along with its origin and the sequence of events that led to it. By balancing training data and identifying the primary services involved, our root cause analysis method consistently identifies 90-98 % of root cause log lines within the top 10 candidates, providing precise and actionable insights for failure mitigation. Our research answers the overarching question of how log analysis methods can be designed and optimized to help DevOps teams resolve failures efficiently. By integrating these three layers, our architecture equips DevOps teams with the necessary methods to enhance IT system reliability.","In der sich schnell entwickelnden Landschaft der Informationstechnologie (IT) sind die Stabilität und Zuverlässigkeit von IT-Systemen und -Diensten von großer Bedeutung. Diese Systeme unterstützen zahlreiche Aspekte des modernen Lebens, aber ihre zunehmende Komplexität stellt DevOps-Teams, die für ihre Wartung verantwortlich sind, vor große Herausforderungen. Die Log-Analyse, eine Kernkomponente der Artificial Intelligence for IT-Operations (AIOps), spielt eine wesentliche Rolle, da sie als eine der Hauptquellen für die Untersuchung von IT-Systemen dient und als Grundlage für eine Fehleruntersuchung benutzt werden kann. Diese Dissertation befasst sich daher mit der Log-Analyse in komplexen IT-Systemen, indem sie eine dreischichtige Architektur vorstellt, die DevOps-Teams bei der Fehleranalyse und Fehlerbehebung unterstützt. Die erste Ebene, Log Investigation, konzentriert sich auf das automatisierte Labeln von Datensätzen und die Klassifikation von Anomalien. Dafür haben wir einerseits eine Methode entwickelt, die Anomalien selbstständig labelt und anderseits eine Taxonomie zur Klassifikation von Anomalien erstellt. Somit gewährleistet diese Ebene eine Auswertung von Anomalieerkennungsmethoden durch die Bereitstellung nahezu perfekt gelabelter Datensätze sowie eine zielgerichtete Auswahl von Anomalieerkennungsmethoden durch die Bereitstellung von verschiedenen Anomalienarten in den Log-Daten. Auf der zweiten Ebene, Anomaly Detection, beschreiben wir eine allgemeine Anomalieerkennungsmethode, die sich an verschiedene Trainingsbedingungen anpassen lässt: unbeaufsichtigt, schwach überwacht oder überwacht. Dabei zeigen unsere Auswertungen auf öffentlichen und industriellen Datensätzen, dass unsere Methode in verschiedenen Trainingsszenarien F1-Werte von 0,98 bis 1,0 erreichen kann und somit eine zuverlässige Anomalieerkennung gewährleistet. Die dritte Ebene befasst sich mit der Ursachenanalyse, wobei irrelevante Anomalien herausgefiltert werden. Dafür erstellen wir automatisch ausbalancierte Trainingsdaten, um unsere Root Cause Analyse Methode zu trainieren. Im Anschluss analysieren wir, welche Services an dem Fehler beteiligt sind und präsentieren die entsprechenden anomalen Log-Zeilen dem DevOps Team. Dabei befinden sich 90-98% der präsentierten Log-Zeilen innerhalb der Top-10-Kandidaten und liefern präzise Erkenntnisse zur Fehlerbehebung. Im Ergebnis beantwortet diese Forschungsarbeit die übergreifende Frage, wie eine Log-Analyse so gestaltet und optimiert werden kann, dass sie DevOps-Teams genügend Details liefert, damit diese Fehler im System beheben können. Durch die Integration dieser drei Ebenen: Log Investigation, Anomaly Detection und Root Cause Analyse, stattet unsere Architektur DevOps-Teams mit den notwendigen Werkzeugen aus, um die Zuverlässigkeit und Leistung von IT-Systemen zu verbessern."],"dc:identifier.uri":["https://depositonce.tu-berlin.de/handle/11303/23198","https://doi.org/10.14279/depositonce-22012"],"dc:language.iso":["en"],"dc:rights.uri":["https://creativecommons.org/licenses/by/4.0/"],"dc:title":["A layered architecture for log analysis in complex IT systems"],"dc:type":["Doctoral Thesis"]},"updated_at":"2026-07-27T21:28:44Z"}