{"id":{"repo_id":"queens","oai_identifier":"oai:queensu.scholaris.ca:1974/35984"},"canonical_url":"https://search.dev.ndltd.org/etd/queens/oai:queensu.scholaris.ca:1974/35984","repository":{"repo_id":"queens","name":"Queens University","base_url":"https://qspace.library.queensu.ca/server/oai/request"},"display":{"title":"Evaluation, Interpretation, and Maintenance of Machine Learning Models for IT Operations","abstract":"AIOps (Artificial Intelligence for IT Operations) solutions leverage the massive data generated during the operation of large-scale systems and machine learning models to assist in managing system operations. While prior studies focus on innovative modeling techniques to improve the performance of AIOps models, how to smoothly transition AIOps solutions from development to production remains an underexplored topic. Since operational data instances often exhibit temporal dependencies, the use of improper model evaluation methods can lead to performance overestimation. Insufficient model maintenance on AIOps solutions incorporated in the production environment can also lead to future performance degradation. In addition, the consistency of model interpretation is impacted by the volatile nature of operational data. These threats pose significant challenges for lab-developed AIOps solutions when deployed in production environments. Therefore, this thesis proposes to explore related techniques for mitigating the challenges and helping practitioners make better decisions for deploying AIOps solutions in dynamic operational environments. We evaluate the impact of different data splitting decisions to understand the data leakage and concept drift challenges in the model evaluation stage. Our findings motivate practitioners to take precautions against using the random data splitting method that could induce data leakage. We assess the factors that impact the consistency of AIOps model interpretations. We propose guidelines for practitioners that help derive more reliable and consistent interpretations from AIOps models. We evaluate model update strategies for maintaining AIOps solutions in terms of performance, updating cost, and stability. Our findings suggest that practitioners consider more sophisticated model update strategies to mitigate the impact of concept drift and to minimize operational cost and performance variation. We examine model selection mechanisms on historical models in maintaining AIOps solutions. Our findings highlight a potential research opportunities for model maintenance beyond the simple retraining and replacing strategies. This thesis helps practitioners manage and mitigate the above-mentioned challenges in the evaluation, interpretation, and maintenance of AIOps solutions to ensure the smooth transition from development to production.","abstract_html":"AIOps (Artificial Intelligence for IT Operations) solutions leverage the massive data generated during the operation of large-scale systems and machine learning models to assist in managing system operations. While prior studies focus on innovative modeling techniques to improve the performance of AIOps models, how to smoothly transition AIOps solutions from development to production remains an underexplored topic. Since operational data instances often exhibit temporal dependencies, the use of improper model evaluation methods can lead to performance overestimation. Insufficient model maintenance on AIOps solutions incorporated in the production environment can also lead to future performance degradation. In addition, the consistency of model interpretation is impacted by the volatile nature of operational data. These threats pose significant challenges for lab-developed AIOps solutions when deployed in production environments. Therefore, this thesis proposes to explore related techniques for mitigating the challenges and helping practitioners make better decisions for deploying AIOps solutions in dynamic operational environments. We evaluate the impact of different data splitting decisions to understand the data leakage and concept drift challenges in the model evaluation stage. Our findings motivate practitioners to take precautions against using the random data splitting method that could induce data leakage. We assess the factors that impact the consistency of AIOps model interpretations. We propose guidelines for practitioners that help derive more reliable and consistent interpretations from AIOps models. We evaluate model update strategies for maintaining AIOps solutions in terms of performance, updating cost, and stability. Our findings suggest that practitioners consider more sophisticated model update strategies to mitigate the impact of concept drift and to minimize operational cost and performance variation. We examine model selection mechanisms on historical models in maintaining AIOps solutions. Our findings highlight a potential research opportunities for model maintenance beyond the simple retraining and replacing strategies. This thesis helps practitioners manage and mitigate the above-mentioned challenges in the evaluation, interpretation, and maintenance of AIOps solutions to ensure the smooth transition from development to production.","abstract_has_math":false,"creators":["Lyu, Yingzhe"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":"Computing","school":null,"contributors":[],"advisors":["Hassan, Ahmed E."],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-10-31","date_published":"2025-10-31","updated_at":"2026-07-27T20:35:43Z","subjects":["Software Engineering","AIOps","Model Maintenance","Failure Prediction","Model Interpretation"],"languages":["eng"],"rights":["Attribution-NonCommercial-ShareAlike 4.0 International"],"rights_urls":["http://creativecommons.org/licenses/by-nc-sa/4.0/"],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/1974/35984","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.department","label":"Department","values":["Computing"]},{"key":"dc:contributor.supervisor","label":"Supervisor","values":["Hassan, Ahmed E."]},{"key":"dc:creator","label":"Author","values":["Lyu, Yingzhe"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2025-10-31T18:53:06Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2025-10-31T18:53:06Z"]},{"key":"dc:date.issued","label":"Date","values":["2025-10-31"]},{"key":"dc:type","label":"Dc Type","values":["thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Software Engineering","AIOps","Model Maintenance","Failure Prediction","Model Interpretation"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Attribution-NonCommercial-ShareAlike 4.0 International"]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://creativecommons.org/licenses/by-nc-sa/4.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/1974/35984"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["AIOps (Artificial Intelligence for IT Operations) solutions leverage the massive data generated during the operation of large-scale systems and machine learning models to assist in managing system operations. While prior studies focus on innovative modeling techniques to improve the performance of AIOps models, how to smoothly transition AIOps solutions from development to production remains an underexplored topic. Since operational data instances often exhibit temporal dependencies, the use of improper model evaluation methods can lead to performance overestimation. Insufficient model maintenance on AIOps solutions incorporated in the production environment can also lead to future performance degradation. In addition, the consistency of model interpretation is impacted by the volatile nature of operational data. These threats pose significant challenges for lab-developed AIOps solutions when deployed in production environments. Therefore, this thesis proposes to explore related techniques for mitigating the challenges and helping practitioners make better decisions for deploying AIOps solutions in dynamic operational environments. We evaluate the impact of different data splitting decisions to understand the data leakage and concept drift challenges in the model evaluation stage. Our findings motivate practitioners to take precautions against using the random data splitting method that could induce data leakage. We assess the factors that impact the consistency of AIOps model interpretations. We propose guidelines for practitioners that help derive more reliable and consistent interpretations from AIOps models. We evaluate model update strategies for maintaining AIOps solutions in terms of performance, updating cost, and stability. Our findings suggest that practitioners consider more sophisticated model update strategies to mitigate the impact of concept drift and to minimize operational cost and performance variation. We examine model selection mechanisms on historical models in maintaining AIOps solutions. Our findings highlight a potential research opportunities for model maintenance beyond the simple retraining and replacing strategies. This thesis helps practitioners manage and mitigate the above-mentioned challenges in the evaluation, interpretation, and maintenance of AIOps solutions to ensure the smooth transition from development to production."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["PhD"]},{"key":"dc:title","label":"Title","values":["Evaluation, Interpretation, and Maintenance of Machine Learning Models for IT Operations"]}]}],"canonical_facts":{"dc:contributor.department":["Computing"],"dc:contributor.supervisor":["Hassan, Ahmed E."],"dc:creator":["Lyu, Yingzhe"],"dc:date.accessioned":["2025-10-31T18:53:06Z"],"dc:date.available":["2025-10-31T18:53:06Z"],"dc:date.issued":["2025-10-31"],"dc:description.abstract":["AIOps (Artificial Intelligence for IT Operations) solutions leverage the massive data generated during the operation of large-scale systems and machine learning models to assist in managing system operations. While prior studies focus on innovative modeling techniques to improve the performance of AIOps models, how to smoothly transition AIOps solutions from development to production remains an underexplored topic. Since operational data instances often exhibit temporal dependencies, the use of improper model evaluation methods can lead to performance overestimation. Insufficient model maintenance on AIOps solutions incorporated in the production environment can also lead to future performance degradation. In addition, the consistency of model interpretation is impacted by the volatile nature of operational data. These threats pose significant challenges for lab-developed AIOps solutions when deployed in production environments. Therefore, this thesis proposes to explore related techniques for mitigating the challenges and helping practitioners make better decisions for deploying AIOps solutions in dynamic operational environments. We evaluate the impact of different data splitting decisions to understand the data leakage and concept drift challenges in the model evaluation stage. Our findings motivate practitioners to take precautions against using the random data splitting method that could induce data leakage. We assess the factors that impact the consistency of AIOps model interpretations. We propose guidelines for practitioners that help derive more reliable and consistent interpretations from AIOps models. We evaluate model update strategies for maintaining AIOps solutions in terms of performance, updating cost, and stability. Our findings suggest that practitioners consider more sophisticated model update strategies to mitigate the impact of concept drift and to minimize operational cost and performance variation. We examine model selection mechanisms on historical models in maintaining AIOps solutions. Our findings highlight a potential research opportunities for model maintenance beyond the simple retraining and replacing strategies. This thesis helps practitioners manage and mitigate the above-mentioned challenges in the evaluation, interpretation, and maintenance of AIOps solutions to ensure the smooth transition from development to production."],"dc:description.degree":["PhD"],"dc:identifier.uri":["https://hdl.handle.net/1974/35984"],"dc:language.iso":["eng"],"dc:rights":["Attribution-NonCommercial-ShareAlike 4.0 International"],"dc:rights.uri":["http://creativecommons.org/licenses/by-nc-sa/4.0/"],"dc:subject":["Software Engineering","AIOps","Model Maintenance","Failure Prediction","Model Interpretation"],"dc:title":["Evaluation, Interpretation, and Maintenance of Machine Learning Models for IT Operations"],"dc:type":["thesis"]},"updated_at":"2026-07-27T20:35:43Z"}