{"id":{"repo_id":"vt","oai_identifier":"oai:vtechworks.lib.vt.edu:10919/141029"},"canonical_url":"https://search.dev.ndltd.org/etd/vt/oai:vtechworks.lib.vt.edu:10919/141029","repository":{"repo_id":"vt","name":"Virginia Tech","base_url":"https://vtechworks.lib.vt.edu/oai/request"},"display":{"title":"Time-to-Event Prediction Using Deep Learning Models: Application to GPU Failure Data","abstract":"Neural network models gain significant popularity in recent years due to their ability to identify complex patterns. In the field of reliability research, efforts are made to develop neural network models for predictive reliability. However, research focused on utilizing neural networks to forecast graphics processing unit (GPU) failures remains limited. Analyzing the reliability of GPUs is crucial for effectively maintaining GPU systems and preventing issues related to GPU failures, such as safety concerns and interruptions in simulations. Most studies concentrate on predicting the remaining lifespan of GPUs, whereas our objective is to predict both lifespan and failure status. Furthermore, there is a growing trend of integrating statistical modeling with neural networks. By incorporating statistical concepts, we can create models that better represent reality and improve interpretability. Additionally, the model efficiently manages complex data structures like hierarchical clusters, which enhances its capacity to generalize. To improve our neural network model, we combine statistical concepts with neural network techniques. Building on these motivations, we outline our research approach as follows. We describe the development of deep learning models for GPU failure prediction, and explain how statistical concepts are integrated to improve model performance. Chapter~ref{cha:genintro} provides a general overview of the research presented in this dissertation. It begins by outlining the motivation behind the study and defining the research problem. Next, it offers a brief summary of the contributions made by this research. The chapter also explains the concept of neural network models, including their training processes. Furthermore, it covers embedding layers and the use of the GPU dataset. Finally, it introduces multitask output for neural networks, which will be utilized in Chapter~ref{cha:NNspatialRE}. Chapter~ref{cha:DL} introduces a deep learning approach for predicting GPU failure time and status. We propose two distinct neural network architectures to address these prediction tasks: TypeEmbedNet for failure type prediction and TimeEmbedNet for failure time prediction. Since our predictors are categorical variables, we employ embedding layers to transform them into lower-dimensional vectors. We utilize the Cross-Entropy Loss function to optimize TypeEmbedNet and the Mean Squared Error (MSE) for TimeEmbedNet. We develop evaluation metrics to comprehensively assess the models' performance in predicting both failure time and status, taking into account the characteristics of an imbalanced dataset. We integrate data frequency, F1 score, recall, and precision into the MSE by applying penalties based on classification outcomes. In Chapter~ref{cha:NNspatialRE}, we introduce a deep learning model for competing risks. We develop a custom loss function for competing risks that incorporates survival theory and is specifically adapted for time prediction. We compare this model to other machine learning approaches and benchmark it against a previously introduced parametric model. Additionally, we test the integration of spatial random-effect embeddings to model GPU failure outcomes. To achieve this, we compute correlation matrices based on two different spatial structures—physical and logical distances—and apply them using Cholesky decomposition. We describe how we assign a learnable embedding to each GPU location, incorporating these correlation matrices. The neural network outperforms both the machine learning models and the parametric model across various MSE metrics. However, adding spatial random effects to the neural network does not result in a significant improvement in predictive performance. Chapter~ref{cha:conclusion} summarizes the key findings of this study and explores several potential avenues for future research.","abstract_html":"Neural network models gain significant popularity in recent years due to their ability to identify complex patterns. In the field of reliability research, efforts are made to develop neural network models for predictive reliability. However, research focused on utilizing neural networks to forecast graphics processing unit (GPU) failures remains limited. Analyzing the reliability of GPUs is crucial for effectively maintaining GPU systems and preventing issues related to GPU failures, such as safety concerns and interruptions in simulations. Most studies concentrate on predicting the remaining lifespan of GPUs, whereas our objective is to predict both lifespan and failure status. Furthermore, there is a growing trend of integrating statistical modeling with neural networks. By incorporating statistical concepts, we can create models that better represent reality and improve interpretability. Additionally, the model efficiently manages complex data structures like hierarchical clusters, which enhances its capacity to generalize. To improve our neural network model, we combine statistical concepts with neural network techniques. Building on these motivations, we outline our research approach as follows. We describe the development of deep learning models for GPU failure prediction, and explain how statistical concepts are integrated to improve model performance. Chapter~ref{cha:genintro} provides a general overview of the research presented in this dissertation. It begins by outlining the motivation behind the study and defining the research problem. Next, it offers a brief summary of the contributions made by this research. The chapter also explains the concept of neural network models, including their training processes. Furthermore, it covers embedding layers and the use of the GPU dataset. Finally, it introduces multitask output for neural networks, which will be utilized in Chapter~ref{cha:NNspatialRE}. Chapter~ref{cha:DL} introduces a deep learning approach for predicting GPU failure time and status. We propose two distinct neural network architectures to address these prediction tasks: TypeEmbedNet for failure type prediction and TimeEmbedNet for failure time prediction. Since our predictors are categorical variables, we employ embedding layers to transform them into lower-dimensional vectors. We utilize the Cross-Entropy Loss function to optimize TypeEmbedNet and the Mean Squared Error (MSE) for TimeEmbedNet. We develop evaluation metrics to comprehensively assess the models&#x27; performance in predicting both failure time and status, taking into account the characteristics of an imbalanced dataset. We integrate data frequency, F1 score, recall, and precision into the MSE by applying penalties based on classification outcomes. In Chapter~ref{cha:NNspatialRE}, we introduce a deep learning model for competing risks. We develop a custom loss function for competing risks that incorporates survival theory and is specifically adapted for time prediction. We compare this model to other machine learning approaches and benchmark it against a previously introduced parametric model. Additionally, we test the integration of spatial random-effect embeddings to model GPU failure outcomes. To achieve this, we compute correlation matrices based on two different spatial structures—physical and logical distances—and apply them using Cholesky decomposition. We describe how we assign a learnable embedding to each GPU location, incorporating these correlation matrices. The neural network outperforms both the machine learning models and the parametric model across various MSE metrics. However, adding spatial random effects to the neural network does not result in a significant improvement in predictive performance. Chapter~ref{cha:conclusion} summarizes the key findings of this study and explores several potential avenues for future research.","abstract_has_math":false,"creators":["Lee, Lina"],"institution":"Virginia Tech","degree_name":"Doctor of Philosophy","degree_level":"doctoral","degree_discipline":"Statistics","degree_department":"Statistics","school":null,"contributors":[],"advisors":[],"committee_chairs":["Hong, Yili"],"committee_members":["Kim, Inyoung","Deng, Xinwei","Du, Pang"],"year":2026,"date_issued":"2026-01-28","date_published":"2026-01-28","updated_at":"2026-07-22T22:19:45Z","subjects":["Competing risk","Deep learning","GPU reliability","Random effect models","Spatial embedding","Time to failure prediction"],"languages":["en"],"rights":["In Copyright"],"rights_urls":["http://rightsstatements.org/vocab/InC/1.0/"],"identifier_entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:45214"],"render_values":[{"text":"vt_gsexam:45214","href":null,"code":true}]}]},"links":{"outbound_url":"https://hdl.handle.net/10919/141029","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.committeechair","label":"Committee Chair","values":["Hong, Yili"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Kim, Inyoung","Deng, Xinwei","Du, Pang"]},{"key":"dc:contributor.department","label":"Department","values":["Statistics"]},{"key":"dc:creator","label":"Author","values":["Lee, Lina"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2026-01-29T09:00:08Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2026-01-29T09:00:08Z"]},{"key":"dc:date.issued","label":"Date","values":["2026-01-28"]},{"key":"dc:publisher","label":"Institution","values":["Virginia Tech"]},{"key":"dc:type","label":"Dc Type","values":["Dissertation"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Statistics"]},{"key":"thesis:degree_level","label":"Degree Level","values":["doctoral"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctor of Philosophy"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Virginia Polytechnic Institute and State University"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Competing risk","Deep learning","GPU reliability","Random effect models","Spatial embedding","Time to failure prediction"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["In Copyright"]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://rightsstatements.org/vocab/InC/1.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:45214"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10919/141029"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Neural network models gain significant popularity in recent years due to their ability to identify complex patterns. In the field of reliability research, efforts are made to develop neural network models for predictive reliability. However, research focused on utilizing neural networks to forecast graphics processing unit (GPU) failures remains limited. Analyzing the reliability of GPUs is crucial for effectively maintaining GPU systems and preventing issues related to GPU failures, such as safety concerns and interruptions in simulations. Most studies concentrate on predicting the remaining lifespan of GPUs, whereas our objective is to predict both lifespan and failure status. Furthermore, there is a growing trend of integrating statistical modeling with neural networks. By incorporating statistical concepts, we can create models that better represent reality and improve interpretability. Additionally, the model efficiently manages complex data structures like hierarchical clusters, which enhances its capacity to generalize. To improve our neural network model, we combine statistical concepts with neural network techniques. Building on these motivations, we outline our research approach as follows. We describe the development of deep learning models for GPU failure prediction, and explain how statistical concepts are integrated to improve model performance. Chapter~ref{cha:genintro} provides a general overview of the research presented in this dissertation. It begins by outlining the motivation behind the study and defining the research problem. Next, it offers a brief summary of the contributions made by this research. The chapter also explains the concept of neural network models, including their training processes. Furthermore, it covers embedding layers and the use of the GPU dataset. Finally, it introduces multitask output for neural networks, which will be utilized in Chapter~ref{cha:NNspatialRE}. Chapter~ref{cha:DL} introduces a deep learning approach for predicting GPU failure time and status. We propose two distinct neural network architectures to address these prediction tasks: TypeEmbedNet for failure type prediction and TimeEmbedNet for failure time prediction. Since our predictors are categorical variables, we employ embedding layers to transform them into lower-dimensional vectors. We utilize the Cross-Entropy Loss function to optimize TypeEmbedNet and the Mean Squared Error (MSE) for TimeEmbedNet. We develop evaluation metrics to comprehensively assess the models' performance in predicting both failure time and status, taking into account the characteristics of an imbalanced dataset. We integrate data frequency, F1 score, recall, and precision into the MSE by applying penalties based on classification outcomes. In Chapter~ref{cha:NNspatialRE}, we introduce a deep learning model for competing risks. We develop a custom loss function for competing risks that incorporates survival theory and is specifically adapted for time prediction. We compare this model to other machine learning approaches and benchmark it against a previously introduced parametric model. Additionally, we test the integration of spatial random-effect embeddings to model GPU failure outcomes. To achieve this, we compute correlation matrices based on two different spatial structures—physical and logical distances—and apply them using Cholesky decomposition. We describe how we assign a learnable embedding to each GPU location, incorporating these correlation matrices. The neural network outperforms both the machine learning models and the parametric model across various MSE metrics. However, adding spatial random effects to the neural network does not result in a significant improvement in predictive performance. Chapter~ref{cha:conclusion} summarizes the key findings of this study and explores several potential avenues for future research."]},{"key":"dc:description.abstractgeneral","label":"General Abstract","values":["Graphics Processing Units, or GPUs, are powerful computer processors used in supercomputers, scientific research, engineering tools, and everyday technologies. When a GPU fails, it can interrupt important work, slow down large systems, and increase operating costs. Being able to predict when a failure might happen and what kind of failure it will be helps organizations prevent unexpected outages and keep their systems running smoothly. This research develops deep learning models that predict both the timing and the type of GPU failure. While many previous studies focus only on estimating how long a GPU will last, this work takes a more complete approach by predicting two outcomes at the same time. Because GPU data contains many categories and complicated patterns, deep learning provides a flexible way to capture these relationships. To make the models more realistic and useful, this research blends ideas from statistics with modern neural networks. It incorporates concepts from survival analysis to reflect the fact that GPUs can fail for different reasons and that some GPUs do not fail during the time they are observed. The models also include information about where each GPU is located inside the system, which helps reveal whether certain areas experience failures more often. Several types of models are tested, including common machine learning methods and a traditional statistical model. The results show that the deep learning approach can naturally represent important features of real failure data, such as multiple potential failure causes and incomplete observations. Although adding spatial information does not substantially change predictions, it offers helpful insight into how failures may cluster within the computing system. Overall, this research demonstrates how combining statistical thinking with deep learning can support better reliability management in large-scale computing environments. The framework developed here can also be adapted to other systems where understanding and anticipating failures is essential."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["Doctor of Philosophy"]},{"key":"dc:format.medium","label":"Dc Format Medium","values":["ETD"]},{"key":"dc:title","label":"Title","values":["Time-to-Event Prediction Using Deep Learning Models: Application to GPU Failure Data"]}]}],"canonical_facts":{"dc:contributor.committeechair":["Hong, Yili"],"dc:contributor.committeemember":["Kim, Inyoung","Deng, Xinwei","Du, Pang"],"dc:contributor.department":["Statistics"],"dc:creator":["Lee, Lina"],"dc:date.accessioned":["2026-01-29T09:00:08Z"],"dc:date.available":["2026-01-29T09:00:08Z"],"dc:date.issued":["2026-01-28"],"dc:description.abstract":["Neural network models gain significant popularity in recent years due to their ability to identify complex patterns. In the field of reliability research, efforts are made to develop neural network models for predictive reliability. However, research focused on utilizing neural networks to forecast graphics processing unit (GPU) failures remains limited. Analyzing the reliability of GPUs is crucial for effectively maintaining GPU systems and preventing issues related to GPU failures, such as safety concerns and interruptions in simulations. Most studies concentrate on predicting the remaining lifespan of GPUs, whereas our objective is to predict both lifespan and failure status. Furthermore, there is a growing trend of integrating statistical modeling with neural networks. By incorporating statistical concepts, we can create models that better represent reality and improve interpretability. Additionally, the model efficiently manages complex data structures like hierarchical clusters, which enhances its capacity to generalize. To improve our neural network model, we combine statistical concepts with neural network techniques. Building on these motivations, we outline our research approach as follows. We describe the development of deep learning models for GPU failure prediction, and explain how statistical concepts are integrated to improve model performance. Chapter~ref{cha:genintro} provides a general overview of the research presented in this dissertation. It begins by outlining the motivation behind the study and defining the research problem. Next, it offers a brief summary of the contributions made by this research. The chapter also explains the concept of neural network models, including their training processes. Furthermore, it covers embedding layers and the use of the GPU dataset. Finally, it introduces multitask output for neural networks, which will be utilized in Chapter~ref{cha:NNspatialRE}. Chapter~ref{cha:DL} introduces a deep learning approach for predicting GPU failure time and status. We propose two distinct neural network architectures to address these prediction tasks: TypeEmbedNet for failure type prediction and TimeEmbedNet for failure time prediction. Since our predictors are categorical variables, we employ embedding layers to transform them into lower-dimensional vectors. We utilize the Cross-Entropy Loss function to optimize TypeEmbedNet and the Mean Squared Error (MSE) for TimeEmbedNet. We develop evaluation metrics to comprehensively assess the models' performance in predicting both failure time and status, taking into account the characteristics of an imbalanced dataset. We integrate data frequency, F1 score, recall, and precision into the MSE by applying penalties based on classification outcomes. In Chapter~ref{cha:NNspatialRE}, we introduce a deep learning model for competing risks. We develop a custom loss function for competing risks that incorporates survival theory and is specifically adapted for time prediction. We compare this model to other machine learning approaches and benchmark it against a previously introduced parametric model. Additionally, we test the integration of spatial random-effect embeddings to model GPU failure outcomes. To achieve this, we compute correlation matrices based on two different spatial structures—physical and logical distances—and apply them using Cholesky decomposition. We describe how we assign a learnable embedding to each GPU location, incorporating these correlation matrices. The neural network outperforms both the machine learning models and the parametric model across various MSE metrics. However, adding spatial random effects to the neural network does not result in a significant improvement in predictive performance. Chapter~ref{cha:conclusion} summarizes the key findings of this study and explores several potential avenues for future research."],"dc:description.abstractgeneral":["Graphics Processing Units, or GPUs, are powerful computer processors used in supercomputers, scientific research, engineering tools, and everyday technologies. When a GPU fails, it can interrupt important work, slow down large systems, and increase operating costs. Being able to predict when a failure might happen and what kind of failure it will be helps organizations prevent unexpected outages and keep their systems running smoothly. This research develops deep learning models that predict both the timing and the type of GPU failure. While many previous studies focus only on estimating how long a GPU will last, this work takes a more complete approach by predicting two outcomes at the same time. Because GPU data contains many categories and complicated patterns, deep learning provides a flexible way to capture these relationships. To make the models more realistic and useful, this research blends ideas from statistics with modern neural networks. It incorporates concepts from survival analysis to reflect the fact that GPUs can fail for different reasons and that some GPUs do not fail during the time they are observed. The models also include information about where each GPU is located inside the system, which helps reveal whether certain areas experience failures more often. Several types of models are tested, including common machine learning methods and a traditional statistical model. The results show that the deep learning approach can naturally represent important features of real failure data, such as multiple potential failure causes and incomplete observations. Although adding spatial information does not substantially change predictions, it offers helpful insight into how failures may cluster within the computing system. Overall, this research demonstrates how combining statistical thinking with deep learning can support better reliability management in large-scale computing environments. The framework developed here can also be adapted to other systems where understanding and anticipating failures is essential."],"dc:description.degree":["Doctor of Philosophy"],"dc:format.medium":["ETD"],"dc:identifier.other":["vt_gsexam:45214"],"dc:identifier.uri":["https://hdl.handle.net/10919/141029"],"dc:language.iso":["en"],"dc:publisher":["Virginia Tech"],"dc:rights":["In Copyright"],"dc:rights.uri":["http://rightsstatements.org/vocab/InC/1.0/"],"dc:subject":["Competing risk","Deep learning","GPU reliability","Random effect models","Spatial embedding","Time to failure prediction"],"dc:title":["Time-to-Event Prediction Using Deep Learning Models: Application to GPU Failure Data"],"dc:type":["Dissertation"],"thesis:degree_discipline":["Statistics"],"thesis:degree_level":["doctoral"],"thesis:degree_name":["Doctor of Philosophy"],"thesis:institution_name":["Virginia Polytechnic Institute and State University"]},"updated_at":"2026-07-22T22:19:45Z"}