{"id":{"repo_id":"vt","oai_identifier":"oai:vtechworks.lib.vt.edu:10919/140814"},"canonical_url":"https://search.dev.ndltd.org/etd/vt/oai:vtechworks.lib.vt.edu:10919/140814","repository":{"repo_id":"vt","name":"Virginia Tech","base_url":"https://vtechworks.lib.vt.edu/oai/request"},"display":{"title":"Statistical Evaluation of Deep Learning for Event Detection in Time Series: Quantifying Uncertainty, Efficiency, and Adaptation with Applications to Seismic Data","abstract":"Rapid developments in deep learning have led to their widespread use in domains that rely on time series, largely because of their strong performance and flexibility. Yet evaluation practices have not kept pace. Deep learning models are often assessed using a few performance metrics computed on benchmark datasets, which ignores important questions about how predictive performance varies with data availability, how uncertainty is communicated in both predictions and aggregate metrics, and how shifting data distributions impact model reliability. Presented as three studies, this dissertation develops principled statistical approaches for deep learning model evaluation that addresses these challenges in the context of time-series-based, scientific problems. The first study introduces an evaluation framework for seismic deep learning models where I assess learning efficiency while mitigating data leakage and quantify benchmark uncertainty by attributing variation to both training stochasticity and data sampling through an expansive design of experiments. The second study compares meta-learning techniques across data regimes and analyzes how consistently they perform under data shift. As part of this study, I contribute SeisTask, a semi-synthetic benchmark dataset with controlled, physically meaningful sources of shift for future study on adaptive learning approaches. The third study provides an empirical comparison of meta-learning and hierarchical Bayesian modeling and highlights their theoretical connection. I compare these methods in terms of interpretability, performance under shift, and predictive uncertainty. In combination, these studies offer statistically grounded evaluations of deep learning models for event detection in time series and show how uncertainty, data requirements, and distributional shift influence model behavior in physical science applications.","abstract_html":"Rapid developments in deep learning have led to their widespread use in domains that rely on time series, largely because of their strong performance and flexibility. Yet evaluation practices have not kept pace. Deep learning models are often assessed using a few performance metrics computed on benchmark datasets, which ignores important questions about how predictive performance varies with data availability, how uncertainty is communicated in both predictions and aggregate metrics, and how shifting data distributions impact model reliability. Presented as three studies, this dissertation develops principled statistical approaches for deep learning model evaluation that addresses these challenges in the context of time-series-based, scientific problems. The first study introduces an evaluation framework for seismic deep learning models where I assess learning efficiency while mitigating data leakage and quantify benchmark uncertainty by attributing variation to both training stochasticity and data sampling through an expansive design of experiments. The second study compares meta-learning techniques across data regimes and analyzes how consistently they perform under data shift. As part of this study, I contribute SeisTask, a semi-synthetic benchmark dataset with controlled, physically meaningful sources of shift for future study on adaptive learning approaches. The third study provides an empirical comparison of meta-learning and hierarchical Bayesian modeling and highlights their theoretical connection. I compare these methods in terms of interpretability, performance under shift, and predictive uncertainty. In combination, these studies offer statistically grounded evaluations of deep learning models for event detection in time series and show how uncertainty, data requirements, and distributional shift influence model behavior in physical science applications.","abstract_has_math":false,"creators":["Myren, Samuel Thomas Wilkins"],"institution":"Virginia Tech","degree_name":"Doctor of Philosophy","degree_level":"doctoral","degree_discipline":"Statistics","degree_department":"Statistics","school":null,"contributors":[],"advisors":[],"committee_chairs":["Higdon, David"],"committee_members":["Flynn, Garrison","Deng, Xinwei","House, Leanna L.","Parikh, Nidhi Kiranbhai"],"year":2026,"date_issued":"2026-01-14","date_published":"2026-01-14","updated_at":"2026-07-22T22:19:01Z","subjects":["Benchmark Variability","Data Leakage","Domain Adaptation","Experimental Design","Model Calibration"],"languages":["en"],"rights":["In Copyright"],"rights_urls":["http://rightsstatements.org/vocab/InC/1.0/"],"identifier_entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:45568"],"render_values":[{"text":"vt_gsexam:45568","href":null,"code":true}]}]},"links":{"outbound_url":"https://hdl.handle.net/10919/140814","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.committeechair","label":"Committee Chair","values":["Higdon, David"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Flynn, Garrison","Deng, Xinwei","House, Leanna L.","Parikh, Nidhi Kiranbhai"]},{"key":"dc:contributor.department","label":"Department","values":["Statistics"]},{"key":"dc:creator","label":"Author","values":["Myren, Samuel Thomas Wilkins"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2026-01-15T09:01:06Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2026-01-15T09:01:06Z"]},{"key":"dc:date.issued","label":"Date","values":["2026-01-14"]},{"key":"dc:publisher","label":"Institution","values":["Virginia Tech"]},{"key":"dc:type","label":"Dc Type","values":["Dissertation"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Statistics"]},{"key":"thesis:degree_level","label":"Degree Level","values":["doctoral"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctor of Philosophy"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Virginia Polytechnic Institute and State University"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Benchmark Variability","Data Leakage","Domain Adaptation","Experimental Design","Model Calibration"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["In Copyright"]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://rightsstatements.org/vocab/InC/1.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:45568"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10919/140814"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Rapid developments in deep learning have led to their widespread use in domains that rely on time series, largely because of their strong performance and flexibility. Yet evaluation practices have not kept pace. Deep learning models are often assessed using a few performance metrics computed on benchmark datasets, which ignores important questions about how predictive performance varies with data availability, how uncertainty is communicated in both predictions and aggregate metrics, and how shifting data distributions impact model reliability. Presented as three studies, this dissertation develops principled statistical approaches for deep learning model evaluation that addresses these challenges in the context of time-series-based, scientific problems. The first study introduces an evaluation framework for seismic deep learning models where I assess learning efficiency while mitigating data leakage and quantify benchmark uncertainty by attributing variation to both training stochasticity and data sampling through an expansive design of experiments. The second study compares meta-learning techniques across data regimes and analyzes how consistently they perform under data shift. As part of this study, I contribute SeisTask, a semi-synthetic benchmark dataset with controlled, physically meaningful sources of shift for future study on adaptive learning approaches. The third study provides an empirical comparison of meta-learning and hierarchical Bayesian modeling and highlights their theoretical connection. I compare these methods in terms of interpretability, performance under shift, and predictive uncertainty. In combination, these studies offer statistically grounded evaluations of deep learning models for event detection in time series and show how uncertainty, data requirements, and distributional shift influence model behavior in physical science applications."]},{"key":"dc:description.abstractgeneral","label":"General Abstract","values":["Deep learning models are widely used in the physical sciences to detect important events in time series, such as earthquakes recorded by seismic sensors. Although these models often perform well, it is not always clear how reliable they are, how much data they truly need, or how they behave when data change in unforeseen ways. These concerns matter in many scientific fields and have practical implications for earthquake monitoring, infrastructure safety, and other applications where accurate event detection is important. This dissertation studies how to evaluate deep learning models in more careful and informative ways. Here, evaluation refers to checking how well a model works, how stable its predictions are, and whether it can be trusted when conditions change. I examine how much model performance varies and why, how efficiently models learn from different amounts of data, and how consistently they perform when the test data no longer resemble what they saw during training. For example, a model trained on signals from one region may struggle when applied to another region with different noise conditions. I also build new datasets and tools that allow these questions to be explored in controlled and meaningful ways. Through a set of case studies in seismology, I compare modern modeling approaches and show which ones are more reliable and stable under different conditions. Overall, this work shows principled ways to compare deep learning models across relevant criteria and highlights how a single performance number hides much of what matters. Since deep learning models form the core of many modern artificial intelligence systems, their evaluation must reflect that nuance. By accounting for uncertainty, data requirements, and changes in the environment, we can better understand when these models can be trusted and how they can be improved for scientific applications."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["Doctor of Philosophy"]},{"key":"dc:format.medium","label":"Dc Format Medium","values":["ETD"]},{"key":"dc:title","label":"Title","values":["Statistical Evaluation of Deep Learning for Event Detection in Time Series: Quantifying Uncertainty, Efficiency, and Adaptation with Applications to Seismic Data"]}]}],"canonical_facts":{"dc:contributor.committeechair":["Higdon, David"],"dc:contributor.committeemember":["Flynn, Garrison","Deng, Xinwei","House, Leanna L.","Parikh, Nidhi Kiranbhai"],"dc:contributor.department":["Statistics"],"dc:creator":["Myren, Samuel Thomas Wilkins"],"dc:date.accessioned":["2026-01-15T09:01:06Z"],"dc:date.available":["2026-01-15T09:01:06Z"],"dc:date.issued":["2026-01-14"],"dc:description.abstract":["Rapid developments in deep learning have led to their widespread use in domains that rely on time series, largely because of their strong performance and flexibility. Yet evaluation practices have not kept pace. Deep learning models are often assessed using a few performance metrics computed on benchmark datasets, which ignores important questions about how predictive performance varies with data availability, how uncertainty is communicated in both predictions and aggregate metrics, and how shifting data distributions impact model reliability. Presented as three studies, this dissertation develops principled statistical approaches for deep learning model evaluation that addresses these challenges in the context of time-series-based, scientific problems. The first study introduces an evaluation framework for seismic deep learning models where I assess learning efficiency while mitigating data leakage and quantify benchmark uncertainty by attributing variation to both training stochasticity and data sampling through an expansive design of experiments. The second study compares meta-learning techniques across data regimes and analyzes how consistently they perform under data shift. As part of this study, I contribute SeisTask, a semi-synthetic benchmark dataset with controlled, physically meaningful sources of shift for future study on adaptive learning approaches. The third study provides an empirical comparison of meta-learning and hierarchical Bayesian modeling and highlights their theoretical connection. I compare these methods in terms of interpretability, performance under shift, and predictive uncertainty. In combination, these studies offer statistically grounded evaluations of deep learning models for event detection in time series and show how uncertainty, data requirements, and distributional shift influence model behavior in physical science applications."],"dc:description.abstractgeneral":["Deep learning models are widely used in the physical sciences to detect important events in time series, such as earthquakes recorded by seismic sensors. Although these models often perform well, it is not always clear how reliable they are, how much data they truly need, or how they behave when data change in unforeseen ways. These concerns matter in many scientific fields and have practical implications for earthquake monitoring, infrastructure safety, and other applications where accurate event detection is important. This dissertation studies how to evaluate deep learning models in more careful and informative ways. Here, evaluation refers to checking how well a model works, how stable its predictions are, and whether it can be trusted when conditions change. I examine how much model performance varies and why, how efficiently models learn from different amounts of data, and how consistently they perform when the test data no longer resemble what they saw during training. For example, a model trained on signals from one region may struggle when applied to another region with different noise conditions. I also build new datasets and tools that allow these questions to be explored in controlled and meaningful ways. Through a set of case studies in seismology, I compare modern modeling approaches and show which ones are more reliable and stable under different conditions. Overall, this work shows principled ways to compare deep learning models across relevant criteria and highlights how a single performance number hides much of what matters. Since deep learning models form the core of many modern artificial intelligence systems, their evaluation must reflect that nuance. By accounting for uncertainty, data requirements, and changes in the environment, we can better understand when these models can be trusted and how they can be improved for scientific applications."],"dc:description.degree":["Doctor of Philosophy"],"dc:format.medium":["ETD"],"dc:identifier.other":["vt_gsexam:45568"],"dc:identifier.uri":["https://hdl.handle.net/10919/140814"],"dc:language.iso":["en"],"dc:publisher":["Virginia Tech"],"dc:rights":["In Copyright"],"dc:rights.uri":["http://rightsstatements.org/vocab/InC/1.0/"],"dc:subject":["Benchmark Variability","Data Leakage","Domain Adaptation","Experimental Design","Model Calibration"],"dc:title":["Statistical Evaluation of Deep Learning for Event Detection in Time Series: Quantifying Uncertainty, Efficiency, and Adaptation with Applications to Seismic Data"],"dc:type":["Dissertation"],"thesis:degree_discipline":["Statistics"],"thesis:degree_level":["doctoral"],"thesis:degree_name":["Doctor of Philosophy"],"thesis:institution_name":["Virginia Polytechnic Institute and State University"]},"updated_at":"2026-07-22T22:19:01Z"}