{"id":{"repo_id":"toronto-retro","oai_identifier":"oai:utoronto.scholaris.ca:1807/150082"},"canonical_url":"https://search.dev.ndltd.org/etd/toronto-retro/oai:utoronto.scholaris.ca:1807/150082","repository":{"repo_id":"toronto-retro","name":"University of Toronto","base_url":"https://utoronto.scholaris.ca/server/oai/request"},"display":{"title":"Empirical Approaches to Challenges in Neural Network Training and Deployment","abstract":"Training and deploying neural networks is a challenging endeavor. During the training phase, we need to decide on what optimizer to use, the exact data to train on, the training curriculum, and many other factors-- all of which affect the final performance. Once the model is trained, we need to make sure that it is ready to be used and works as intended.This thesis aims to address these challenges in two parts. The first part focuses on gaining insights into what changes in the training pipeline affect the final performance. We start by showing empirically that more popular adaptive optimizers do not perform worse than optimizers they can approximate with a practical amount of hyperparameter tuning, which contrasts with conventional wisdom that adaptive methods generalize worse than non-adaptive methods. We next focus on multi-task learning, and show that using a particular ordering of training data can improve performance on the data-scarce tasks that are prone to overfitting. The second part focuses on the goal of understanding models' capabilities and risks. One capability that we study in detail is large language models' ability to piece together latent knowledge scattered across training documents, and use this knowledge without explicitly reasoning about it. We argue that this behavior needs to be studied and monitored due to potential safety risks of models uncovering and using implicit harmful knowledge without it being noticed by us. Finally, we introduce a framework for verifying the training data of a model, which can provide important information for determining the capabilities and potential risks of the model.","abstract_html":"Training and deploying neural networks is a challenging endeavor. During the training phase, we need to decide on what optimizer to use, the exact data to train on, the training curriculum, and many other factors-- all of which affect the final performance. Once the model is trained, we need to make sure that it is ready to be used and works as intended.This thesis aims to address these challenges in two parts. The first part focuses on gaining insights into what changes in the training pipeline affect the final performance. We start by showing empirically that more popular adaptive optimizers do not perform worse than optimizers they can approximate with a practical amount of hyperparameter tuning, which contrasts with conventional wisdom that adaptive methods generalize worse than non-adaptive methods. We next focus on multi-task learning, and show that using a particular ordering of training data can improve performance on the data-scarce tasks that are prone to overfitting. The second part focuses on the goal of understanding models&#x27; capabilities and risks. One capability that we study in detail is large language models&#x27; ability to piece together latent knowledge scattered across training documents, and use this knowledge without explicitly reasoning about it. We argue that this behavior needs to be studied and monitored due to potential safety risks of models uncovering and using implicit harmful knowledge without it being noticed by us. Finally, we introduce a framework for verifying the training data of a model, which can provide important information for determining the capabilities and potential risks of the model.","abstract_has_math":false,"creators":["Choi, Dami"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":"Computer Science","school":null,"contributors":[],"advisors":["Duvenaud, David K","Maddison, Chris J"],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-10","date_published":"2025-10","updated_at":"2026-07-27T21:28:20Z","subjects":["Machine Learning"],"languages":[],"rights":["Attribution 4.0 International"],"rights_urls":["http://creativecommons.org/licenses/by/4.0/"],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/1807/150082","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Duvenaud, David K","Maddison, Chris J"]},{"key":"dc:contributor.department","label":"Department","values":["Computer Science"]},{"key":"dc:creator","label":"Author","values":["Choi, Dami"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-10"]},{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2025-11-28T17:04:39Z"]},{"key":"dc:date.issued","label":"Date","values":["2025-10"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Machine Learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["Attribution 4.0 International"]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://creativecommons.org/licenses/by/4.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/1807/150082"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Training and deploying neural networks is a challenging endeavor. During the training phase, we need to decide on what optimizer to use, the exact data to train on, the training curriculum, and many other factors-- all of which affect the final performance. Once the model is trained, we need to make sure that it is ready to be used and works as intended.This thesis aims to address these challenges in two parts. The first part focuses on gaining insights into what changes in the training pipeline affect the final performance. We start by showing empirically that more popular adaptive optimizers do not perform worse than optimizers they can approximate with a practical amount of hyperparameter tuning, which contrasts with conventional wisdom that adaptive methods generalize worse than non-adaptive methods. We next focus on multi-task learning, and show that using a particular ordering of training data can improve performance on the data-scarce tasks that are prone to overfitting. The second part focuses on the goal of understanding models' capabilities and risks. One capability that we study in detail is large language models' ability to piece together latent knowledge scattered across training documents, and use this knowledge without explicitly reasoning about it. We argue that this behavior needs to be studied and monitored due to potential safety risks of models uncovering and using implicit harmful knowledge without it being noticed by us. Finally, we introduce a framework for verifying the training data of a model, which can provide important information for determining the capabilities and potential risks of the model."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["Ph.D."]},{"key":"dc:title","label":"Title","values":["Empirical Approaches to Challenges in Neural Network Training and Deployment"]}]}],"canonical_facts":{"dc:contributor.advisor":["Duvenaud, David K","Maddison, Chris J"],"dc:contributor.department":["Computer Science"],"dc:creator":["Choi, Dami"],"dc:date":["2025-10"],"dc:date.accessioned":["2025-11-28T17:04:39Z"],"dc:date.issued":["2025-10"],"dc:description.abstract":["Training and deploying neural networks is a challenging endeavor. During the training phase, we need to decide on what optimizer to use, the exact data to train on, the training curriculum, and many other factors-- all of which affect the final performance. Once the model is trained, we need to make sure that it is ready to be used and works as intended.This thesis aims to address these challenges in two parts. The first part focuses on gaining insights into what changes in the training pipeline affect the final performance. We start by showing empirically that more popular adaptive optimizers do not perform worse than optimizers they can approximate with a practical amount of hyperparameter tuning, which contrasts with conventional wisdom that adaptive methods generalize worse than non-adaptive methods. We next focus on multi-task learning, and show that using a particular ordering of training data can improve performance on the data-scarce tasks that are prone to overfitting. The second part focuses on the goal of understanding models' capabilities and risks. One capability that we study in detail is large language models' ability to piece together latent knowledge scattered across training documents, and use this knowledge without explicitly reasoning about it. We argue that this behavior needs to be studied and monitored due to potential safety risks of models uncovering and using implicit harmful knowledge without it being noticed by us. Finally, we introduce a framework for verifying the training data of a model, which can provide important information for determining the capabilities and potential risks of the model."],"dc:description.degree":["Ph.D."],"dc:identifier.uri":["https://hdl.handle.net/1807/150082"],"dc:rights":["Attribution 4.0 International"],"dc:rights.uri":["http://creativecommons.org/licenses/by/4.0/"],"dc:subject":["Machine Learning"],"dc:title":["Empirical Approaches to Challenges in Neural Network Training and Deployment"],"dc:type":["Thesis"]},"updated_at":"2026-07-27T21:28:20Z"}