{"id":{"repo_id":"toronto-retro","oai_identifier":"oai:utoronto.scholaris.ca:1807/142699"},"canonical_url":"https://search.dev.ndltd.org/etd/toronto-retro/oai:utoronto.scholaris.ca:1807/142699","repository":{"repo_id":"toronto-retro","name":"University of Toronto","base_url":"https://utoronto.scholaris.ca/server/oai/request"},"display":{"title":"Beyond Gradients: Using Curvature Information for Deep Learning","abstract":"This thesis investigates optimization and interpretability techniques for deep learning, extending beyond gradient-based methods to incorporate curvature information. We begin by addressing the limitations of first-order optimization methods like gradient descent, which can exhibit slow convergence in directions of low curvature within the loss landscape. To address these challenges, we introduce a novel framework that dynamically learns and integrates curvature information throughout optimization, significantly accelerating convergence. Our investigation extends to bilevel problems, focusing on hyperparameter optimization and influence functions. Bilevel optimization involves nested optimization problems where the outer-level objective depends on the optimal solution to an inner-level optimization. In gradient-based bilevel optimization, a key challenge is accurately capturing how outer-level variable updates affect the inner-level optimization problem, where inner-level curvature information provides essential sensitivity measures. For hyperparameter optimization, we improve existing hypernetwork-based approaches for learning curvature information by addressing subtle training pathologies. Examining the bilevel formulation of influence functions reveals that their application to neural networks approximates a different quantity than previously believed. We identify key limitations, including their inability to account for optimizers' implicit bias or multi-stage training processes. We present an unrolled differentiation formulation to address these shortcomings, providing a more comprehensive framework for analyzing a training data point's impact on model predictions. The thesis concludes by discussing limitations and future research directions for the proposed techniques.","abstract_html":"This thesis investigates optimization and interpretability techniques for deep learning, extending beyond gradient-based methods to incorporate curvature information. We begin by addressing the limitations of first-order optimization methods like gradient descent, which can exhibit slow convergence in directions of low curvature within the loss landscape. To address these challenges, we introduce a novel framework that dynamically learns and integrates curvature information throughout optimization, significantly accelerating convergence. Our investigation extends to bilevel problems, focusing on hyperparameter optimization and influence functions. Bilevel optimization involves nested optimization problems where the outer-level objective depends on the optimal solution to an inner-level optimization. In gradient-based bilevel optimization, a key challenge is accurately capturing how outer-level variable updates affect the inner-level optimization problem, where inner-level curvature information provides essential sensitivity measures. For hyperparameter optimization, we improve existing hypernetwork-based approaches for learning curvature information by addressing subtle training pathologies. Examining the bilevel formulation of influence functions reveals that their application to neural networks approximates a different quantity than previously believed. We identify key limitations, including their inability to account for optimizers&#x27; implicit bias or multi-stage training processes. We present an unrolled differentiation formulation to address these shortcomings, providing a more comprehensive framework for analyzing a training data point&#x27;s impact on model predictions. The thesis concludes by discussing limitations and future research directions for the proposed techniques.","abstract_has_math":false,"creators":["Bae, Juhan"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":"Computer Science","school":null,"contributors":[],"advisors":["Grosse, Roger"],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-03","date_published":"2025-03","updated_at":"2026-07-27T21:28:07Z","subjects":[],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/1807/142699","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Grosse, Roger"]},{"key":"dc:contributor.department","label":"Department","values":["Computer Science"]},{"key":"dc:creator","label":"Author","values":["Bae, Juhan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-03"]},{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2025-04-18T15:20:43Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2025-04-18T15:20:43Z"]},{"key":"dc:date.issued","label":"Date","values":["2025-03"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/1807/142699"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["This thesis investigates optimization and interpretability techniques for deep learning, extending beyond gradient-based methods to incorporate curvature information. We begin by addressing the limitations of first-order optimization methods like gradient descent, which can exhibit slow convergence in directions of low curvature within the loss landscape. To address these challenges, we introduce a novel framework that dynamically learns and integrates curvature information throughout optimization, significantly accelerating convergence. Our investigation extends to bilevel problems, focusing on hyperparameter optimization and influence functions. Bilevel optimization involves nested optimization problems where the outer-level objective depends on the optimal solution to an inner-level optimization. In gradient-based bilevel optimization, a key challenge is accurately capturing how outer-level variable updates affect the inner-level optimization problem, where inner-level curvature information provides essential sensitivity measures. For hyperparameter optimization, we improve existing hypernetwork-based approaches for learning curvature information by addressing subtle training pathologies. Examining the bilevel formulation of influence functions reveals that their application to neural networks approximates a different quantity than previously believed. We identify key limitations, including their inability to account for optimizers' implicit bias or multi-stage training processes. We present an unrolled differentiation formulation to address these shortcomings, providing a more comprehensive framework for analyzing a training data point's impact on model predictions. The thesis concludes by discussing limitations and future research directions for the proposed techniques."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["Ph.D."]},{"key":"dc:title","label":"Title","values":["Beyond Gradients: Using Curvature Information for Deep Learning"]}]}],"canonical_facts":{"dc:contributor.advisor":["Grosse, Roger"],"dc:contributor.department":["Computer Science"],"dc:creator":["Bae, Juhan"],"dc:date":["2025-03"],"dc:date.accessioned":["2025-04-18T15:20:43Z"],"dc:date.available":["2025-04-18T15:20:43Z"],"dc:date.issued":["2025-03"],"dc:description.abstract":["This thesis investigates optimization and interpretability techniques for deep learning, extending beyond gradient-based methods to incorporate curvature information. We begin by addressing the limitations of first-order optimization methods like gradient descent, which can exhibit slow convergence in directions of low curvature within the loss landscape. To address these challenges, we introduce a novel framework that dynamically learns and integrates curvature information throughout optimization, significantly accelerating convergence. Our investigation extends to bilevel problems, focusing on hyperparameter optimization and influence functions. Bilevel optimization involves nested optimization problems where the outer-level objective depends on the optimal solution to an inner-level optimization. In gradient-based bilevel optimization, a key challenge is accurately capturing how outer-level variable updates affect the inner-level optimization problem, where inner-level curvature information provides essential sensitivity measures. For hyperparameter optimization, we improve existing hypernetwork-based approaches for learning curvature information by addressing subtle training pathologies. Examining the bilevel formulation of influence functions reveals that their application to neural networks approximates a different quantity than previously believed. We identify key limitations, including their inability to account for optimizers' implicit bias or multi-stage training processes. We present an unrolled differentiation formulation to address these shortcomings, providing a more comprehensive framework for analyzing a training data point's impact on model predictions. The thesis concludes by discussing limitations and future research directions for the proposed techniques."],"dc:description.degree":["Ph.D."],"dc:identifier.uri":["https://hdl.handle.net/1807/142699"],"dc:title":["Beyond Gradients: Using Curvature Information for Deep Learning"],"dc:type":["Thesis"]},"updated_at":"2026-07-27T21:28:07Z"}