{"id":{"repo_id":"cambridge","oai_identifier":"oai:www.repository.cam.ac.uk:1810/372888"},"canonical_url":"https://search.dev.ndltd.org/etd/cambridge/oai:www.repository.cam.ac.uk:1810/372888","repository":{"repo_id":"cambridge","name":"Cambridge University","base_url":"https://api.repository.cam.ac.uk/server/oai/request"},"display":{"title":"Advances in Optimisation of Model Parameters and Hyperparameters for Neural Networks","abstract":"Machine Learning has exploded in popularity in recent years, and now sees use in a huge variety of applications, from advertisement targeting and image generation to the recent proliferation of Large Language Models. At the heart of each setting is the problem of optimising model parameters to minimise some loss metric, so the chosen optimisation algorithm plays a fundamental role in the training process — both through the optimisation logic itself, and the auxiliary *hyperparameters* which configure the optimiser’s behaviour. Moreover, Machine Learning tasks often demand unique properties from optimisers, distinct from those explored in classical optimisation research. In this Thesis, we explore three novel techniques for optimising these parameters and hyperparameters, each sharing the theme of an awareness of the underlying curvature of the optimisation space. We begin with a study of Hyperparameter Optimisation. Many existing algorithms require complete training runs to evaluate each proposed hyperparameter configuration, which carry considerable computational cost. Methods based on hypergradients use only one training pass, but these either cannot be applied to arbitrary optimiser hyperparameters (such as learning rates and momenta) or suffer considerable additional training time. In an extension to these methods, we develop an approximate hypergradient-based hyperparameter optimiser, which is applicable to any continuous hyperparameter appearing in a differentiable model weight update. Our algorithm requires only one training episode (with no restarts), has a motivating argument for convergence to the true hypergradient, and scales to optimising independent learning rates for each model parameter, which we demonstrate in a wide-ranging empirical study. This contribution advances our ability to effectively configure a parameter optimiser for Neural Network training. Having developed a technique for selecting the hyperparameters controlling the behaviour of an optimisation algorithm, we turn to the model parameter optimisation task itself to explore which algorithmic properties might be useful in this application. The intractably-large Hessian matrices of many Machine Learning models make it challenging to apply the second-order quasi-Newton methods which are popular in the broader continuous optimisation literature. Attempts to address non-convexity in the problem, for instance by modifying eigenvalues as in Saddle-Free Newton methods, only exacerbate the issue of intractability. Addressing both these concerns, we propose the first (to our knowledge) efficiently-scalable optimisation algorithm to asymptotically use the exact, eigenvalue-modified inverse Hessian. Our method uses a power series to principally square-root and invert the squared Hessian, then precondition a gradient vector, all without explicitly computing or storing the Hessian. A truncation of this infinite series yields an optimisation algorithm which is competitive with common first- and second-order approaches, a claim which our experiments verify. However, while performing this work, we noted unexpected differences between the computational efficiency of first-order, gradient-based methods and the theoretical efficiency of second-order, curvature-based methods. In particular, we were intrigued by the relative strength and popularity of the Adam algorithm and the performant but fragile nature of the K-FAC algorithm. Inspired by this behaviour, we conclude this Thesis by seeking to unify the benefits of both approaches, combining the stabilising heuristics of second-order methods (such as Levenberg-Marquardt damping) with the efficient update direction selection of first-order methods. Our resulting algorithm recasts the Adam optimiser from a second-order optimisation perspective by adopting features from the K-FAC optimiser, and raises interesting questions about the dynamics of both popular methods; we explore these in an empirical validation over a range of settings. To conclude, we summarise our contributions and explore some possible future research directions they raise. We briefly consider the imminent challenges faced by the Machine Learning research community, and make a case for an ethical, human-centred emphasis on potential developments.","abstract_html":"Machine Learning has exploded in popularity in recent years, and now sees use in a huge variety of applications, from advertisement targeting and image generation to the recent proliferation of Large Language Models. At the heart of each setting is the problem of optimising model parameters to minimise some loss metric, so the chosen optimisation algorithm plays a fundamental role in the training process — both through the optimisation logic itself, and the auxiliary *hyperparameters* which configure the optimiser’s behaviour. Moreover, Machine Learning tasks often demand unique properties from optimisers, distinct from those explored in classical optimisation research. In this Thesis, we explore three novel techniques for optimising these parameters and hyperparameters, each sharing the theme of an awareness of the underlying curvature of the optimisation space. We begin with a study of Hyperparameter Optimisation. Many existing algorithms require complete training runs to evaluate each proposed hyperparameter configuration, which carry considerable computational cost. Methods based on hypergradients use only one training pass, but these either cannot be applied to arbitrary optimiser hyperparameters (such as learning rates and momenta) or suffer considerable additional training time. In an extension to these methods, we develop an approximate hypergradient-based hyperparameter optimiser, which is applicable to any continuous hyperparameter appearing in a differentiable model weight update. Our algorithm requires only one training episode (with no restarts), has a motivating argument for convergence to the true hypergradient, and scales to optimising independent learning rates for each model parameter, which we demonstrate in a wide-ranging empirical study. This contribution advances our ability to effectively configure a parameter optimiser for Neural Network training. Having developed a technique for selecting the hyperparameters controlling the behaviour of an optimisation algorithm, we turn to the model parameter optimisation task itself to explore which algorithmic properties might be useful in this application. The intractably-large Hessian matrices of many Machine Learning models make it challenging to apply the second-order quasi-Newton methods which are popular in the broader continuous optimisation literature. Attempts to address non-convexity in the problem, for instance by modifying eigenvalues as in Saddle-Free Newton methods, only exacerbate the issue of intractability. Addressing both these concerns, we propose the first (to our knowledge) efficiently-scalable optimisation algorithm to asymptotically use the exact, eigenvalue-modified inverse Hessian. Our method uses a power series to principally square-root and invert the squared Hessian, then precondition a gradient vector, all without explicitly computing or storing the Hessian. A truncation of this infinite series yields an optimisation algorithm which is competitive with common first- and second-order approaches, a claim which our experiments verify. However, while performing this work, we noted unexpected differences between the computational efficiency of first-order, gradient-based methods and the theoretical efficiency of second-order, curvature-based methods. In particular, we were intrigued by the relative strength and popularity of the Adam algorithm and the performant but fragile nature of the K-FAC algorithm. Inspired by this behaviour, we conclude this Thesis by seeking to unify the benefits of both approaches, combining the stabilising heuristics of second-order methods (such as Levenberg-Marquardt damping) with the efficient update direction selection of first-order methods. Our resulting algorithm recasts the Adam optimiser from a second-order optimisation perspective by adopting features from the K-FAC optimiser, and raises interesting questions about the dynamics of both popular methods; we explore these in an empirical validation over a range of settings. To conclude, we summarise our contributions and explore some possible future research directions they raise. We briefly consider the imminent challenges faced by the Machine Learning research community, and make a case for an ethical, human-centred emphasis on potential developments.","abstract_has_math":false,"creators":["Clarke, Ross MacKinnon"],"institution":"University of Cambridge","degree_name":"Doctor of Philosophy (PhD)","degree_level":"Doctoral","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Hernández-Lobato, José Miguel"],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-01-03","date_published":"2024-01-03","updated_at":"2026-07-22T22:24:14Z","subjects":["Adam","Artificial Intelligence","Hyperparameter Optimisation","K-FAC","Machine Learning","Neural Networks","Optimisation"],"languages":["eng"],"rights":[],"rights_urls":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/544c547a-8595-4819-8cd6-f305e8ff5509/download","https://www.rioxx.net/licenses/all-rights-reserved/"],"identifier_entries":[]},"links":{"outbound_url":"https://doi.org/10.17863/CAM.111541","outbound_label":"DOI","outbound_source":"dc:identifier.doi"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Hernández-Lobato, José Miguel"]},{"key":"dc:contributor.sponsor","label":"Sponsor","values":["Funding from Google DeepMind via a grant awarded to my Supervisor"]},{"key":"dc:creator","label":"Author","values":["Clarke, Ross MacKinnon"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2024-01-03"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["University of Cambridge"]},{"key":"dc:relation.isreferencedby.uri","label":"Dc Relation Isreferencedby URI","values":["https://www.repository.cam.ac.uk/handle/1810/372888"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["Doctoral"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["Doctor of Philosophy (PhD)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Adam","Artificial Intelligence","Hyperparameter Optimisation","K-FAC","Machine Learning","Neural Networks","Optimisation"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/544c547a-8595-4819-8cd6-f305e8ff5509/download","https://www.rioxx.net/licenses/all-rights-reserved/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.17863/CAM.111541"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/a5ff129d-3a07-4e06-8e71-c2a909c6c5f2/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Machine Learning has exploded in popularity in recent years, and now sees use in a huge variety of applications, from advertisement targeting and image generation to the recent proliferation of Large Language Models. At the heart of each setting is the problem of optimising model parameters to minimise some loss metric, so the chosen optimisation algorithm plays a fundamental role in the training process — both through the optimisation logic itself, and the auxiliary *hyperparameters* which configure the optimiser’s behaviour. Moreover, Machine Learning tasks often demand unique properties from optimisers, distinct from those explored in classical optimisation research. In this Thesis, we explore three novel techniques for optimising these parameters and hyperparameters, each sharing the theme of an awareness of the underlying curvature of the optimisation space. We begin with a study of Hyperparameter Optimisation. Many existing algorithms require complete training runs to evaluate each proposed hyperparameter configuration, which carry considerable computational cost. Methods based on hypergradients use only one training pass, but these either cannot be applied to arbitrary optimiser hyperparameters (such as learning rates and momenta) or suffer considerable additional training time. In an extension to these methods, we develop an approximate hypergradient-based hyperparameter optimiser, which is applicable to any continuous hyperparameter appearing in a differentiable model weight update. Our algorithm requires only one training episode (with no restarts), has a motivating argument for convergence to the true hypergradient, and scales to optimising independent learning rates for each model parameter, which we demonstrate in a wide-ranging empirical study. This contribution advances our ability to effectively configure a parameter optimiser for Neural Network training. Having developed a technique for selecting the hyperparameters controlling the behaviour of an optimisation algorithm, we turn to the model parameter optimisation task itself to explore which algorithmic properties might be useful in this application. The intractably-large Hessian matrices of many Machine Learning models make it challenging to apply the second-order quasi-Newton methods which are popular in the broader continuous optimisation literature. Attempts to address non-convexity in the problem, for instance by modifying eigenvalues as in Saddle-Free Newton methods, only exacerbate the issue of intractability. Addressing both these concerns, we propose the first (to our knowledge) efficiently-scalable optimisation algorithm to asymptotically use the exact, eigenvalue-modified inverse Hessian. Our method uses a power series to principally square-root and invert the squared Hessian, then precondition a gradient vector, all without explicitly computing or storing the Hessian. A truncation of this infinite series yields an optimisation algorithm which is competitive with common first- and second-order approaches, a claim which our experiments verify. However, while performing this work, we noted unexpected differences between the computational efficiency of first-order, gradient-based methods and the theoretical efficiency of second-order, curvature-based methods. In particular, we were intrigued by the relative strength and popularity of the Adam algorithm and the performant but fragile nature of the K-FAC algorithm. Inspired by this behaviour, we conclude this Thesis by seeking to unify the benefits of both approaches, combining the stabilising heuristics of second-order methods (such as Levenberg-Marquardt damping) with the efficient update direction selection of first-order methods. Our resulting algorithm recasts the Adam optimiser from a second-order optimisation perspective by adopting features from the K-FAC optimiser, and raises interesting questions about the dynamics of both popular methods; we explore these in an empirical validation over a range of settings. To conclude, we summarise our contributions and explore some possible future research directions they raise. We briefly consider the imminent challenges faced by the Machine Learning research community, and make a case for an ethical, human-centred emphasis on potential developments."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["74f529d222cb1b4afdf5984b5f48ba93","87eda9de84448d1f82354d60eee3eb5f"]},{"key":"dc:title","label":"Title","values":["Advances in Optimisation of Model Parameters and Hyperparameters for Neural Networks"]}]}],"canonical_facts":{"dc:contributor.advisor":["Hernández-Lobato, José Miguel"],"dc:contributor.sponsor":["Funding from Google DeepMind via a grant awarded to my Supervisor"],"dc:creator":["Clarke, Ross MacKinnon"],"dc:date.issued":["2024-01-03"],"dc:description.abstract":["Machine Learning has exploded in popularity in recent years, and now sees use in a huge variety of applications, from advertisement targeting and image generation to the recent proliferation of Large Language Models. At the heart of each setting is the problem of optimising model parameters to minimise some loss metric, so the chosen optimisation algorithm plays a fundamental role in the training process — both through the optimisation logic itself, and the auxiliary *hyperparameters* which configure the optimiser’s behaviour. Moreover, Machine Learning tasks often demand unique properties from optimisers, distinct from those explored in classical optimisation research. In this Thesis, we explore three novel techniques for optimising these parameters and hyperparameters, each sharing the theme of an awareness of the underlying curvature of the optimisation space. We begin with a study of Hyperparameter Optimisation. Many existing algorithms require complete training runs to evaluate each proposed hyperparameter configuration, which carry considerable computational cost. Methods based on hypergradients use only one training pass, but these either cannot be applied to arbitrary optimiser hyperparameters (such as learning rates and momenta) or suffer considerable additional training time. In an extension to these methods, we develop an approximate hypergradient-based hyperparameter optimiser, which is applicable to any continuous hyperparameter appearing in a differentiable model weight update. Our algorithm requires only one training episode (with no restarts), has a motivating argument for convergence to the true hypergradient, and scales to optimising independent learning rates for each model parameter, which we demonstrate in a wide-ranging empirical study. This contribution advances our ability to effectively configure a parameter optimiser for Neural Network training. Having developed a technique for selecting the hyperparameters controlling the behaviour of an optimisation algorithm, we turn to the model parameter optimisation task itself to explore which algorithmic properties might be useful in this application. The intractably-large Hessian matrices of many Machine Learning models make it challenging to apply the second-order quasi-Newton methods which are popular in the broader continuous optimisation literature. Attempts to address non-convexity in the problem, for instance by modifying eigenvalues as in Saddle-Free Newton methods, only exacerbate the issue of intractability. Addressing both these concerns, we propose the first (to our knowledge) efficiently-scalable optimisation algorithm to asymptotically use the exact, eigenvalue-modified inverse Hessian. Our method uses a power series to principally square-root and invert the squared Hessian, then precondition a gradient vector, all without explicitly computing or storing the Hessian. A truncation of this infinite series yields an optimisation algorithm which is competitive with common first- and second-order approaches, a claim which our experiments verify. However, while performing this work, we noted unexpected differences between the computational efficiency of first-order, gradient-based methods and the theoretical efficiency of second-order, curvature-based methods. In particular, we were intrigued by the relative strength and popularity of the Adam algorithm and the performant but fragile nature of the K-FAC algorithm. Inspired by this behaviour, we conclude this Thesis by seeking to unify the benefits of both approaches, combining the stabilising heuristics of second-order methods (such as Levenberg-Marquardt damping) with the efficient update direction selection of first-order methods. Our resulting algorithm recasts the Adam optimiser from a second-order optimisation perspective by adopting features from the K-FAC optimiser, and raises interesting questions about the dynamics of both popular methods; we explore these in an empirical validation over a range of settings. To conclude, we summarise our contributions and explore some possible future research directions they raise. We briefly consider the imminent challenges faced by the Machine Learning research community, and make a case for an ethical, human-centred emphasis on potential developments."],"dc:format.checksum.md5":["74f529d222cb1b4afdf5984b5f48ba93","87eda9de84448d1f82354d60eee3eb5f"],"dc:identifier.doi":["https://doi.org/10.17863/CAM.111541"],"dc:identifier.uri":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/a5ff129d-3a07-4e06-8e71-c2a909c6c5f2/download"],"dc:language":["eng"],"dc:publisher.institution":["University of Cambridge"],"dc:relation.isreferencedby.uri":["https://www.repository.cam.ac.uk/handle/1810/372888"],"dc:rights":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/544c547a-8595-4819-8cd6-f305e8ff5509/download","https://www.rioxx.net/licenses/all-rights-reserved/"],"dc:subject":["Adam","Artificial Intelligence","Hyperparameter Optimisation","K-FAC","Machine Learning","Neural Networks","Optimisation"],"dc:title":["Advances in Optimisation of Model Parameters and Hyperparameters for Neural Networks"],"dc:type":["Thesis"],"dc:type.qualificationlevel":["Doctoral"],"dc:type.qualificationname":["Doctor of Philosophy (PhD)"]},"updated_at":"2026-07-22T22:24:14Z"}