{"id":{"repo_id":"cambridge","oai_identifier":"oai:www.repository.cam.ac.uk:1810/367808"},"canonical_url":"https://search.dev.ndltd.org/etd/cambridge/oai:www.repository.cam.ac.uk:1810/367808","repository":{"repo_id":"cambridge","name":"Cambridge University","base_url":"https://api.repository.cam.ac.uk/server/oai/request"},"display":{"title":"Scalable Bayesian Inference in the Era of Deep Learning: From Gaussian Processes to Deep Neural Networks","abstract":"Large neural networks trained on large datasets have become the dominant paradigm in machine learning. These systems rely on maximum likelihood point estimates of their parameters, precluding them from expressing model uncertainty. This may result in overconfident predictions and it prevents the use of deep learning models for sequential decision making. This thesis develops scalable methods to equip neural networks with model uncertainty. To achieve this, we do not try to fight progress in deep learning but instead borrow ideas from this field to make probabilistic methods more scalable. In particular, we leverage the linearised Laplace approximation to equip pre-trained neural networks with the uncertainty estimates provided by their tangent linear models. This turns the problem of Bayesian inference in neural networks into one of Bayesian inference in conjugate Gaussian-linear models. Alas, the cost of this remains cubic in either the number of network parameters or in the number of observations times output dimensions. By assumption, neither are tractable. We address this intractability by using stochastic gradient descent (SGD)---the workhorse algorithm of deep learning---to perform posterior sampling in linear models and their convex duals: Gaussian processes. With this, we turn back to linearised neural networks, finding the linearised Laplace approximation to present a number of incompatibilities with modern deep learning practices---namely, stochastic optimisation, early stopping and normalisation layers---when used for hyperparameter learning. We resolve these and construct a sample-based EM algorithm for scalable hyperparameter learning with linearised neural networks. We apply the above methods to perform linearised neural network inference with {ResNet-50} (25M parameters) trained on Imagenet (1.2M observations and 1000 output dimensions). To the best of our knowledge, this is the first time Bayesian inference has been performed in this real-world-scaled setting without assuming some degree of independence across network weights. Additionally, we apply our methods to estimate uncertainty for 3d tomographic reconstructions obtained with the deep image prior network, also a first. We conclude by using the linearised deep image prior to adaptively choose sequences of scanning angles that produce higher quality tomographic reconstructions while applying less radiation dosage.","abstract_html":"Large neural networks trained on large datasets have become the dominant paradigm in machine learning. These systems rely on maximum likelihood point estimates of their parameters, precluding them from expressing model uncertainty. This may result in overconfident predictions and it prevents the use of deep learning models for sequential decision making. This thesis develops scalable methods to equip neural networks with model uncertainty. To achieve this, we do not try to fight progress in deep learning but instead borrow ideas from this field to make probabilistic methods more scalable. In particular, we leverage the linearised Laplace approximation to equip pre-trained neural networks with the uncertainty estimates provided by their tangent linear models. This turns the problem of Bayesian inference in neural networks into one of Bayesian inference in conjugate Gaussian-linear models. Alas, the cost of this remains cubic in either the number of network parameters or in the number of observations times output dimensions. By assumption, neither are tractable. We address this intractability by using stochastic gradient descent (SGD)---the workhorse algorithm of deep learning---to perform posterior sampling in linear models and their convex duals: Gaussian processes. With this, we turn back to linearised neural networks, finding the linearised Laplace approximation to present a number of incompatibilities with modern deep learning practices---namely, stochastic optimisation, early stopping and normalisation layers---when used for hyperparameter learning. We resolve these and construct a sample-based EM algorithm for scalable hyperparameter learning with linearised neural networks. We apply the above methods to perform linearised neural network inference with {ResNet-50} (25M parameters) trained on Imagenet (1.2M observations and 1000 output dimensions). To the best of our knowledge, this is the first time Bayesian inference has been performed in this real-world-scaled setting without assuming some degree of independence across network weights. Additionally, we apply our methods to estimate uncertainty for 3d tomographic reconstructions obtained with the deep image prior network, also a first. We conclude by using the linearised deep image prior to adaptively choose sequences of scanning angles that produce higher quality tomographic reconstructions while applying less radiation dosage.","abstract_has_math":false,"creators":["Antoran, Javier"],"institution":"University of Cambridge","degree_name":"Doctor of Philosophy (PhD)","degree_level":"Doctoral","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Hernandez-Lobato, Jose Miguel"],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-02-09","date_published":"2024-02-09","updated_at":"2026-07-22T22:23:57Z","subjects":["Approximate Inference","Artificial Intelligence","Bayesian Inference","Computed tomography","Gaussian Process","Laplace Approximation","Linearised Laplace","Linearized Laplace","Machine Learning","Sample based inference","Stochastic Gradient Descent","Variational Inference"],"languages":["eng"],"rights":[],"rights_urls":["https://www.repository.cam.ac.uk/bitstreams/b5ef5ca7-30b7-4acd-aa78-ec2da15f85de/download","https://www.rioxx.net/licenses/all-rights-reserved/"],"identifier_entries":[]},"links":{"outbound_url":"https://doi.org/10.17863/CAM.108253","outbound_label":"DOI","outbound_source":"dc:identifier.doi"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Hernandez-Lobato, Jose Miguel"]},{"key":"dc:contributor.sponsor","label":"Sponsor","values":["Javier Antorán acknowledges support from Microsoft Research, through its PhD Scholarship Programme, and from the EPSRC"]},{"key":"dc:creator","label":"Author","values":["Antoran, Javier"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2024-02-09"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["University of Cambridge"]},{"key":"dc:relation.isreferencedby.uri","label":"Dc Relation Isreferencedby URI","values":["https://www.repository.cam.ac.uk/handle/1810/367808"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["Doctoral"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["Doctor of Philosophy (PhD)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Approximate Inference","Artificial Intelligence","Bayesian Inference","Computed tomography","Gaussian Process","Laplace Approximation","Linearised Laplace","Linearized Laplace","Machine Learning","Sample based inference","Stochastic Gradient Descent","Variational Inference"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["https://www.repository.cam.ac.uk/bitstreams/b5ef5ca7-30b7-4acd-aa78-ec2da15f85de/download","https://www.rioxx.net/licenses/all-rights-reserved/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.17863/CAM.108253"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://www.repository.cam.ac.uk/bitstreams/203dc060-fc67-4772-9980-e4c8cc7d0e69/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Large neural networks trained on large datasets have become the dominant paradigm in machine learning. These systems rely on maximum likelihood point estimates of their parameters, precluding them from expressing model uncertainty. This may result in overconfident predictions and it prevents the use of deep learning models for sequential decision making. This thesis develops scalable methods to equip neural networks with model uncertainty. To achieve this, we do not try to fight progress in deep learning but instead borrow ideas from this field to make probabilistic methods more scalable. In particular, we leverage the linearised Laplace approximation to equip pre-trained neural networks with the uncertainty estimates provided by their tangent linear models. This turns the problem of Bayesian inference in neural networks into one of Bayesian inference in conjugate Gaussian-linear models. Alas, the cost of this remains cubic in either the number of network parameters or in the number of observations times output dimensions. By assumption, neither are tractable. We address this intractability by using stochastic gradient descent (SGD)---the workhorse algorithm of deep learning---to perform posterior sampling in linear models and their convex duals: Gaussian processes. With this, we turn back to linearised neural networks, finding the linearised Laplace approximation to present a number of incompatibilities with modern deep learning practices---namely, stochastic optimisation, early stopping and normalisation layers---when used for hyperparameter learning. We resolve these and construct a sample-based EM algorithm for scalable hyperparameter learning with linearised neural networks. We apply the above methods to perform linearised neural network inference with {ResNet-50} (25M parameters) trained on Imagenet (1.2M observations and 1000 output dimensions). To the best of our knowledge, this is the first time Bayesian inference has been performed in this real-world-scaled setting without assuming some degree of independence across network weights. Additionally, we apply our methods to estimate uncertainty for 3d tomographic reconstructions obtained with the deep image prior network, also a first. We conclude by using the linearised deep image prior to adaptively choose sequences of scanning angles that produce higher quality tomographic reconstructions while applying less radiation dosage."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["325b021abb7c95c550ffa323d766d2a6","87eda9de84448d1f82354d60eee3eb5f"]},{"key":"dc:title","label":"Title","values":["Scalable Bayesian Inference in the Era of Deep Learning: From Gaussian Processes to Deep Neural Networks"]}]}],"canonical_facts":{"dc:contributor.advisor":["Hernandez-Lobato, Jose Miguel"],"dc:contributor.sponsor":["Javier Antorán acknowledges support from Microsoft Research, through its PhD Scholarship Programme, and from the EPSRC"],"dc:creator":["Antoran, Javier"],"dc:date.issued":["2024-02-09"],"dc:description.abstract":["Large neural networks trained on large datasets have become the dominant paradigm in machine learning. These systems rely on maximum likelihood point estimates of their parameters, precluding them from expressing model uncertainty. This may result in overconfident predictions and it prevents the use of deep learning models for sequential decision making. This thesis develops scalable methods to equip neural networks with model uncertainty. To achieve this, we do not try to fight progress in deep learning but instead borrow ideas from this field to make probabilistic methods more scalable. In particular, we leverage the linearised Laplace approximation to equip pre-trained neural networks with the uncertainty estimates provided by their tangent linear models. This turns the problem of Bayesian inference in neural networks into one of Bayesian inference in conjugate Gaussian-linear models. Alas, the cost of this remains cubic in either the number of network parameters or in the number of observations times output dimensions. By assumption, neither are tractable. We address this intractability by using stochastic gradient descent (SGD)---the workhorse algorithm of deep learning---to perform posterior sampling in linear models and their convex duals: Gaussian processes. With this, we turn back to linearised neural networks, finding the linearised Laplace approximation to present a number of incompatibilities with modern deep learning practices---namely, stochastic optimisation, early stopping and normalisation layers---when used for hyperparameter learning. We resolve these and construct a sample-based EM algorithm for scalable hyperparameter learning with linearised neural networks. We apply the above methods to perform linearised neural network inference with {ResNet-50} (25M parameters) trained on Imagenet (1.2M observations and 1000 output dimensions). To the best of our knowledge, this is the first time Bayesian inference has been performed in this real-world-scaled setting without assuming some degree of independence across network weights. Additionally, we apply our methods to estimate uncertainty for 3d tomographic reconstructions obtained with the deep image prior network, also a first. We conclude by using the linearised deep image prior to adaptively choose sequences of scanning angles that produce higher quality tomographic reconstructions while applying less radiation dosage."],"dc:format.checksum.md5":["325b021abb7c95c550ffa323d766d2a6","87eda9de84448d1f82354d60eee3eb5f"],"dc:identifier.doi":["https://doi.org/10.17863/CAM.108253"],"dc:identifier.uri":["https://www.repository.cam.ac.uk/bitstreams/203dc060-fc67-4772-9980-e4c8cc7d0e69/download"],"dc:language":["eng"],"dc:publisher.institution":["University of Cambridge"],"dc:relation.isreferencedby.uri":["https://www.repository.cam.ac.uk/handle/1810/367808"],"dc:rights":["https://www.repository.cam.ac.uk/bitstreams/b5ef5ca7-30b7-4acd-aa78-ec2da15f85de/download","https://www.rioxx.net/licenses/all-rights-reserved/"],"dc:subject":["Approximate Inference","Artificial Intelligence","Bayesian Inference","Computed tomography","Gaussian Process","Laplace Approximation","Linearised Laplace","Linearized Laplace","Machine Learning","Sample based inference","Stochastic Gradient Descent","Variational Inference"],"dc:title":["Scalable Bayesian Inference in the Era of Deep Learning: From Gaussian Processes to Deep Neural Networks"],"dc:type":["Thesis"],"dc:type.qualificationlevel":["Doctoral"],"dc:type.qualificationname":["Doctor of Philosophy (PhD)"]},"updated_at":"2026-07-22T22:23:57Z"}