{"id":{"repo_id":"cambridge","oai_identifier":"oai:www.repository.cam.ac.uk:1810/379315"},"canonical_url":"https://search.dev.ndltd.org/etd/cambridge/oai:www.repository.cam.ac.uk:1810/379315","repository":{"repo_id":"cambridge","name":"Cambridge University","base_url":"https://api.repository.cam.ac.uk/server/oai/request"},"display":{"title":"Symmetry and Other Structures: Topics in Nonparametric Regression","abstract":"In regression we measure covariate samples X_i and response variables Y_i and seek to estimate the conditional mean, or regression function, f(x) = E(Y_i ∣ X_i = x) within some hypothesis class F of possible regression functions. Estimation of a nonparametric regression function f ∶ [0, 1]^d → R becomes increasingly difficult as the dimension of the covariate space increases. Specifically, the number of samples n required to ensure that sup_{f∈F} E(∥f_n - f∥2) ≤ η for a fixed η > 0 grows exponentially in d for all estimators fˆ_n in many usual problem contexts. To overcome this, we often seek to find structure within the function f we can exploit. The two main structures that have been previously studied are Sparsity, where f depends only on s < d of the covariates’ coordinates, and Additivity, where f is a sum of functions on each of the covariates’ coordinates independently. In this thesis we present a new framework for lower dimensional structures that can capture much of the same benefits but now in non-linear situations. This framework exploits the structure of the level sets of the regression function. If we know the level set [x]_f = {u ∈ [0, 1]^d ∶ f(x) = f(u)} then we can average any sufficiently local estimator fn over this level set to obtain an estimator with a pointwise risk that is bounded in terms of d − dim[x]_f rather than d. In general it is hard to learn [x]_f ; however, even a simple plug in estimator (i.e., that uses the level sets of some estimate pn of f) can be shown to achieve these faster rates (up to log factors) under some strong conditions on f and demonstrate remarkable performance improvements in some finite sample experiments. Instead of placing conditions on f, this thesis gives sufficient conditions on structured hypothe- sis classes for the level sets [x]_f under which we can achieve the same fast rates of convergence. This key idea is to generalise sparsity, multi-index models, and other non-linear lower dimensional patterns in f in terms of transformations of the covariate space: Ψ_f ⊆ {Bijections ψ∶X→X} (1) for which f ○ ψ = f for all ψ ∈ Ψ_f . Functions that satisfy this property have Symmetry: in the same way that we can translate a sparse function along the unimportant axes we can transform f using Ψ_f without effect. Any such f must have level sets [x]_f that contain the sets {ψ(x) ∶ ψ ∈ Ψ_f }, which we can seek to exploit. We show that an oracle that knows any symmetry group Ψ_f of f performs with worst case L_2 risk decay that depends on a smaller dimension s ≤ d, rather than the dimension d of the original covariate space. The regression function f can have many possible symmetry groups, but when f is continuous there is always a unique maximal symmetry group Ψ_f within the classes considered. We then show that we can use an M-estimator of this maximal Ψ_f in order to achieve the dimension reduction adaptively under some technical conditions on Ψ_f and the hypothesis class of symmetries considered. We then show that we can give an inferential framework for this maximal Ψ, giving both hypothesis tests for: H_0(Ψ)∶f○ψ=f, for all ψ∈Ψ for any transformation group Ψ, as well as confidence regions for the maximal Ψ in the hypothesis class of possible Ψ we consider. This also leads to another estimator of Ψ that has the computational advantage of not estimating f directly. We demonstrate the performance of these tools both on synthetic datasets and on real data from both the Earth’s magnetic field and temporospatial locations of sunspots.","abstract_html":"In regression we measure covariate samples X_i and response variables Y_i and seek to estimate the conditional mean, or regression function, f(x) = E(Y_i ∣ X_i = x) within some hypothesis class F of possible regression functions. Estimation of a nonparametric regression function f ∶ [0, 1]^d → R becomes increasingly difficult as the dimension of the covariate space increases. Specifically, the number of samples n required to ensure that sup_{f∈F} E(∥f_n - f∥2) ≤ η for a fixed η &gt; 0 grows exponentially in d for all estimators fˆ_n in many usual problem contexts. To overcome this, we often seek to find structure within the function f we can exploit. The two main structures that have been previously studied are Sparsity, where f depends only on s &lt; d of the covariates’ coordinates, and Additivity, where f is a sum of functions on each of the covariates’ coordinates independently. In this thesis we present a new framework for lower dimensional structures that can capture much of the same benefits but now in non-linear situations. This framework exploits the structure of the level sets of the regression function. If we know the level set [x]_f = {u ∈ [0, 1]^d ∶ f(x) = f(u)} then we can average any sufficiently local estimator fn over this level set to obtain an estimator with a pointwise risk that is bounded in terms of d − dim[x]_f rather than d. In general it is hard to learn [x]_f ; however, even a simple plug in estimator (i.e., that uses the level sets of some estimate pn of f) can be shown to achieve these faster rates (up to log factors) under some strong conditions on f and demonstrate remarkable performance improvements in some finite sample experiments. Instead of placing conditions on f, this thesis gives sufficient conditions on structured hypothe- sis classes for the level sets [x]_f under which we can achieve the same fast rates of convergence. This key idea is to generalise sparsity, multi-index models, and other non-linear lower dimensional patterns in f in terms of transformations of the covariate space: Ψ_f ⊆ {Bijections ψ∶X→X} (1) for which f ○ ψ = f for all ψ ∈ Ψ_f . Functions that satisfy this property have Symmetry: in the same way that we can translate a sparse function along the unimportant axes we can transform f using Ψ_f without effect. Any such f must have level sets [x]_f that contain the sets {ψ(x) ∶ ψ ∈ Ψ_f }, which we can seek to exploit. We show that an oracle that knows any symmetry group Ψ_f of f performs with worst case L_2 risk decay that depends on a smaller dimension s ≤ d, rather than the dimension d of the original covariate space. The regression function f can have many possible symmetry groups, but when f is continuous there is always a unique maximal symmetry group Ψ_f within the classes considered. We then show that we can use an M-estimator of this maximal Ψ_f in order to achieve the dimension reduction adaptively under some technical conditions on Ψ_f and the hypothesis class of symmetries considered. We then show that we can give an inferential framework for this maximal Ψ, giving both hypothesis tests for: H_0(Ψ)∶f○ψ=f, for all ψ∈Ψ for any transformation group Ψ, as well as confidence regions for the maximal Ψ in the hypothesis class of possible Ψ we consider. This also leads to another estimator of Ψ that has the computational advantage of not estimating f directly. We demonstrate the performance of these tools both on synthetic datasets and on real data from both the Earth’s magnetic field and temporospatial locations of sunspots.","abstract_has_math":false,"creators":["Christie, Louis"],"institution":"University of Cambridge","degree_name":"Doctor of Philosophy (PhD)","degree_level":"Doctoral","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Aston, John"],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-09-07","date_published":"2024-09-07","updated_at":"2026-07-22T22:24:14Z","subjects":["Nonparametric Regression","Invariant Models","Dimension Reduction"],"languages":["eng"],"rights":[],"rights_urls":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/d8d3ff0e-4d41-4468-b6c6-f956edc7f28f/download","http://purl.org/NET/rdflicense/allrightsreserved"],"identifier_entries":[]},"links":{"outbound_url":"https://doi.org/10.17863/CAM.115448","outbound_label":"DOI","outbound_source":"dc:identifier.doi"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Aston, John"]},{"key":"dc:contributor.sponsor","label":"Sponsor","values":["The Harding Distinguished Postgraduate Scholars Programme"]},{"key":"dc:creator","label":"Author","values":["Christie, Louis"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2024-09-07"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["University of Cambridge"]},{"key":"dc:relation.isreferencedby.uri","label":"Dc Relation Isreferencedby URI","values":["https://www.repository.cam.ac.uk/handle/1810/379315"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["Doctoral"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["Doctor of Philosophy (PhD)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Nonparametric Regression","Invariant Models","Dimension Reduction"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/d8d3ff0e-4d41-4468-b6c6-f956edc7f28f/download","http://purl.org/NET/rdflicense/allrightsreserved"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.17863/CAM.115448"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/b7fabc31-7129-4927-a268-b9ad3a2eec32/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["In regression we measure covariate samples X_i and response variables Y_i and seek to estimate the conditional mean, or regression function, f(x) = E(Y_i ∣ X_i = x) within some hypothesis class F of possible regression functions. Estimation of a nonparametric regression function f ∶ [0, 1]^d → R becomes increasingly difficult as the dimension of the covariate space increases. Specifically, the number of samples n required to ensure that sup_{f∈F} E(∥f_n - f∥2) ≤ η for a fixed η > 0 grows exponentially in d for all estimators fˆ_n in many usual problem contexts. To overcome this, we often seek to find structure within the function f we can exploit. The two main structures that have been previously studied are Sparsity, where f depends only on s < d of the covariates’ coordinates, and Additivity, where f is a sum of functions on each of the covariates’ coordinates independently. In this thesis we present a new framework for lower dimensional structures that can capture much of the same benefits but now in non-linear situations. This framework exploits the structure of the level sets of the regression function. If we know the level set [x]_f = {u ∈ [0, 1]^d ∶ f(x) = f(u)} then we can average any sufficiently local estimator fn over this level set to obtain an estimator with a pointwise risk that is bounded in terms of d − dim[x]_f rather than d. In general it is hard to learn [x]_f ; however, even a simple plug in estimator (i.e., that uses the level sets of some estimate pn of f) can be shown to achieve these faster rates (up to log factors) under some strong conditions on f and demonstrate remarkable performance improvements in some finite sample experiments. Instead of placing conditions on f, this thesis gives sufficient conditions on structured hypothe- sis classes for the level sets [x]_f under which we can achieve the same fast rates of convergence. This key idea is to generalise sparsity, multi-index models, and other non-linear lower dimensional patterns in f in terms of transformations of the covariate space: Ψ_f ⊆ {Bijections ψ∶X→X} (1) for which f ○ ψ = f for all ψ ∈ Ψ_f . Functions that satisfy this property have Symmetry: in the same way that we can translate a sparse function along the unimportant axes we can transform f using Ψ_f without effect. Any such f must have level sets [x]_f that contain the sets {ψ(x) ∶ ψ ∈ Ψ_f }, which we can seek to exploit. We show that an oracle that knows any symmetry group Ψ_f of f performs with worst case L_2 risk decay that depends on a smaller dimension s ≤ d, rather than the dimension d of the original covariate space. The regression function f can have many possible symmetry groups, but when f is continuous there is always a unique maximal symmetry group Ψ_f within the classes considered. We then show that we can use an M-estimator of this maximal Ψ_f in order to achieve the dimension reduction adaptively under some technical conditions on Ψ_f and the hypothesis class of symmetries considered. We then show that we can give an inferential framework for this maximal Ψ, giving both hypothesis tests for: H_0(Ψ)∶f○ψ=f, for all ψ∈Ψ for any transformation group Ψ, as well as confidence regions for the maximal Ψ in the hypothesis class of possible Ψ we consider. This also leads to another estimator of Ψ that has the computational advantage of not estimating f directly. We demonstrate the performance of these tools both on synthetic datasets and on real data from both the Earth’s magnetic field and temporospatial locations of sunspots."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["4e8f72cf8754c70257a6c6fb0e0bd548","87eda9de84448d1f82354d60eee3eb5f"]},{"key":"dc:title","label":"Title","values":["Symmetry and Other Structures: Topics in Nonparametric Regression"]}]}],"canonical_facts":{"dc:contributor.advisor":["Aston, John"],"dc:contributor.sponsor":["The Harding Distinguished Postgraduate Scholars Programme"],"dc:creator":["Christie, Louis"],"dc:date.issued":["2024-09-07"],"dc:description.abstract":["In regression we measure covariate samples X_i and response variables Y_i and seek to estimate the conditional mean, or regression function, f(x) = E(Y_i ∣ X_i = x) within some hypothesis class F of possible regression functions. Estimation of a nonparametric regression function f ∶ [0, 1]^d → R becomes increasingly difficult as the dimension of the covariate space increases. Specifically, the number of samples n required to ensure that sup_{f∈F} E(∥f_n - f∥2) ≤ η for a fixed η > 0 grows exponentially in d for all estimators fˆ_n in many usual problem contexts. To overcome this, we often seek to find structure within the function f we can exploit. The two main structures that have been previously studied are Sparsity, where f depends only on s < d of the covariates’ coordinates, and Additivity, where f is a sum of functions on each of the covariates’ coordinates independently. In this thesis we present a new framework for lower dimensional structures that can capture much of the same benefits but now in non-linear situations. This framework exploits the structure of the level sets of the regression function. If we know the level set [x]_f = {u ∈ [0, 1]^d ∶ f(x) = f(u)} then we can average any sufficiently local estimator fn over this level set to obtain an estimator with a pointwise risk that is bounded in terms of d − dim[x]_f rather than d. In general it is hard to learn [x]_f ; however, even a simple plug in estimator (i.e., that uses the level sets of some estimate pn of f) can be shown to achieve these faster rates (up to log factors) under some strong conditions on f and demonstrate remarkable performance improvements in some finite sample experiments. Instead of placing conditions on f, this thesis gives sufficient conditions on structured hypothe- sis classes for the level sets [x]_f under which we can achieve the same fast rates of convergence. This key idea is to generalise sparsity, multi-index models, and other non-linear lower dimensional patterns in f in terms of transformations of the covariate space: Ψ_f ⊆ {Bijections ψ∶X→X} (1) for which f ○ ψ = f for all ψ ∈ Ψ_f . Functions that satisfy this property have Symmetry: in the same way that we can translate a sparse function along the unimportant axes we can transform f using Ψ_f without effect. Any such f must have level sets [x]_f that contain the sets {ψ(x) ∶ ψ ∈ Ψ_f }, which we can seek to exploit. We show that an oracle that knows any symmetry group Ψ_f of f performs with worst case L_2 risk decay that depends on a smaller dimension s ≤ d, rather than the dimension d of the original covariate space. The regression function f can have many possible symmetry groups, but when f is continuous there is always a unique maximal symmetry group Ψ_f within the classes considered. We then show that we can use an M-estimator of this maximal Ψ_f in order to achieve the dimension reduction adaptively under some technical conditions on Ψ_f and the hypothesis class of symmetries considered. We then show that we can give an inferential framework for this maximal Ψ, giving both hypothesis tests for: H_0(Ψ)∶f○ψ=f, for all ψ∈Ψ for any transformation group Ψ, as well as confidence regions for the maximal Ψ in the hypothesis class of possible Ψ we consider. This also leads to another estimator of Ψ that has the computational advantage of not estimating f directly. We demonstrate the performance of these tools both on synthetic datasets and on real data from both the Earth’s magnetic field and temporospatial locations of sunspots."],"dc:format.checksum.md5":["4e8f72cf8754c70257a6c6fb0e0bd548","87eda9de84448d1f82354d60eee3eb5f"],"dc:identifier.doi":["https://doi.org/10.17863/CAM.115448"],"dc:identifier.uri":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/b7fabc31-7129-4927-a268-b9ad3a2eec32/download"],"dc:language":["eng"],"dc:publisher.institution":["University of Cambridge"],"dc:relation.isreferencedby.uri":["https://www.repository.cam.ac.uk/handle/1810/379315"],"dc:rights":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/d8d3ff0e-4d41-4468-b6c6-f956edc7f28f/download","http://purl.org/NET/rdflicense/allrightsreserved"],"dc:subject":["Nonparametric Regression","Invariant Models","Dimension Reduction"],"dc:title":["Symmetry and Other Structures: Topics in Nonparametric Regression"],"dc:type":["Thesis"],"dc:type.qualificationlevel":["Doctoral"],"dc:type.qualificationname":["Doctor of Philosophy (PhD)"]},"updated_at":"2026-07-22T22:24:14Z"}