{"id":{"repo_id":"cambridge","oai_identifier":"oai:www.repository.cam.ac.uk:1810/380985"},"canonical_url":"https://search.dev.ndltd.org/etd/cambridge/oai:www.repository.cam.ac.uk:1810/380985","repository":{"repo_id":"cambridge","name":"Cambridge University","base_url":"https://api.repository.cam.ac.uk/server/oai/request"},"display":{"title":"Dynamic factor analysis with dependent Gaussian processes for high-dimensional biomarker trajectories","abstract":"The increasing availability of high-dimensional, longitudinal measures of biomarkers can facilitate understanding of biological mechanisms, as required for precision medicine. Biological knowledge suggests that it may be best to describe complex diseases at the level of underlying pathways, which may interact with one another. We propose a Bayesian approach that allows for characterising such correlation among different pathways through dependent Gaussian processes (DGP) and mapping the observed high-dimensional gene expression trajectories into unobserved low-dimensional pathway expression trajectories via Bayesian sparse factor analysis in Chapter 2. Our proposal is the first attempt to relax the classical assumption of independent factors for longitudinal data and has demonstrated a superior performance in recovering the shape of pathway expression trajectories, revealing the relationships between biomarkers and pathways, and predicting biomarker expressions (closer point estimates and narrower predictive intervals), as demonstrated through simulations and application to a real-world H3N2 dataset. To fit the model, we propose a Monte Carlo expectation maximization (MCEM) scheme that can be implemented conveniently by combining a standard Markov Chain Monte Carlo sampler and an R package GPFDA, which returns the maximum likelihood estimates of DGP hyperparameters. The modular structure of MCEM makes it generalizable to other complex models involving the DGP model component. Our R package DGP4LCF that implements the proposed approach is available on CRAN. We extend the proposed method in two ways in Chapter 3, so that it is applicable to a wider range of settings. First, to address the overfitting issue of DGP caused by sparsely collected longitudinal data (i.e., number of observed times is few), we introduce a roughness penalty upon DGP to guarantee an appropriate level of smoothness. Second, we develop a stochastic expectation maximization algorithm to further speed up the inference of DGP hyperparameters. These improvements can be naturally incorporated into our initial analysis framework in Chapter 2, and have proven useful in practice, as we will show in the simulation study and application to a real-world COVID-19 data application in Chapter 3. Application to measurements of metabolites in this example discovers a biomarker named taurine that has been receiving an increasing amount of attention clinically, yet its role has been ignored in a previous analysis of the data. We have developed an R package DFA4SIL (dynamic factor analysis for sparse and irregular longitudinal data) that implements the proposed method. We further extend our model by incorporating clinical outcomes into analysis in Chapter 4. Compared to models without information on clinical labels, this outcome-guided factor analysis is able to explicitly identify pathways predictive of the clinical outcome of interest and improve the estimation of the underlying factor structure, as we will show through an application to COVID-19 data in Chapter 4.","abstract_html":"The increasing availability of high-dimensional, longitudinal measures of biomarkers can facilitate understanding of biological mechanisms, as required for precision medicine. Biological knowledge suggests that it may be best to describe complex diseases at the level of underlying pathways, which may interact with one another. We propose a Bayesian approach that allows for characterising such correlation among different pathways through dependent Gaussian processes (DGP) and mapping the observed high-dimensional gene expression trajectories into unobserved low-dimensional pathway expression trajectories via Bayesian sparse factor analysis in Chapter 2. Our proposal is the first attempt to relax the classical assumption of independent factors for longitudinal data and has demonstrated a superior performance in recovering the shape of pathway expression trajectories, revealing the relationships between biomarkers and pathways, and predicting biomarker expressions (closer point estimates and narrower predictive intervals), as demonstrated through simulations and application to a real-world H3N2 dataset. To fit the model, we propose a Monte Carlo expectation maximization (MCEM) scheme that can be implemented conveniently by combining a standard Markov Chain Monte Carlo sampler and an R package GPFDA, which returns the maximum likelihood estimates of DGP hyperparameters. The modular structure of MCEM makes it generalizable to other complex models involving the DGP model component. Our R package DGP4LCF that implements the proposed approach is available on CRAN. We extend the proposed method in two ways in Chapter 3, so that it is applicable to a wider range of settings. First, to address the overfitting issue of DGP caused by sparsely collected longitudinal data (i.e., number of observed times is few), we introduce a roughness penalty upon DGP to guarantee an appropriate level of smoothness. Second, we develop a stochastic expectation maximization algorithm to further speed up the inference of DGP hyperparameters. These improvements can be naturally incorporated into our initial analysis framework in Chapter 2, and have proven useful in practice, as we will show in the simulation study and application to a real-world COVID-19 data application in Chapter 3. Application to measurements of metabolites in this example discovers a biomarker named taurine that has been receiving an increasing amount of attention clinically, yet its role has been ignored in a previous analysis of the data. We have developed an R package DFA4SIL (dynamic factor analysis for sparse and irregular longitudinal data) that implements the proposed method. We further extend our model by incorporating clinical outcomes into analysis in Chapter 4. Compared to models without information on clinical labels, this outcome-guided factor analysis is able to explicitly identify pathways predictive of the clinical outcome of interest and improve the estimation of the underlying factor structure, as we will show through an application to COVID-19 data in Chapter 4.","abstract_has_math":false,"creators":["Cai, Jiachen"],"institution":"University of Cambridge","degree_name":"Doctor of Philosophy (PhD)","degree_level":"Doctoral","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Tom, Brian","Goudie, Robert"],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-08-27","date_published":"2024-08-27","updated_at":"2026-07-22T22:24:06Z","subjects":["Dependent Gaussian processes","High-dimensional biomarker trajectories","Monte Carlo expectation maximization","Multivariate longitudinal data","Pathways","Sparse factor analysis","Stochastic expectation maximization"],"languages":["eng"],"rights":[],"rights_urls":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/4ffe66d3-4e6e-4df1-84ee-34d2ebb5a563/download","http://purl.org/NET/rdflicense/allrightsreserved"],"identifier_entries":[]},"links":{"outbound_url":"https://doi.org/10.17863/CAM.116381","outbound_label":"DOI","outbound_source":"dc:identifier.doi"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Tom, Brian","Goudie, Robert"]},{"key":"dc:contributor.sponsor","label":"Sponsor","values":["Medical Research Council Studentship"]},{"key":"dc:creator","label":"Author","values":["Cai, Jiachen"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2024-08-27"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["University of Cambridge"]},{"key":"dc:relation.isreferencedby.uri","label":"Dc Relation Isreferencedby URI","values":["https://www.repository.cam.ac.uk/handle/1810/380985"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["Doctoral"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["Doctor of Philosophy (PhD)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Dependent Gaussian processes","High-dimensional biomarker trajectories","Monte Carlo expectation maximization","Multivariate longitudinal data","Pathways","Sparse factor analysis","Stochastic expectation maximization"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/4ffe66d3-4e6e-4df1-84ee-34d2ebb5a563/download","http://purl.org/NET/rdflicense/allrightsreserved"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.17863/CAM.116381"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/1b7792f7-2691-4e22-b36b-9c6c3ed941f8/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["The increasing availability of high-dimensional, longitudinal measures of biomarkers can facilitate understanding of biological mechanisms, as required for precision medicine. Biological knowledge suggests that it may be best to describe complex diseases at the level of underlying pathways, which may interact with one another. We propose a Bayesian approach that allows for characterising such correlation among different pathways through dependent Gaussian processes (DGP) and mapping the observed high-dimensional gene expression trajectories into unobserved low-dimensional pathway expression trajectories via Bayesian sparse factor analysis in Chapter 2. Our proposal is the first attempt to relax the classical assumption of independent factors for longitudinal data and has demonstrated a superior performance in recovering the shape of pathway expression trajectories, revealing the relationships between biomarkers and pathways, and predicting biomarker expressions (closer point estimates and narrower predictive intervals), as demonstrated through simulations and application to a real-world H3N2 dataset. To fit the model, we propose a Monte Carlo expectation maximization (MCEM) scheme that can be implemented conveniently by combining a standard Markov Chain Monte Carlo sampler and an R package GPFDA, which returns the maximum likelihood estimates of DGP hyperparameters. The modular structure of MCEM makes it generalizable to other complex models involving the DGP model component. Our R package DGP4LCF that implements the proposed approach is available on CRAN. We extend the proposed method in two ways in Chapter 3, so that it is applicable to a wider range of settings. First, to address the overfitting issue of DGP caused by sparsely collected longitudinal data (i.e., number of observed times is few), we introduce a roughness penalty upon DGP to guarantee an appropriate level of smoothness. Second, we develop a stochastic expectation maximization algorithm to further speed up the inference of DGP hyperparameters. These improvements can be naturally incorporated into our initial analysis framework in Chapter 2, and have proven useful in practice, as we will show in the simulation study and application to a real-world COVID-19 data application in Chapter 3. Application to measurements of metabolites in this example discovers a biomarker named taurine that has been receiving an increasing amount of attention clinically, yet its role has been ignored in a previous analysis of the data. We have developed an R package DFA4SIL (dynamic factor analysis for sparse and irregular longitudinal data) that implements the proposed method. We further extend our model by incorporating clinical outcomes into analysis in Chapter 4. Compared to models without information on clinical labels, this outcome-guided factor analysis is able to explicitly identify pathways predictive of the clinical outcome of interest and improve the estimation of the underlying factor structure, as we will show through an application to COVID-19 data in Chapter 4."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["1b303ff658ec757a2c1dde4bd5a7e538","87eda9de84448d1f82354d60eee3eb5f"]},{"key":"dc:title","label":"Title","values":["Dynamic factor analysis with dependent Gaussian processes for high-dimensional biomarker trajectories"]}]}],"canonical_facts":{"dc:contributor.advisor":["Tom, Brian","Goudie, Robert"],"dc:contributor.sponsor":["Medical Research Council Studentship"],"dc:creator":["Cai, Jiachen"],"dc:date.issued":["2024-08-27"],"dc:description.abstract":["The increasing availability of high-dimensional, longitudinal measures of biomarkers can facilitate understanding of biological mechanisms, as required for precision medicine. Biological knowledge suggests that it may be best to describe complex diseases at the level of underlying pathways, which may interact with one another. We propose a Bayesian approach that allows for characterising such correlation among different pathways through dependent Gaussian processes (DGP) and mapping the observed high-dimensional gene expression trajectories into unobserved low-dimensional pathway expression trajectories via Bayesian sparse factor analysis in Chapter 2. Our proposal is the first attempt to relax the classical assumption of independent factors for longitudinal data and has demonstrated a superior performance in recovering the shape of pathway expression trajectories, revealing the relationships between biomarkers and pathways, and predicting biomarker expressions (closer point estimates and narrower predictive intervals), as demonstrated through simulations and application to a real-world H3N2 dataset. To fit the model, we propose a Monte Carlo expectation maximization (MCEM) scheme that can be implemented conveniently by combining a standard Markov Chain Monte Carlo sampler and an R package GPFDA, which returns the maximum likelihood estimates of DGP hyperparameters. The modular structure of MCEM makes it generalizable to other complex models involving the DGP model component. Our R package DGP4LCF that implements the proposed approach is available on CRAN. We extend the proposed method in two ways in Chapter 3, so that it is applicable to a wider range of settings. First, to address the overfitting issue of DGP caused by sparsely collected longitudinal data (i.e., number of observed times is few), we introduce a roughness penalty upon DGP to guarantee an appropriate level of smoothness. Second, we develop a stochastic expectation maximization algorithm to further speed up the inference of DGP hyperparameters. These improvements can be naturally incorporated into our initial analysis framework in Chapter 2, and have proven useful in practice, as we will show in the simulation study and application to a real-world COVID-19 data application in Chapter 3. Application to measurements of metabolites in this example discovers a biomarker named taurine that has been receiving an increasing amount of attention clinically, yet its role has been ignored in a previous analysis of the data. We have developed an R package DFA4SIL (dynamic factor analysis for sparse and irregular longitudinal data) that implements the proposed method. We further extend our model by incorporating clinical outcomes into analysis in Chapter 4. Compared to models without information on clinical labels, this outcome-guided factor analysis is able to explicitly identify pathways predictive of the clinical outcome of interest and improve the estimation of the underlying factor structure, as we will show through an application to COVID-19 data in Chapter 4."],"dc:format.checksum.md5":["1b303ff658ec757a2c1dde4bd5a7e538","87eda9de84448d1f82354d60eee3eb5f"],"dc:identifier.doi":["https://doi.org/10.17863/CAM.116381"],"dc:identifier.uri":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/1b7792f7-2691-4e22-b36b-9c6c3ed941f8/download"],"dc:language":["eng"],"dc:publisher.institution":["University of Cambridge"],"dc:relation.isreferencedby.uri":["https://www.repository.cam.ac.uk/handle/1810/380985"],"dc:rights":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/4ffe66d3-4e6e-4df1-84ee-34d2ebb5a563/download","http://purl.org/NET/rdflicense/allrightsreserved"],"dc:subject":["Dependent Gaussian processes","High-dimensional biomarker trajectories","Monte Carlo expectation maximization","Multivariate longitudinal data","Pathways","Sparse factor analysis","Stochastic expectation maximization"],"dc:title":["Dynamic factor analysis with dependent Gaussian processes for high-dimensional biomarker trajectories"],"dc:type":["Thesis"],"dc:type.qualificationlevel":["Doctoral"],"dc:type.qualificationname":["Doctor of Philosophy (PhD)"]},"updated_at":"2026-07-22T22:24:06Z"}