{"id":{"repo_id":"cambridge","oai_identifier":"oai:www.repository.cam.ac.uk:1810/379741"},"canonical_url":"https://search.dev.ndltd.org/etd/cambridge/oai:www.repository.cam.ac.uk:1810/379741","repository":{"repo_id":"cambridge","name":"Cambridge University","base_url":"https://api.repository.cam.ac.uk/server/oai/request"},"display":{"title":"Modelling multiple disease risk factors for microsimulation studies","abstract":"Longitudinal microsimulation is a common technique used in studies that aim to evaluate and compare health policies to reduce the risk of chronic diseases. This requires simulating realistic trajectories of multiple risk factors for a synthetic population. A variety of methods have been used to model the progression through time of multiple disease risk factors, such as blood pressure and smoking status, in microsimulations, based on a variety of kinds of data. However, the relative merits of different methods have rarely been discussed, the common statistical principles underlying them haven’t been described, and the principles that might help us decide between different methods aren't clear. This thesis gives a review of the methods that have been used for the purpose of generating synthetic longitudinal data sets, in contexts where different forms of data are available (either longitudinal or cross-sectional data), and the different assumptions they make. We also describe cross-validation methods and graphical diagnostics that can be used to compare diverse methods in practice. We illustrate the methods, and how they are compared, in the context of the cardiovascular risk factor data used in a microsimulation model for mid-life health checks. We first describe both parametric and non-parametric methods for simulating a cross-sectional dataset. These methods may be used to simulate a baseline population for microsimulation, and as the basis of the more complex methods that are required to simulate longitudinal data. We describe a number of different parametric methods, which are distinguished by how they decompose the multivariate distribution of the risk factors. The non-parametric methods are based on stratified sampling, and we develop a novel method for defining strata based on regression trees. We then describe how we can build on those methods to simulate longitudinal data on multiple risk factors, firstly in a situation when longitudinal data is available. In this situation we have observations at multiple time points on the same individuals, meaning there is direct information about how an individual's risk factors change over time. We described how both Markov models and random effects models can be fitted to longitudinal data in order to simulate synthetic longitudinal data. We also describe methods that can be used to simulate longitudinal data when only serial cross-sectional data is available. While cross-sectional data can describe how the risk distribution of a population will change with age, the extra challenge here is to simulate how each synthetic individual's trajectory of risk factors is expected to differ from that of other individuals. These methods all employ in some way the assumption of ``rank stability'', that is, the assumption that an individual's rank or quantile for a given risk factor, compared to other individuals stays constant over time. Our case study research demonstrated there are a number of models which produce realistic synthetic longitudinal data based on either longitudinal or cross-sectional data. These models are flexible, easy to implement and based on plausible assumptions, and we have demonstrated practical tools for model selection.","abstract_html":"Longitudinal microsimulation is a common technique used in studies that aim to evaluate and compare health policies to reduce the risk of chronic diseases. This requires simulating realistic trajectories of multiple risk factors for a synthetic population. A variety of methods have been used to model the progression through time of multiple disease risk factors, such as blood pressure and smoking status, in microsimulations, based on a variety of kinds of data. However, the relative merits of different methods have rarely been discussed, the common statistical principles underlying them haven’t been described, and the principles that might help us decide between different methods aren&#x27;t clear. This thesis gives a review of the methods that have been used for the purpose of generating synthetic longitudinal data sets, in contexts where different forms of data are available (either longitudinal or cross-sectional data), and the different assumptions they make. We also describe cross-validation methods and graphical diagnostics that can be used to compare diverse methods in practice. We illustrate the methods, and how they are compared, in the context of the cardiovascular risk factor data used in a microsimulation model for mid-life health checks. We first describe both parametric and non-parametric methods for simulating a cross-sectional dataset. These methods may be used to simulate a baseline population for microsimulation, and as the basis of the more complex methods that are required to simulate longitudinal data. We describe a number of different parametric methods, which are distinguished by how they decompose the multivariate distribution of the risk factors. The non-parametric methods are based on stratified sampling, and we develop a novel method for defining strata based on regression trees. We then describe how we can build on those methods to simulate longitudinal data on multiple risk factors, firstly in a situation when longitudinal data is available. In this situation we have observations at multiple time points on the same individuals, meaning there is direct information about how an individual&#x27;s risk factors change over time. We described how both Markov models and random effects models can be fitted to longitudinal data in order to simulate synthetic longitudinal data. We also describe methods that can be used to simulate longitudinal data when only serial cross-sectional data is available. While cross-sectional data can describe how the risk distribution of a population will change with age, the extra challenge here is to simulate how each synthetic individual&#x27;s trajectory of risk factors is expected to differ from that of other individuals. These methods all employ in some way the assumption of ``rank stability&#x27;&#x27;, that is, the assumption that an individual&#x27;s rank or quantile for a given risk factor, compared to other individuals stays constant over time. Our case study research demonstrated there are a number of models which produce realistic synthetic longitudinal data based on either longitudinal or cross-sectional data. These models are flexible, easy to implement and based on plausible assumptions, and we have demonstrated practical tools for model selection.","abstract_has_math":false,"creators":["Church, Oliver"],"institution":"University of Cambridge","degree_name":"Doctor of Philosophy (PhD)","degree_level":"Doctoral","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Jackson, Christopher"],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-09-05","date_published":"2024-09-05","updated_at":"2026-07-22T22:24:21Z","subjects":["cardiovascular","chronic diseases","longitudinal","microsimulation","multivariate","risk factors"],"languages":["eng"],"rights":[],"rights_urls":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/3c109561-e3aa-4976-af71-57c383d76719/download","http://purl.org/NET/rdflicense/allrightsreserved"],"identifier_entries":[]},"links":{"outbound_url":"https://doi.org/10.17863/CAM.115710","outbound_label":"DOI","outbound_source":"dc:identifier.doi"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Jackson, Christopher"]},{"key":"dc:creator","label":"Author","values":["Church, Oliver"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2024-09-05"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["University of Cambridge"]},{"key":"dc:relation.isreferencedby.uri","label":"Dc Relation Isreferencedby URI","values":["https://www.repository.cam.ac.uk/handle/1810/379741"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["Doctoral"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["Doctor of Philosophy (PhD)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["cardiovascular","chronic diseases","longitudinal","microsimulation","multivariate","risk factors"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/3c109561-e3aa-4976-af71-57c383d76719/download","http://purl.org/NET/rdflicense/allrightsreserved"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.17863/CAM.115710"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/825552c9-fadf-4166-a3e3-a7c94e4d9699/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Longitudinal microsimulation is a common technique used in studies that aim to evaluate and compare health policies to reduce the risk of chronic diseases. This requires simulating realistic trajectories of multiple risk factors for a synthetic population. A variety of methods have been used to model the progression through time of multiple disease risk factors, such as blood pressure and smoking status, in microsimulations, based on a variety of kinds of data. However, the relative merits of different methods have rarely been discussed, the common statistical principles underlying them haven’t been described, and the principles that might help us decide between different methods aren't clear. This thesis gives a review of the methods that have been used for the purpose of generating synthetic longitudinal data sets, in contexts where different forms of data are available (either longitudinal or cross-sectional data), and the different assumptions they make. We also describe cross-validation methods and graphical diagnostics that can be used to compare diverse methods in practice. We illustrate the methods, and how they are compared, in the context of the cardiovascular risk factor data used in a microsimulation model for mid-life health checks. We first describe both parametric and non-parametric methods for simulating a cross-sectional dataset. These methods may be used to simulate a baseline population for microsimulation, and as the basis of the more complex methods that are required to simulate longitudinal data. We describe a number of different parametric methods, which are distinguished by how they decompose the multivariate distribution of the risk factors. The non-parametric methods are based on stratified sampling, and we develop a novel method for defining strata based on regression trees. We then describe how we can build on those methods to simulate longitudinal data on multiple risk factors, firstly in a situation when longitudinal data is available. In this situation we have observations at multiple time points on the same individuals, meaning there is direct information about how an individual's risk factors change over time. We described how both Markov models and random effects models can be fitted to longitudinal data in order to simulate synthetic longitudinal data. We also describe methods that can be used to simulate longitudinal data when only serial cross-sectional data is available. While cross-sectional data can describe how the risk distribution of a population will change with age, the extra challenge here is to simulate how each synthetic individual's trajectory of risk factors is expected to differ from that of other individuals. These methods all employ in some way the assumption of ``rank stability'', that is, the assumption that an individual's rank or quantile for a given risk factor, compared to other individuals stays constant over time. Our case study research demonstrated there are a number of models which produce realistic synthetic longitudinal data based on either longitudinal or cross-sectional data. These models are flexible, easy to implement and based on plausible assumptions, and we have demonstrated practical tools for model selection."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["d90a95e81f7064819f95ba73b116f7cf","87eda9de84448d1f82354d60eee3eb5f"]},{"key":"dc:title","label":"Title","values":["Modelling multiple disease risk factors for microsimulation studies"]}]}],"canonical_facts":{"dc:contributor.advisor":["Jackson, Christopher"],"dc:creator":["Church, Oliver"],"dc:date.issued":["2024-09-05"],"dc:description.abstract":["Longitudinal microsimulation is a common technique used in studies that aim to evaluate and compare health policies to reduce the risk of chronic diseases. This requires simulating realistic trajectories of multiple risk factors for a synthetic population. A variety of methods have been used to model the progression through time of multiple disease risk factors, such as blood pressure and smoking status, in microsimulations, based on a variety of kinds of data. However, the relative merits of different methods have rarely been discussed, the common statistical principles underlying them haven’t been described, and the principles that might help us decide between different methods aren't clear. This thesis gives a review of the methods that have been used for the purpose of generating synthetic longitudinal data sets, in contexts where different forms of data are available (either longitudinal or cross-sectional data), and the different assumptions they make. We also describe cross-validation methods and graphical diagnostics that can be used to compare diverse methods in practice. We illustrate the methods, and how they are compared, in the context of the cardiovascular risk factor data used in a microsimulation model for mid-life health checks. We first describe both parametric and non-parametric methods for simulating a cross-sectional dataset. These methods may be used to simulate a baseline population for microsimulation, and as the basis of the more complex methods that are required to simulate longitudinal data. We describe a number of different parametric methods, which are distinguished by how they decompose the multivariate distribution of the risk factors. The non-parametric methods are based on stratified sampling, and we develop a novel method for defining strata based on regression trees. We then describe how we can build on those methods to simulate longitudinal data on multiple risk factors, firstly in a situation when longitudinal data is available. In this situation we have observations at multiple time points on the same individuals, meaning there is direct information about how an individual's risk factors change over time. We described how both Markov models and random effects models can be fitted to longitudinal data in order to simulate synthetic longitudinal data. We also describe methods that can be used to simulate longitudinal data when only serial cross-sectional data is available. While cross-sectional data can describe how the risk distribution of a population will change with age, the extra challenge here is to simulate how each synthetic individual's trajectory of risk factors is expected to differ from that of other individuals. These methods all employ in some way the assumption of ``rank stability'', that is, the assumption that an individual's rank or quantile for a given risk factor, compared to other individuals stays constant over time. Our case study research demonstrated there are a number of models which produce realistic synthetic longitudinal data based on either longitudinal or cross-sectional data. These models are flexible, easy to implement and based on plausible assumptions, and we have demonstrated practical tools for model selection."],"dc:format.checksum.md5":["d90a95e81f7064819f95ba73b116f7cf","87eda9de84448d1f82354d60eee3eb5f"],"dc:identifier.doi":["https://doi.org/10.17863/CAM.115710"],"dc:identifier.uri":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/825552c9-fadf-4166-a3e3-a7c94e4d9699/download"],"dc:language":["eng"],"dc:publisher.institution":["University of Cambridge"],"dc:relation.isreferencedby.uri":["https://www.repository.cam.ac.uk/handle/1810/379741"],"dc:rights":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/3c109561-e3aa-4976-af71-57c383d76719/download","http://purl.org/NET/rdflicense/allrightsreserved"],"dc:subject":["cardiovascular","chronic diseases","longitudinal","microsimulation","multivariate","risk factors"],"dc:title":["Modelling multiple disease risk factors for microsimulation studies"],"dc:type":["Thesis"],"dc:type.qualificationlevel":["Doctoral"],"dc:type.qualificationname":["Doctor of Philosophy (PhD)"]},"updated_at":"2026-07-22T22:24:21Z"}