{"id":{"repo_id":"venda","oai_identifier":"oai:univendspace.univen.ac.za:11602/1552"},"canonical_url":"https://search.dev.ndltd.org/etd/venda/oai:univendspace.univen.ac.za:11602/1552","repository":{"repo_id":"venda","name":"University of Venda","base_url":"https://univendspace.univen.ac.za/server/oai/request"},"display":{"title":"Variable selection in discrete survival models","abstract":"Selection of variables is vital in high dimensional statistical modelling as it aims to identify the right subset model. However, variable selection for discrete survival analysis poses many challenges due to a complicated data structure. Survival data might have unobserved heterogeneity leading to biased estimates when not taken into account. Conventional variable selection methods have stability problems. A simulation approach was used to assess and compare the performance of Least Absolute Shrinkage and Selection Operator (Lasso) and gradient boosting on discrete survival data. Parameter related mean squared errors (MSEs) and false positive rates suggest Lasso performs better than gradient boosting. Frailty models outperform discrete survival models that do not account for unobserved heterogeneity. The two methods were also applied on Zimbabwe Demographic Health Survey (ZDHS) 2016 data on age at first marriage and did not select exactly the same variables. Gradient boosting retained more variables into the model. Place of residence, highest educational level attained and age cohort are the major influential factors of age at first marriage in Zimbabwe based on Lasso.","abstract_html":"Selection of variables is vital in high dimensional statistical modelling as it aims to identify the right subset model. However, variable selection for discrete survival analysis poses many challenges due to a complicated data structure. Survival data might have unobserved heterogeneity leading to biased estimates when not taken into account. Conventional variable selection methods have stability problems. A simulation approach was used to assess and compare the performance of Least Absolute Shrinkage and Selection Operator (Lasso) and gradient boosting on discrete survival data. Parameter related mean squared errors (MSEs) and false positive rates suggest Lasso performs better than gradient boosting. Frailty models outperform discrete survival models that do not account for unobserved heterogeneity. The two methods were also applied on Zimbabwe Demographic Health Survey (ZDHS) 2016 data on age at first marriage and did not select exactly the same variables. Gradient boosting retained more variables into the model. Place of residence, highest educational level attained and age cohort are the major influential factors of age at first marriage in Zimbabwe based on Lasso.","abstract_has_math":false,"creators":["Mabvuu, Coster"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Bere, A.","Sigauke, C."],"committee_chairs":[],"committee_members":[],"year":2020,"date_issued":"2020-02-27","date_published":"2020-02-27","updated_at":"2026-07-27T21:57:55Z","subjects":["Boosting","Discrete-time hazard model","Lasso","Penalised variable selection methods","Unobservrd heterogeneity"],"languages":["en"],"rights":["University of Venda"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/11602/1552","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Bere, A.","Sigauke, C."]},{"key":"dc:creator","label":"Author","values":["Mabvuu, Coster"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2020"]},{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2020-09-29T19:33:45Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2020-09-29T19:33:45Z"]},{"key":"dc:date.issued","label":"Date","values":["2020-02-27"]},{"key":"dc:type","label":"Dc Type","values":["Dissertation"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Boosting","Discrete-time hazard model","Lasso","Penalised variable selection methods","Unobservrd heterogeneity"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["University of Venda"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["http://hdl.handle.net/11602/1552"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["MSc (Statistics)","Department of Statistics"]},{"key":"dc:description.abstract","label":"Abstract","values":["Selection of variables is vital in high dimensional statistical modelling as it aims to identify the right subset model. However, variable selection for discrete survival analysis poses many challenges due to a complicated data structure. Survival data might have unobserved heterogeneity leading to biased estimates when not taken into account. Conventional variable selection methods have stability problems. A simulation approach was used to assess and compare the performance of Least Absolute Shrinkage and Selection Operator (Lasso) and gradient boosting on discrete survival data. Parameter related mean squared errors (MSEs) and false positive rates suggest Lasso performs better than gradient boosting. Frailty models outperform discrete survival models that do not account for unobserved heterogeneity. The two methods were also applied on Zimbabwe Demographic Health Survey (ZDHS) 2016 data on age at first marriage and did not select exactly the same variables. Gradient boosting retained more variables into the model. Place of residence, highest educational level attained and age cohort are the major influential factors of age at first marriage in Zimbabwe based on Lasso."]},{"key":"dc:title","label":"Title","values":["Variable selection in discrete survival models"]}]}],"canonical_facts":{"dc:contributor.advisor":["Bere, A.","Sigauke, C."],"dc:creator":["Mabvuu, Coster"],"dc:date":["2020"],"dc:date.accessioned":["2020-09-29T19:33:45Z"],"dc:date.available":["2020-09-29T19:33:45Z"],"dc:date.issued":["2020-02-27"],"dc:description":["MSc (Statistics)","Department of Statistics"],"dc:description.abstract":["Selection of variables is vital in high dimensional statistical modelling as it aims to identify the right subset model. However, variable selection for discrete survival analysis poses many challenges due to a complicated data structure. Survival data might have unobserved heterogeneity leading to biased estimates when not taken into account. Conventional variable selection methods have stability problems. A simulation approach was used to assess and compare the performance of Least Absolute Shrinkage and Selection Operator (Lasso) and gradient boosting on discrete survival data. Parameter related mean squared errors (MSEs) and false positive rates suggest Lasso performs better than gradient boosting. Frailty models outperform discrete survival models that do not account for unobserved heterogeneity. The two methods were also applied on Zimbabwe Demographic Health Survey (ZDHS) 2016 data on age at first marriage and did not select exactly the same variables. Gradient boosting retained more variables into the model. Place of residence, highest educational level attained and age cohort are the major influential factors of age at first marriage in Zimbabwe based on Lasso."],"dc:identifier.uri":["http://hdl.handle.net/11602/1552"],"dc:language.iso":["en"],"dc:rights":["University of Venda"],"dc:subject":["Boosting","Discrete-time hazard model","Lasso","Penalised variable selection methods","Unobservrd heterogeneity"],"dc:title":["Variable selection in discrete survival models"],"dc:type":["Dissertation"]},"updated_at":"2026-07-27T21:57:55Z"}