{"id":{"repo_id":"helsinki","oai_identifier":"oai:helda.helsinki.fi:10138/598194"},"canonical_url":"https://search.dev.ndltd.org/etd/helsinki/oai:helda.helsinki.fi:10138/598194","repository":{"repo_id":"helsinki","name":"University of Helsinki","base_url":"https://helda.helsinki.fi/server/oai/request"},"display":{"title":"Bayesian approach for modeling zero-inflated plant percent covers using spatial left-censored beta regression","abstract":"Species distribution models are tools to combine species observations to environmental data. These models learn about the species' ecological preferences and can be used to predict their distributions in unsampled locations or under climate change scenarios. This information is essential to learn about species' responses to climate change and further to aid conservation decisions. Typical challenge with ecological data is the large number of zeros: most of the time species are not observed at the sampling location. Zeros can be thought to arise from ecological reasons or due to randomness. This problem of zero inflation is well studied for discrete positive data, such as counts and occurrences. When sampling plant species, it is logical to record their percent cover, leading to observations between zero and one. With improving technologies, this kind of data is becoming increasingly popular while the need to model vegetation communities is high. However, zero-inflated models for continuous positive data have not been widely studied. One way to build models for percent cover data is to use beta distribution, utilizing left-censoring to create probability mass for zero observations. In this study, alternative versions of left-censored beta regression are implemented and their behavior studied when applied to real data from the Baltic Sea. Specifically, the model is extended by introducing spatial random effects and a separate process for zeros arising from a sites unsuitability for a species. This study finds that including spatial random effects improves the predictive performance of the model, whereas accounting for two sources of zeros does not make a big difference. It is also found that modeling the precision of the beta distribution with covariates is highly beneficial for predictive ability and essential to make the model catch the patterns in the observed data. This result is considered important, since the existing studies typically assume a constant precision parameter across sampling locations. On top of the predictive performance, this study shows that the model specification can have a large effect on the identified hotspot areas. This could further affect the conservation actions informed by these models. The result serves as a reminder for proper model comparison and model assessment when doing Bayesian analysis.","abstract_html":"Species distribution models are tools to combine species observations to environmental data. These models learn about the species&#x27; ecological preferences and can be used to predict their distributions in unsampled locations or under climate change scenarios. This information is essential to learn about species&#x27; responses to climate change and further to aid conservation decisions. Typical challenge with ecological data is the large number of zeros: most of the time species are not observed at the sampling location. Zeros can be thought to arise from ecological reasons or due to randomness. This problem of zero inflation is well studied for discrete positive data, such as counts and occurrences. When sampling plant species, it is logical to record their percent cover, leading to observations between zero and one. With improving technologies, this kind of data is becoming increasingly popular while the need to model vegetation communities is high. However, zero-inflated models for continuous positive data have not been widely studied. One way to build models for percent cover data is to use beta distribution, utilizing left-censoring to create probability mass for zero observations. In this study, alternative versions of left-censored beta regression are implemented and their behavior studied when applied to real data from the Baltic Sea. Specifically, the model is extended by introducing spatial random effects and a separate process for zeros arising from a sites unsuitability for a species. This study finds that including spatial random effects improves the predictive performance of the model, whereas accounting for two sources of zeros does not make a big difference. It is also found that modeling the precision of the beta distribution with covariates is highly beneficial for predictive ability and essential to make the model catch the patterns in the observed data. This result is considered important, since the existing studies typically assume a constant precision parameter across sampling locations. On top of the predictive performance, this study shows that the model specification can have a large effect on the identified hotspot areas. This could further affect the conservation actions informed by these models. The result serves as a reminder for proper model comparison and model assessment when doing Bayesian analysis.","abstract_has_math":false,"creators":["Pietilä, Juho"],"institution":"Helsingin yliopisto","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-06-26","date_published":"2025-06-26","updated_at":"2026-07-27T19:55:56Z","subjects":["species distribution modeling","percent cover data","beta regression","Bayesian hierarchical models","spatial random effects","zero inflation","left-censoring"],"languages":["eng"],"rights":["CC BY 4.0"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/10138/598194","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["Pietilä, Juho"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2025-06-26T07:21:19Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2025-06-26T07:21:19Z"]},{"key":"dc:date.issued","label":"Date","values":["2025-06-26"]},{"key":"dc:publisher","label":"Institution","values":["Helsingin yliopisto","University of Helsinki","Helsingfors universitet"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["species distribution modeling","percent cover data","beta regression","Bayesian hierarchical models","spatial random effects","zero inflation","left-censoring"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["CC BY 4.0"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["http://hdl.handle.net/10138/598194"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Species distribution models are tools to combine species observations to environmental data. These models learn about the species' ecological preferences and can be used to predict their distributions in unsampled locations or under climate change scenarios. This information is essential to learn about species' responses to climate change and further to aid conservation decisions. Typical challenge with ecological data is the large number of zeros: most of the time species are not observed at the sampling location. Zeros can be thought to arise from ecological reasons or due to randomness. This problem of zero inflation is well studied for discrete positive data, such as counts and occurrences. When sampling plant species, it is logical to record their percent cover, leading to observations between zero and one. With improving technologies, this kind of data is becoming increasingly popular while the need to model vegetation communities is high. However, zero-inflated models for continuous positive data have not been widely studied. One way to build models for percent cover data is to use beta distribution, utilizing left-censoring to create probability mass for zero observations. In this study, alternative versions of left-censored beta regression are implemented and their behavior studied when applied to real data from the Baltic Sea. Specifically, the model is extended by introducing spatial random effects and a separate process for zeros arising from a sites unsuitability for a species. This study finds that including spatial random effects improves the predictive performance of the model, whereas accounting for two sources of zeros does not make a big difference. It is also found that modeling the precision of the beta distribution with covariates is highly beneficial for predictive ability and essential to make the model catch the patterns in the observed data. This result is considered important, since the existing studies typically assume a constant precision parameter across sampling locations. On top of the predictive performance, this study shows that the model specification can have a large effect on the identified hotspot areas. This could further affect the conservation actions informed by these models. The result serves as a reminder for proper model comparison and model assessment when doing Bayesian analysis."]},{"key":"dc:title","label":"Title","values":["Bayesian approach for modeling zero-inflated plant percent covers using spatial left-censored beta regression"]}]}],"canonical_facts":{"dc:creator":["Pietilä, Juho"],"dc:date.accessioned":["2025-06-26T07:21:19Z"],"dc:date.available":["2025-06-26T07:21:19Z"],"dc:date.issued":["2025-06-26"],"dc:description.abstract":["Species distribution models are tools to combine species observations to environmental data. These models learn about the species' ecological preferences and can be used to predict their distributions in unsampled locations or under climate change scenarios. This information is essential to learn about species' responses to climate change and further to aid conservation decisions. Typical challenge with ecological data is the large number of zeros: most of the time species are not observed at the sampling location. Zeros can be thought to arise from ecological reasons or due to randomness. This problem of zero inflation is well studied for discrete positive data, such as counts and occurrences. When sampling plant species, it is logical to record their percent cover, leading to observations between zero and one. With improving technologies, this kind of data is becoming increasingly popular while the need to model vegetation communities is high. However, zero-inflated models for continuous positive data have not been widely studied. One way to build models for percent cover data is to use beta distribution, utilizing left-censoring to create probability mass for zero observations. In this study, alternative versions of left-censored beta regression are implemented and their behavior studied when applied to real data from the Baltic Sea. Specifically, the model is extended by introducing spatial random effects and a separate process for zeros arising from a sites unsuitability for a species. This study finds that including spatial random effects improves the predictive performance of the model, whereas accounting for two sources of zeros does not make a big difference. It is also found that modeling the precision of the beta distribution with covariates is highly beneficial for predictive ability and essential to make the model catch the patterns in the observed data. This result is considered important, since the existing studies typically assume a constant precision parameter across sampling locations. On top of the predictive performance, this study shows that the model specification can have a large effect on the identified hotspot areas. This could further affect the conservation actions informed by these models. The result serves as a reminder for proper model comparison and model assessment when doing Bayesian analysis."],"dc:identifier.uri":["http://hdl.handle.net/10138/598194"],"dc:language.iso":["eng"],"dc:publisher":["Helsingin yliopisto","University of Helsinki","Helsingfors universitet"],"dc:rights":["CC BY 4.0"],"dc:subject":["species distribution modeling","percent cover data","beta regression","Bayesian hierarchical models","spatial random effects","zero inflation","left-censoring"],"dc:title":["Bayesian approach for modeling zero-inflated plant percent covers using spatial left-censored beta regression"]},"updated_at":"2026-07-27T19:55:56Z"}