Back to results

University of Arkansas

Statistical Modeling for High-dimensional Compositional data with Applications to the Human Microbiome

Abstract

dc:description.abstract

<p>Compositional data refer to the data that lie on a simplex, which are common in many scientific domains such as genomics, geology, and economics. As the components in a composition must sum to one, traditional tests based on unconstrained data become inappropriate, and new statistical methods are needed to analyze this special type of data. This dissertation is motivated by some statistical problems arising in the analysis of compositional data. In particular, we focus on the high-dimensional and over-dispersed setting, where the dimensionality of compositions is greater than the sample size and the dispersion parameter is moderate or large. In this dissertation, we consider a general problem of testing for the compositional difference between K populations. We propose a new Bayesian hypothesis, together with a nonparametric and distance-based testing method. Furthermore, we utilize multiple variable-selecting models, including LASSO, elastic net, ridge regression and cumulative logit model, to identify the most important subset of variables. This dissertation is structured as follows: </p> <p>Chapter 1 introduces the compositional microbiome data, and then briefly review different statistical tests and model to be used in our framework, including distance correlation, LASSO, Ridge regression, elastic net, cumulative logit and adjacent-category logit model. </p> <p>Chapter 2 then presents our new statistical test together with two real world applications form human microbiome study. We first formulate a hypothesis from the Bayesian point of view and suggest a nonparametric test based on inter-point distance to evaluate statistical significance. Unlike most existing tests for compositional data, the distance-based method is more sensitive to the compositional difference than the mean-based method, especially when the data are over-dispersed or zero-inflated. It does not rely on any data transformation, sparsity assumption or regularity conditions on the covariance matrix, but directly analyzes the compositions. The performance of this method is evaluated using simulation studies. We apply this new procedure to two human microbiome datasets including a throat microbiome dataset and an intestinal microbiome data. </p> <p>In addition to the overall testing, we also want to identify a small subset of variables that distinguish different populations. Chapter 3 introduces the procedure to select most significant variables (bacteria or genus) using LASSO, Ridge regression, elastic net, cumulative logit model and adjacent-category logit models. Chapter 4 validates our findings from Chapter 3 and presents visualizations using multi-dimensional scaling (MDS). </p> <p>Chapter 5 discusses and concludes the dissertation with some future perspectives. </p>

Degree

thesis:*
Name thesis:degree_name
Doctor of Philosophy in Mathematics (PhD)
Level thesis:degree_level
Dissertation
Year dc:date.available
2021

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Dao, Thy
Advisor dc:contributor.advisor
  • Zhang, Qingyang
Contributors dc:contributor
  • Kaman, Tulin
  • Lee-Bartlett, Jung Ae

Subjects

dc:subject × 8

Identifiers

dc:identifier.*
Repository record dc:identifier
https://scholarworks.uark.edu/etd/4137
OAI identifier oai:identifier
oai:scholarworks.uark.edu:etd-5687

Chain of custody

source
Harvested from
University of Arkansas
Base URL
scholarworks.uark.edu/do/oai/
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Dao, Thy. Statistical Modeling for High-dimensional Compositional data with Applications to the Human Microbiome. Dissertation thesis, 2021. https://scholarworks.uark.edu/etd/4137