Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 12 of 12 for “"High-dimensional data analysis"”.
-
Scalable sparsity structure learning using Bayesian methods
Learning sparsity pattern in high dimension is a great challenge in both implementation and theory. In this thesis we develop scalable Bayesian algorithms based on EM algorithm and variational inference to learn sparsity structure in various models. Estimation consistency and selection consistency …
-
Nonparametric variable selection and dimension reduction methods and their applications in pharmacogenomics
… it is common to collect large volumes of data in many fields with an extensive amount of variables, but often a small or moderate number of samples. For example, in the analysis of genomic data, the number of genes can be very large, varying from tens of thousands to several millions, …
-
Robust and Constrained Dimension Reduction
"The well-known ""curse of dimensionality"" makes high-dimensional data analysis unusually challenging. Dimension reduction plays a valuable role in enabling certain statistical analyses performed in a parsimonious way. The canonical correlation (CANCOR) method developed by Fung et al. (2002) is …
-
Cluster Analysis in High Dimensions: Robustness, Privacy, and Beyond
Cluster analysis focuses on understanding the cluster structure of data, and is perhaps one of the most important subfields in high-dimensional data analysis. Traditionally, cluster analysis focuses on partitioning data into closely related groups, such as in k-means clustering and learning mixture …
-
Tackling Computational Challenges in High-Throughput RNA Interference Screening
… while reducing cost and decreasing time. High-throughput RNAi screening (HTS) has been widely accepted and used in a variety of biomedical and biological research projects as the first step to identifying novel drug targets or pathway components. Huge data sets are being generated, but …
-
Tree-based Methods for Learning Probability Distributions
… task in statistics but challenging if a data distribution of our interest is complicated and high-dimensional. Addressing this challenging problem is the main topic of this thesis, and mainly discussed herein are two types of new tree-based methods: a single-tree method and an ensemble …
-
Model-Free Variable Screening, Sparse Regression Analysis and Other Applications with Optimal Transformations
… methods play important roles in modeling high dimensional data. Variable screening is the process of filtering out irrelevant variables, with the aim to reduce the dimensionality from ultrahigh to high while retaining all important variables. Variable selection is the process of selecting …
-
Applications of low-rank matrix recovery methods in computer vision
The ubiquitous availability of high-dimensional data such as images and videos has generated a lot of interest in high-dimensional data analysis. One of the key issues that needs to be addressed in real applications is the presence of large-magnitude non-Gaussian errors. For image data, the problem …
-
Bayesian and Information-Theoretic Learning of High Dimensional Data
… of sparseness is harnessed to learn a low dimensional representation of high dimensional data. This sparseness assumption is exploited in multiple ways. In the Bayesian Elastic Net, a small number of correlated features are identified for the response variable. In the sparse Factor Analysis …
-
Cyber-physical data distillation in a sensor rich world
… will be dominated by vast amount of sensing data. This thesis attacks a grand challenge in this vision: how to extract human-consumable information from such data? The prime contribution of this work is a framework called FusionSuite that facilitates the development of future applications and …
-
Algorithmic advances in learning from large dimensional matrices and scientific data
… a range of questions in machine learning and data analysis related to large dimensional matrices and scientific data. Two key research objectives connect the different parts of the thesis: (a) development of fast, efficient, and scalable algorithms for machine learning which handle large …
-
Low-rank estimation and embedding learning: theory and applications
In many real-world applications of data mining, datasets can be represented using matrices, where rows of the matrix correspond to objects (or data instances) and columns to features (or attributes). Often the datasets are in high-dimensional feature space. For example, in the vector space model of …