Florida State University
Mathematical Study on the Expressive Power of Machine Learning and Applications in Optimal Filtering Problems
Abstract
dc:descriptionThis dissertation investigates the practical expressive power of machine learning models as general function approximators and their practical applications in optimal filtering problems. The practical expressive power of these models depends on the choice of \textit{model structure} (deep or shallow), which sets the upper limit of expressive power, and the \textit{trainability} of the model, which enables the realization of the expressive power in practical applications, such as regression and classification. Our primary objective is to enhance the practical expressive power and overall performance of neural networks (NNs) by improving their trainability. In addition to traditional machine learning tasks, we specifically focus on the optimal filtering problem, which holds significant importance in science and engineering. When addressing the representation of filtering density in optimal filtering problems, we adopt a similar rationale for practical expressive power by considering the \textit{model structure}, which determines the theoretical capacity for storing probability distributions, and the \textit{trainability}, which determines the optimization cost for utilizing the expressive power of the model to approximate the filtering density. By carefully balancing these two factors, our aim is to enhance the overall performance of optimal filtering problems. The first part of the dissertation focuses on enhancing the practical expressive power of neural networks by improving their trainability. To achieve this, we first introduce the concept of the general activation transform (GAT) as a generalization of the traditional elementwise activation function in the function space. The GAT can be derived from the general deep neural networks (GDNNs) framework, in which the affine transformations in traditional deep neural networks (DNNs) are replaced with integral transforms, with a finite-rank parameterization of the weight integral kernels. The GAT maps the input vector into the function space by treating the input vector as coefficients of the input basis function, performing activation on the state function, and generating the output through integration with the output basis function. We specifically focus on GAT-ReLU, a variant of GAT with ReLU activation, which exhibits enriched activation patterns and smoother solutions compared to traditional ReLU activation. By selecting proper basis functions (continuous and high-variation), GAT-ReLU enhances the practical expressive power of the model, which is demonstrated by many numerical experiments. Apart from GAT, we introduce transferable neural networks (TransNet) for solving partial differential equations (PDEs) by reducing the optimization difficulty through the idea of transfer learning. The construction of transferable neural feature spaces involves re-parameterization of hidden neurons and auxiliary functions. Theoretical analysis validates the high quality of the resulting feature space, characterized by uniformly distributed neurons. Extensive numerical experiments demonstrate the effectiveness and improved transferability of our proposed method across different PDEs. The second part of this dissertation concentrates on addressing the optimal filtering problem by different choices of representations for filtering density. This involves considering the model's theoretical capacity for representing probability distributions and the associated optimization cost for approximating the distributions. By carefully balancing these factors, we aim to develop effective approaches for solving the optimal filtering problem. We first introduce an adaptive kernel method that utilizes a Gaussian mixture model (GMM) to represent the probability density function (PDF) within the Bayesian framework. The state dynamics are formulated using the Fokker-Planck equation, and the PDE operator is decomposed into linear and nonlinear parts. The linear part is analytically incorporated into the GMM model, while the nonlinear part is approximated using an adaptive kernel algorithm. Bayesian inference is applied to each Gaussian component in the GMM to incorporate observational data. We also provide many numerical examples to validate the performance of the GMM representation in solving nonlinear filtering problem. In addition, We also explore the use of score-based models for solving optimal filtering problems, where the score function (gradient of the logarithm of the PDF) plays a central role in representing the filtering density. We train score-based generative models to approximate samples from the prior distribution, and by combining the learned score functions of the prior and likelihood, we can directly generate samples from the posterior distribution. Score-based models leverage the high capacity of deep neural networks (DNNs) to effectively capture non-Gaussian and high-dimensional distributions. Our numerical experiments demonstrate the effectiveness of this approach, showcasing its ability to handle high-dimensional (up to 100 dimensions) and highly nonlinear filtering problems effectively.
Degree
thesis:*- Grantor dc:publisher
- Florida State University
- Year dc:date
- 2023
Author and committee
dc:creator, dc:contributor.*- Contributors dc:contributor
-
- Zhang, Zezhong (author)
- Bao, Feng (professor directing dissertation)
- Liu, Xiuwen, 1966- (university representative)
- Farhat, Aseel (committee member)
- Ökten, Giray (committee member)
- Zhu, Lingjiong (committee member)
- Florida State University (degree granting institution)
- College of Arts and Sciences (degree granting college)
- Department of Mathematics (degree granting department)
Subjects
dc:subject × 1Rights
- Language dc:language
- English
Identifiers
dc:identifier.*- Identifier
-
fsu:927814
iid: Zhang_fsu_0071E_18168 - OAI identifier oai:identifier
- oai:diginole.lib.fsu.edu:fsu_927814