National University of Singapore
INVERSE APPROXIMATION THEORY OF RECURRENT MODELS FOR LEARNING SEQUENCES
Abstract
dc:description.abstractLearning long-term relationships is a challenging task in sequence modelling. Despite numerous empirical results demonstrating the difficulty of recurrent models in learning long-term relationships, this dissertation presents a series of theoretical studies on the learning of long-term memories using recurrent models. First, based on the concept of a generalized memory function over nonlinear functional sequences, we prove that adding nonlinear activation does not change the asymptotic exponential memory decay pattern. Then we prove the universal approximation property for state-space models. A similar inverse approximation result is established for state-space models, indicating that despite their high efficiency, changing the method of incorporating nonlinear activation fails to alleviate the memory decay constraint observed in the first part. Based on the proofs, we identify that suitable reparameterizations are the key to stably approximation. A class of stable reparameterizations enables state-space models to achieve stable approximation for any target with decaying memory.
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- WANG SHIDA