Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 20 of 112 for “"Speech processing"”.
-
On speech processing
… important ways of communication between humans. Speech processing refers to the technology of speech signal transformations for more efficient storage and transmission, for enhanced intelligibility and ease of assimilation. According to the different purpose as above, technologies on speech …
-
Self-Supervised Learning for Speech Processing
… learning algorithms on large amounts of labeled speech data have achieved remarkable performance on various spoken language processing applications, often being the state of the arts on the corresponding leaderboards. However, the fact that training these systems relies on large amounts of …
-
Eigenstructure based speech processing in noise
Thesis (M.Eng.)--Massachusetts Institute of Technology, Dept. of Electrical Engineering and Computer Science, 1998.
-
Neural Enhancement Strategies for Robust Speech Processing
In real-world scenarios, speech signals are often contaminated with environmental noises, and reverberation, which degrades speech quality and intelligibility. Lately, the development of deep learning algorithms has marked milestones in speech- based research fields e.g. speech recognition, spoken …
-
Application of generative models in speech processing tasks
Generative probabilistic and neural models of the speech signal are shown to be effective in speech synthesis and speech enhancement, where generating natural and clean speech is the goal. This thesis develops two probabilistic signal processing algorithms based on the source-filter model of speech …
-
Attention-Based Encoder-Decoder Models for Speech Processing
Speech processing is one of the key components of machine perception. It covers a wide range of topics and plays an important role in many real-world applications. Many speech processing problems are modelled using sequence-to-sequence models. More recently, the Attention-Based Encoder-Decoder …
-
Adversarial Attacks on Natural Language and Speech Processing Models
… in computer vision, natural language processing, and speech processing. Despite their high performance, deep learning models are vulnerable to adversarial attacks. A deliberate and specific perturbation of a clean input sample can create an adversarial example, which, when processed by …
-
The benefits of acoustic perceptual information for speech processing systems
… frame-synchronized framework has dominated many speech processing systems, such as ASR and AED targeting human speech activities. These systems have little consideration for the science behind speech and treat the task as a simple statistical classification. The framework also assumes each …
-
Statistical Model Based Multi-Microphone Speech Processing: Toward Overcoming Mismatch Problem
In this thesis, a joint optimal method for clean speech estimation and ASR in a mismatched condition will be described with a unified speech model under a generalized expectation maximization (GEM) scheme. From this perspective, multi-microphone optimal speech estimation can be interpreted as …
-
Unsupervised speech processing with applications to query-by-example spoken term detection
… searching and extracting useful information from speech data in a completely unsupervised setting. In many real world speech processing problems, obtaining annotated data is not cost and time effective. We therefore ask how much can we learn from speech data without any transcription. To address …
-
Speech processing with less supervision : learning from weak labels and multiple modalities
… learning has achieved great success in speech processing with powerful neural network models and vast quantities of in-domain labeled data. However, collecting a labeled dataset covering all domains can be either expensive due to the diversity of speech or almost impossible for some …
-
Talker-specific adaptation: how listeners learn and use indexical information during speech processing
Humans’ ability to understand speech is remarkable in that, despite large amounts of inter-talker variability due to factors such as pitch, speech rate, and accents, we are usually able to understand what is being said quickly and with little conscious effort. However, there is still much to be …
-
Application of signal estimation from the modified short-time Fourier transform to speech processing
Thesis (M.S.)--Massachusetts Institute of Technology, Dept. of Electrical Engineering and Computer Science, 1985.
-
Representation of speech in the primary auditory cortex and its implications for robust speech processing
Speech has evolved as a primary form of communication between humans. This most used means of communication has been the subject of intense study for years, but there is still a lot that we do not know about it. It is an oft repeated fact, that even the performance of the best speech processing …
-
Exploring the neural mechanisms underlying the speech processing and phonological deficits that characterise individuals with dyslexia
… a broad spectrum of cognitive and sensory processing abilities. Analyses focused on the following neural measures: neural phase entrainment (delta, theta, beta, low gamma), cerebro-acoustic coherence (delta and theta), band power (delta, theta, beta, low gamma), event-related potentials …
-
Taking attention away from the auditory modality : investigations of the effect on speech processing using machine learning
Real-world speech processing often takes place in complex multisensory environments. Listeners may need to prioritize sensory inputs from modalities other than audition. Selective attention is thought to be critical in selecting the sensory modality most relevant to the task at hand. Two critical …
-
Cardiovascular Reactivity to Speech Processing and Cold Pressor Stress: Evidence for Sex Differences in Dynamic Functional Cerebral Laterality
… were also able to identify significantly more speech sounds presented to the left ear than men, and they were able to dynamically increase accuracy at the targeted ear identified within each focus group (left or right). Speech sounds processing (dichotic listening task) significantly decreased …
Page 1 of 6