Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 20 of 304 for “"Audio-visual"”.
-
Interpreting Electroacoustic Audio-visual Music
… the process of composition for electroacoustic audio-visual music. These are fixed media works in which sound and image materials are accessed, generated, explored and configured in creation of a musically informed audio-visual expression. Within the process of composition, the composer must …
-
Environmental audio-visual context awareness
Thesis (M.Eng.)--Massachusetts Institute of Technology, Dept. of Electrical Engineering and Computer Science, 1999.
-
Towards Integrated Audio-Visual Learning: From Vision-to-Audio Generation to a Unified Audio-Visual Framework
The interplay between audio and visual signals, rich in correlations across various scales, significantly impacts human perception and drives a consistent demand for audio-visual applications across fields such as video production, animation, and virtual reality. Historically, the creation and …
-
Learning digits via joint audio-visual representations
… intermediary text transcriptions in correlating visual and audio inputs from the environment; rather, they directly make connections between what they see and what they hear, sometimes even across languages! In this thesis, we present weakly-supervised models for learning representations of …
-
Audio-visual frameworks for design process representation
… In the first part of this study, the use of audio-visual interfaces as tools for representing the design process is proposed. The idea is to understand, through simulation, what beneficial effects a process based on multiple feedbacks can potentially have on the actual design. As such, five …
-
A Study of Teacher Constructed Audio-Visual Aids
… study is to determine which teacher constructed audio-visual aids, as evaluated by a sample of audio-visual coordinators, are considered to be the most valuable for classroom use and which materials and equipment are needed for the construction and utilization of these aids.
-
Information Fusion for Robust Audio -Visual Speech Recognition
… technique that can efficiently incorporate both visual and acoustic speech signals to achieve better speech recognition accuracy than that achieved by either of the single modalities alone. The proposed fusion schemes have been tested in different situations. The experiment results show …
-
Self-Supervised Audio-Visual Speech Diarization and Recognition
… word recognition accuracy. The proposed Audio-Visual Cocktail (AVC) HuBERT model extends video input dimenions, lengthens feature size, and adds projection layers to split outputs into corresponding speakers. A complementary synthesized dataset is constructed by mixing audio and video …
-
Learning visual models from paired audio-visual examples
… bustle of a busy café, our days are filled with visual experiences that are accompanied by distinctive sounds. In this thesis, we show that these sounds can provide a rich training signal for learning visual models. First, we propose the task of predicting the sound that an object makes when …
-
Multimodal Fusion With Applications to Audio -Visual Speech Recognition
… characteristics of the intermodal couplings in audio-visual speech recognition and in multichannel biometrics defy a universal fusion method for both applications. For audio-visual speech modeling, we propose a novel sensory fusion method based on the coupled hidden Markov models (CHMMs). The …
-
Audio-Visual Scene Analysis With Application in Sports Video
… the thesis presents our work on ""single-channel audio source separation"" using generative probabilistic models. The approach can be potentially very useful as a pre-processing step for the key audio ""objects"" detection."
-
Efficient audio-visual representations for reasoning and synthesis tasks
Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2023-04-12 without embargo terms
-
Robust audio-visual person verification using Web-camera video
This thesis examines the challenge of robust audio-visual person verification using data recorded in multiple environments with various lighting conditions, irregular visual backgrounds, and diverse background noise. Audio-visual person verification could prove to be very useful in both physical …
-
Discovering Audio-Visual Associations in Narrated Videos of Human Activities
… correct associations between video scenes and audio utterances in an unsupervised way despite the imperfect correlation between the video and audio. The algorithm outperforms standard supervised learning algorithms. Among other things, this research shows that the performance of the algorithm …
-
Statistical modeling and analysis of audio-visual association in speech
… between associated and non-associated audio-visual observations. Various methods for modeling our audio-visual observations and ways of carrying out this test are studied and their relative performance is compared. We discuss issues that arise from the inherently high dimensional nature …
-
Preference between audio-visual recorded performance and audio-only recorded performance
… whether subjects have a preference for an audio-only (AO) presentation or an audio-visual (AV) presentation of the same piece of music. The research was conducted in two parts as a pilot study and as a main study. In the pilot study, the adult subjects were directed to the website YouTube …
-
The use of audio-visual aids in music education in California
… in 1945-46, I became interested in the use of audio-visual aids by the Army in teaching soldiers various Army procedures. I was subjected to an Army course in the use of Audio-visual aids, and later I designed visual aids for use in the Quartermaster School. At that time I compared the use of …
-
Audio-visual football video analysis, from structure detection to attention analysis
… replay and break were identified by low level visual features. A four-state hidden Markov model was trained to simulate transition processes among these shot classes. Since attack structures are the longest repetitive temporal unit in a sports video, a suffix tree is proposed to find the longest …
Page 1 of 16