Abstract
dc:description.abstractThis thesis encloses five publications which describe technologies for recording, analysing, manipulating, and reproducing spatial sound scenes, which confront many of the challenges associated with the development of systems capable of delivering high quality audio within virtual reality and augmented hearing contexts. The technologies detailed herein operate based upon microphone array signals, which have been transformed into the time-frequency domain. Through the adoption of an assumed sound-field model, an input sound scene may be parameterised and decomposed, which permits the optional manipulation and subsequent reproduction of the sound scene over an arbitrary playback setup. This type of processing often leads to a high degree of playback flexibility and perceived spatial accuracy, which would otherwise be unattainable when using signal-independent and non-parametric alternatives. The first contribution of this thesis concerns the parameterisation and rendering of microphone array room impulse responses, such that the spatial characteristics of a measured space may be imparted onto a monophonic input signal and reproduced over a target loudspeaker setup. The second contribution explores a parametric method for converting microphone array signals into the popular Ambisonics format, while placing specific emphasis on the use of microphone arrays that are mounted onto irregular/non-spherical geometries; such as head-worn devices, which may find application within future augmented reality contexts. The third contribution also concerns a head-worn microphone array, but instead utilised microphones that are sensitive to ultrasonic frequencies. The intention is for ultrasonic sound sources to be captured by the array and then down pitch-shifted to the audible range, while being spatialised in the same direction that the sound arrived from. A number of spatial audio effects and sound-field modification tools were then explored in the fourth contribution, which operate based upon Ambisonic signals as input and involve the use of a parametric rendering framework. The final contribution concerns the use of a distributed arrangement of multiple Ambisonic receivers, which may be used to capture the sound scene from multiple perspectives. Subsequent analysis and decomposition of the sound scene, into its individual components, enables reproduction at different positions; thus, allowing a listener to navigate through the recorded sound scene.
Degree
thesis:*- Department dc:contributor.department
- Signaalinkäsittelyn ja akustiikan laitos
- Grantor dc:publisher
- Aalto University
- Year dc:date.issued
- 2023
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- McCormack, Leo
- Advisors dc:contributor.advisor
-
- Politis, Archontis, Prof., Tampere University, Finland
- Pulkki, Ville, Prof., Aalto University, Finland
- Pulkki, Ville, Prof., Aalto University, Department of Information and Communications Engineering, Finland
- Contributors dc:contributor
-
- Aalto-yliopisto
- Aalto University
Rights
- Language dc:language.iso
- en
Identifiers
dc:identifier.*- Repository record dc:identifier.uri
- https://aaltodoc.aalto.fi/handle/123456789/120569