George Mason University
Securing Voice Processing Systems from Malicious Audio Attacks
Abstract
In an era where voice processing systems have become an integral part of our daily lives, it is of paramount importance to ensure their security and protection against malicious voice attacks. This dissertation focuses on understanding the vulnerabilities inherent in voice processing systems and presenting effective defense methods to safeguard these critical systems from exploitation. Our exploration begins by harnessing the power of the spectrum compensation technique, which serves as a cornerstone in identifying potential attack surfaces and devising appropriate defense strategies. Through meticulous analysis, we uncover a previously unknown form of voice attack - modulated replay attack. This insidious attack leverages the spectrum compensation technique to neutralize spectrum distortion caused by loudspeakers, rendering conventional spectrum-based defense methods ineffective. Our research does not stop at exposing this threat. We propose a robust defense method that combines both frequency-domain and time-domain features to effectively combat modulated replay attacks. Expanding our horizons beyond the frequency and time domains, we explore the physical properties of voice propagation. In the context of driverless car scenarios, we harness these physical properties as well as time-frequency properties to identify legitimate voice commands, enhancing the security of voice interactions in vehicular environments. In our pursuit of comprehensive defense strategies, we shed light on the versatility of spectrum compensation technique. Not only does it provide attack support for modulated replay attacks, but it also delivers defense capability in defeating spectrum reduction attacks - a practical menace to content moderation systems. The structure of this dissertation is thoughtfully designed to guide readers through the intricacy of voice processing system vulnerabilities and their corresponding defenses. In Chapter 1, we commence with an illuminating chapter that provides essential background information and a thorough review of the existing literature about attacks against voice processing systems. Building upon this foundation, Chapter 2 unveils the modulated replay attacks, showcases the spectrum compensation technique, and presents our innovative two-domain defense method. Chapter 3 represents an in-vehicle secure automatic speech recognition system. This system has leveraged multiple audio attributes, including time-domain features, frequency features, and the physical properties of voice propagation, to ensure the security of voice commands within the dynamic environment of driverless cars. Chapter 4 introduces an acoustic compensation system as a new defense against the spectrum reduction attacks that are a big threat to the content moderation systems. Through the deployment of this system, we reinforce the robustness and reliability of content filtering processes, maintaining the sanctity of online platforms. Finally, in Chapter 5, we conclude our work by summarizing the insights gained and emphasizing the importance of ongoing research. We outline potential future directions that may further strengthen voice processing systems against new emerging threats, ensuring the continued security and trustworthiness of these critical technologies. It is our fervent hope that this dissertation serves as an inspiring resource, enlightening readers about the vulnerabilities of voice processing systems and empowering them with the knowledge to develop effective defense strategies. We invite you to embark on this exploration, fostering a safer and more secure landscape for voice-based interactions.
Author and committee
dc:creator, dc:contributor.*- Author
-
- Wang, Shu
Subjects
dc:subject × 6Identifiers
dc:identifier.*- Identifier
- hdl:1920/13731
- OAI identifier oai:identifier
- oai:MARS:1920/13731