{"id":{"repo_id":"texas","oai_identifier":"oai:repositories.lib.utexas.edu:2152/65752"},"canonical_url":"https://search.dev.ndltd.org/etd/texas/oai:repositories.lib.utexas.edu:2152/65752","repository":{"repo_id":"texas","name":"University of Texas","base_url":"https://repositories.lib.utexas.edu/server/oai/request"},"display":{"title":"Taking attention away from the auditory modality : investigations of the effect on speech processing using machine learning","abstract":"Real-world speech processing often takes place in complex multisensory environments. Listeners may need to prioritize sensory inputs from modalities other than audition. Selective attention is thought to be critical in selecting the sensory modality most relevant to the task at hand. Two critical research questions have driven crossmodal attention research thus far: first, how early does crossmodal attention influences processing in the unattended modality? Second, is there a limitation in attentional resources between sensory modalities? Set within the context of this prior work, this dissertation aims to examine the effects of crossmodal attention on speech processing when sensory inputs from vision are prioritized. In study 1, we demonstrate that modulating visual perceptual load can impact the early sensory representation of linguistically-relevant pitch contours (Mandarin tones), a suprasegmental feature that is critical to the percept of lexical tones. Further, we provide novel evidence that the impact of the visual load is highly dependent on the predictability of the incoming speech stream. In study 2, we utilized ecologically valid, continuous speech and tested the extent to which dividing attention to a visual task affects neural processing of speech signals. We show that dividing attention between auditory and visual tasks leads to both behavioral and electrophysiological costs in the processing of continuous speech stimuli. The results also demonstrate that the neural encoding of suprasegmental features (e.g., envelope and fundamental frequency) in continuous speech is modulated by diverting attention away from the auditory modality. In contrast, the neural encoding of segmental features (e.g., phonetic features) may be unaffected by taking attention away from the auditory stream. The theoretical and practical implications of the two studies are discussed.","abstract_html":"Real-world speech processing often takes place in complex multisensory environments. Listeners may need to prioritize sensory inputs from modalities other than audition. Selective attention is thought to be critical in selecting the sensory modality most relevant to the task at hand. Two critical research questions have driven crossmodal attention research thus far: first, how early does crossmodal attention influences processing in the unattended modality? Second, is there a limitation in attentional resources between sensory modalities? Set within the context of this prior work, this dissertation aims to examine the effects of crossmodal attention on speech processing when sensory inputs from vision are prioritized. In study 1, we demonstrate that modulating visual perceptual load can impact the early sensory representation of linguistically-relevant pitch contours (Mandarin tones), a suprasegmental feature that is critical to the percept of lexical tones. Further, we provide novel evidence that the impact of the visual load is highly dependent on the predictability of the incoming speech stream. In study 2, we utilized ecologically valid, continuous speech and tested the extent to which dividing attention to a visual task affects neural processing of speech signals. We show that dividing attention between auditory and visual tasks leads to both behavioral and electrophysiological costs in the processing of continuous speech stimuli. The results also demonstrate that the neural encoding of suprasegmental features (e.g., envelope and fundamental frequency) in continuous speech is modulated by diverting attention away from the auditory modality. In contrast, the neural encoding of segmental features (e.g., phonetic features) may be unaffected by taking attention away from the auditory stream. The theoretical and practical implications of the two studies are discussed.","abstract_has_math":false,"creators":["Xie, Zilong"],"institution":"The University of Texas at Austin","degree_name":"Doctor of Philosophy","degree_level":"Doctoral","degree_discipline":"Communication Sciences &amp; Disorders","degree_department":null,"school":null,"contributors":[],"advisors":["Chandrasekaran, Bharath"],"committee_chairs":[],"committee_members":["Beevers, Christopher G","Champlin, Craig A","Liu, Chang"],"year":2018,"date_issued":"2018-06-22","date_published":"2018-06-22","updated_at":"2026-07-24T05:01:12Z","subjects":["Crossmodal attention","Speech processing","Suprasegmental features","Segmental features","Machine learning"],"languages":["en"],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["doi:10.15781/T20000H8N"],"render_values":[{"text":"doi:10.15781/T20000H8N","href":"https://doi.org/10.15781/T20000H8N","code":true}]}]},"links":{"outbound_url":"http://hdl.handle.net/2152/65752","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Chandrasekaran, Bharath"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Beevers, Christopher G","Champlin, Craig A","Liu, Chang"]},{"key":"dc:creator","label":"Author","values":["Xie, Zilong"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2018-07-24T16:52:27Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2018-07-24T16:52:27Z"]},{"key":"dc:date.issued","label":"Date","values":["2018-06-22"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Communication Sciences &amp; Disorders"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Doctoral"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctor of Philosophy"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["The University of Texas at Austin"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Crossmodal attention","Speech processing","Suprasegmental features","Segmental features","Machine learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["doi:10.15781/T20000H8N"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["http://hdl.handle.net/2152/65752"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Real-world speech processing often takes place in complex multisensory environments. Listeners may need to prioritize sensory inputs from modalities other than audition. Selective attention is thought to be critical in selecting the sensory modality most relevant to the task at hand. Two critical research questions have driven crossmodal attention research thus far: first, how early does crossmodal attention influences processing in the unattended modality? Second, is there a limitation in attentional resources between sensory modalities? Set within the context of this prior work, this dissertation aims to examine the effects of crossmodal attention on speech processing when sensory inputs from vision are prioritized. In study 1, we demonstrate that modulating visual perceptual load can impact the early sensory representation of linguistically-relevant pitch contours (Mandarin tones), a suprasegmental feature that is critical to the percept of lexical tones. Further, we provide novel evidence that the impact of the visual load is highly dependent on the predictability of the incoming speech stream. In study 2, we utilized ecologically valid, continuous speech and tested the extent to which dividing attention to a visual task affects neural processing of speech signals. We show that dividing attention between auditory and visual tasks leads to both behavioral and electrophysiological costs in the processing of continuous speech stimuli. The results also demonstrate that the neural encoding of suprasegmental features (e.g., envelope and fundamental frequency) in continuous speech is modulated by diverting attention away from the auditory modality. In contrast, the neural encoding of segmental features (e.g., phonetic features) may be unaffected by taking attention away from the auditory stream. The theoretical and practical implications of the two studies are discussed."]},{"key":"dc:format.mimetype","label":"Dc Format Mimetype","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Taking attention away from the auditory modality : investigations of the effect on speech processing using machine learning"]}]}],"canonical_facts":{"dc:contributor.advisor":["Chandrasekaran, Bharath"],"dc:contributor.committeemember":["Beevers, Christopher G","Champlin, Craig A","Liu, Chang"],"dc:creator":["Xie, Zilong"],"dc:date.accessioned":["2018-07-24T16:52:27Z"],"dc:date.available":["2018-07-24T16:52:27Z"],"dc:date.issued":["2018-06-22"],"dc:description.abstract":["Real-world speech processing often takes place in complex multisensory environments. Listeners may need to prioritize sensory inputs from modalities other than audition. Selective attention is thought to be critical in selecting the sensory modality most relevant to the task at hand. Two critical research questions have driven crossmodal attention research thus far: first, how early does crossmodal attention influences processing in the unattended modality? Second, is there a limitation in attentional resources between sensory modalities? Set within the context of this prior work, this dissertation aims to examine the effects of crossmodal attention on speech processing when sensory inputs from vision are prioritized. In study 1, we demonstrate that modulating visual perceptual load can impact the early sensory representation of linguistically-relevant pitch contours (Mandarin tones), a suprasegmental feature that is critical to the percept of lexical tones. Further, we provide novel evidence that the impact of the visual load is highly dependent on the predictability of the incoming speech stream. In study 2, we utilized ecologically valid, continuous speech and tested the extent to which dividing attention to a visual task affects neural processing of speech signals. We show that dividing attention between auditory and visual tasks leads to both behavioral and electrophysiological costs in the processing of continuous speech stimuli. The results also demonstrate that the neural encoding of suprasegmental features (e.g., envelope and fundamental frequency) in continuous speech is modulated by diverting attention away from the auditory modality. In contrast, the neural encoding of segmental features (e.g., phonetic features) may be unaffected by taking attention away from the auditory stream. The theoretical and practical implications of the two studies are discussed."],"dc:format.mimetype":["application/pdf"],"dc:identifier":["doi:10.15781/T20000H8N"],"dc:identifier.uri":["http://hdl.handle.net/2152/65752"],"dc:language.iso":["en"],"dc:subject":["Crossmodal attention","Speech processing","Suprasegmental features","Segmental features","Machine learning"],"dc:title":["Taking attention away from the auditory modality : investigations of the effect on speech processing using machine learning"],"dc:type":["Thesis"],"thesis:degree_discipline":["Communication Sciences &amp; Disorders"],"thesis:degree_level":["Doctoral"],"thesis:degree_name":["Doctor of Philosophy"],"thesis:institution_name":["The University of Texas at Austin"]},"updated_at":"2026-07-24T05:01:12Z"}