{"id":{"repo_id":"wlv","oai_identifier":"oai:wlv.openrepository.com:2436/626132"},"canonical_url":"https://search.dev.ndltd.org/etd/wlv/oai:wlv.openrepository.com:2436/626132","repository":{"repo_id":"wlv","name":"University of Wolverhampton","base_url":"https://wlv.openrepository.com/server/oai/request"},"display":{"title":"Speech enhancement using multisensory cooperative computing","abstract":"This thesis investigates novel approaches for audio-visual speech enhancement (AVSE) through biologically inspired architectures, lightweight multimodal learning, and hybrid classical–deep frameworks. It advances three complementary themes of the speech enhancement problem. First, Multisensory Cooperative Computing (MCC) is introduced, a deep neural architecture inspired by the context-sensitive processing of two-point layer five pyramidal neurons (L5PCs). Unlike conventional point neuron models that indiscriminately process inputs, MCC adaptively filters and amplifies only contextually salient signals via a dendritic gating mechanism. Implemented on Xilinx Ultra- Scale+ MPSoC hardware, the system achieves substantial energy savings, up to 62% in semi-supervised settings and 1250× fewer energy demands per feedforward in supervised modes, by suppressing redundant synaptic activity. MCC establishes a paradigm for energy-efficient, high-capacity neuromorphic computing suited to real-time audio-visual learning. Second, to address AVSE on resource-constrained edge devices and the challenges of real-world noise, a novel target mask, the Ideal Smoothed Mask (ISM), is proposed. ISM combines morphological and spectral filtering for robust speech separation. A transfer-learning fusion framework maps visual lip movements to speech representations with enhanced temporal modelling. Nonlinear transfer functions and a multi-objective loss incorporating mutual information strengthen cross-modal attention. The resulting system improves generalisation, reduces mask complexity, and supports real-time enhancement on constrained hardware. Third, a lightweight AVSE framework is presented that merges classical spectral subtraction with visual speech detection. A CNN–LSTM module classifies short lip sequences into speech/no-speech labels to guide noise estimation and subtraction, overcoming the unreliability of audio-only voice activity detection (VAD) at low SNR. By isolating noise-only segments using robust lip-based cues, the approach preserves the interpretability and efficiency of classical methods while achieving substantial perceptual gains. Collectively, these contributions provide a unified vision for context-aware, resource-efficient, and explainable speech enhancement, bridging deep learning with neuro-inspired design and practical deployment.","abstract_html":"This thesis investigates novel approaches for audio-visual speech enhancement (AVSE) through biologically inspired architectures, lightweight multimodal learning, and hybrid classical–deep frameworks. It advances three complementary themes of the speech enhancement problem. First, Multisensory Cooperative Computing (MCC) is introduced, a deep neural architecture inspired by the context-sensitive processing of two-point layer five pyramidal neurons (L5PCs). Unlike conventional point neuron models that indiscriminately process inputs, MCC adaptively filters and amplifies only contextually salient signals via a dendritic gating mechanism. Implemented on Xilinx Ultra- Scale+ MPSoC hardware, the system achieves substantial energy savings, up to 62% in semi-supervised settings and 1250× fewer energy demands per feedforward in supervised modes, by suppressing redundant synaptic activity. MCC establishes a paradigm for energy-efficient, high-capacity neuromorphic computing suited to real-time audio-visual learning. Second, to address AVSE on resource-constrained edge devices and the challenges of real-world noise, a novel target mask, the Ideal Smoothed Mask (ISM), is proposed. ISM combines morphological and spectral filtering for robust speech separation. A transfer-learning fusion framework maps visual lip movements to speech representations with enhanced temporal modelling. Nonlinear transfer functions and a multi-objective loss incorporating mutual information strengthen cross-modal attention. The resulting system improves generalisation, reduces mask complexity, and supports real-time enhancement on constrained hardware. Third, a lightweight AVSE framework is presented that merges classical spectral subtraction with visual speech detection. A CNN–LSTM module classifies short lip sequences into speech/no-speech labels to guide noise estimation and subtraction, overcoming the unreliability of audio-only voice activity detection (VAD) at low SNR. By isolating noise-only segments using robust lip-based cues, the approach preserves the interpretability and efficiency of classical methods while achieving substantial perceptual gains. Collectively, these contributions provide a unified vision for context-aware, resource-efficient, and explainable speech enhancement, bridging deep learning with neuro-inspired design and practical deployment.","abstract_has_math":false,"creators":["Ahmed, Khubaib"],"institution":"University of Wolverhampton","degree_name":"PhD","degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Frommholz, Ingo","Adeel, Ahsan"],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025","date_published":"2025","updated_at":"2026-07-24T06:09:29Z","subjects":["audio-visual speech enhancement","mask-based speech enhancement","lip-reading","multimodal","spectral analysis","STFT","short-time fourier transform","Mel-spectrum","attention","visual speech detection","multisensory cooperative computing"],"languages":[],"rights":[],"rights_urls":["https://wlv.openrepository.com/bitstreams/84d14ce7-4994-4abd-ae32-8227a4dc6f7a/download"],"identifier_entries":[]},"links":{"outbound_url":null,"outbound_label":null,"outbound_source":null},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Frommholz, Ingo","Adeel, Ahsan"]},{"key":"dc:contributor.sponsor","label":"Sponsor","values":["EPSRC"]},{"key":"dc:creator","label":"Author","values":["Ahmed, Khubaib"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2025"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["University of Wolverhampton"]},{"key":"dc:relation.isreferencedby","label":"Dc Relation Isreferencedby","values":["https://wlv.openrepository.com/handle/2436/626132"]},{"key":"dc:type","label":"Dc Type","values":["Thesis or dissertation"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["PhD"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["audio-visual speech enhancement","mask-based speech enhancement","lip-reading","multimodal","spectral analysis","STFT","short-time fourier transform","Mel-spectrum","attention","visual speech detection","multisensory cooperative computing"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["https://wlv.openrepository.com/bitstreams/84d14ce7-4994-4abd-ae32-8227a4dc6f7a/download"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://wlv.openrepository.com/bitstreams/10409722-dcb2-4885-a8f9-149ea06d8336/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["This thesis investigates novel approaches for audio-visual speech enhancement (AVSE) through biologically inspired architectures, lightweight multimodal learning, and hybrid classical–deep frameworks. It advances three complementary themes of the speech enhancement problem. First, Multisensory Cooperative Computing (MCC) is introduced, a deep neural architecture inspired by the context-sensitive processing of two-point layer five pyramidal neurons (L5PCs). Unlike conventional point neuron models that indiscriminately process inputs, MCC adaptively filters and amplifies only contextually salient signals via a dendritic gating mechanism. Implemented on Xilinx Ultra- Scale+ MPSoC hardware, the system achieves substantial energy savings, up to 62% in semi-supervised settings and 1250× fewer energy demands per feedforward in supervised modes, by suppressing redundant synaptic activity. MCC establishes a paradigm for energy-efficient, high-capacity neuromorphic computing suited to real-time audio-visual learning. Second, to address AVSE on resource-constrained edge devices and the challenges of real-world noise, a novel target mask, the Ideal Smoothed Mask (ISM), is proposed. ISM combines morphological and spectral filtering for robust speech separation. A transfer-learning fusion framework maps visual lip movements to speech representations with enhanced temporal modelling. Nonlinear transfer functions and a multi-objective loss incorporating mutual information strengthen cross-modal attention. The resulting system improves generalisation, reduces mask complexity, and supports real-time enhancement on constrained hardware. Third, a lightweight AVSE framework is presented that merges classical spectral subtraction with visual speech detection. A CNN–LSTM module classifies short lip sequences into speech/no-speech labels to guide noise estimation and subtraction, overcoming the unreliability of audio-only voice activity detection (VAD) at low SNR. By isolating noise-only segments using robust lip-based cues, the approach preserves the interpretability and efficiency of classical methods while achieving substantial perceptual gains. Collectively, these contributions provide a unified vision for context-aware, resource-efficient, and explainable speech enhancement, bridging deep learning with neuro-inspired design and practical deployment."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["7d782afa79ad5d923e703c87e30c2a64","6fac97cc6f90c0d988b145078d045870","79cc9c558089c2829a28510d0296eb06"]},{"key":"dc:title","label":"Title","values":["Speech enhancement using multisensory cooperative computing"]}]}],"canonical_facts":{"dc:contributor.advisor":["Frommholz, Ingo","Adeel, Ahsan"],"dc:contributor.sponsor":["EPSRC"],"dc:creator":["Ahmed, Khubaib"],"dc:date.issued":["2025"],"dc:description.abstract":["This thesis investigates novel approaches for audio-visual speech enhancement (AVSE) through biologically inspired architectures, lightweight multimodal learning, and hybrid classical–deep frameworks. It advances three complementary themes of the speech enhancement problem. First, Multisensory Cooperative Computing (MCC) is introduced, a deep neural architecture inspired by the context-sensitive processing of two-point layer five pyramidal neurons (L5PCs). Unlike conventional point neuron models that indiscriminately process inputs, MCC adaptively filters and amplifies only contextually salient signals via a dendritic gating mechanism. Implemented on Xilinx Ultra- Scale+ MPSoC hardware, the system achieves substantial energy savings, up to 62% in semi-supervised settings and 1250× fewer energy demands per feedforward in supervised modes, by suppressing redundant synaptic activity. MCC establishes a paradigm for energy-efficient, high-capacity neuromorphic computing suited to real-time audio-visual learning. Second, to address AVSE on resource-constrained edge devices and the challenges of real-world noise, a novel target mask, the Ideal Smoothed Mask (ISM), is proposed. ISM combines morphological and spectral filtering for robust speech separation. A transfer-learning fusion framework maps visual lip movements to speech representations with enhanced temporal modelling. Nonlinear transfer functions and a multi-objective loss incorporating mutual information strengthen cross-modal attention. The resulting system improves generalisation, reduces mask complexity, and supports real-time enhancement on constrained hardware. Third, a lightweight AVSE framework is presented that merges classical spectral subtraction with visual speech detection. A CNN–LSTM module classifies short lip sequences into speech/no-speech labels to guide noise estimation and subtraction, overcoming the unreliability of audio-only voice activity detection (VAD) at low SNR. By isolating noise-only segments using robust lip-based cues, the approach preserves the interpretability and efficiency of classical methods while achieving substantial perceptual gains. Collectively, these contributions provide a unified vision for context-aware, resource-efficient, and explainable speech enhancement, bridging deep learning with neuro-inspired design and practical deployment."],"dc:format.checksum.md5":["7d782afa79ad5d923e703c87e30c2a64","6fac97cc6f90c0d988b145078d045870","79cc9c558089c2829a28510d0296eb06"],"dc:identifier.uri":["https://wlv.openrepository.com/bitstreams/10409722-dcb2-4885-a8f9-149ea06d8336/download"],"dc:publisher.institution":["University of Wolverhampton"],"dc:relation.isreferencedby":["https://wlv.openrepository.com/handle/2436/626132"],"dc:rights":["https://wlv.openrepository.com/bitstreams/84d14ce7-4994-4abd-ae32-8227a4dc6f7a/download"],"dc:subject":["audio-visual speech enhancement","mask-based speech enhancement","lip-reading","multimodal","spectral analysis","STFT","short-time fourier transform","Mel-spectrum","attention","visual speech detection","multisensory cooperative computing"],"dc:title":["Speech enhancement using multisensory cooperative computing"],"dc:type":["Thesis or dissertation"],"dc:type.qualificationname":["PhD"]},"updated_at":"2026-07-24T06:09:29Z"}