Back to results

Università degli studi di Trento

Information Fusion Approaches for Distant Speech Recognition in a Multi-microphone Setting

Abstract

dc:description

It is a well known fact that high quality Automatic Speech Recognition is still difficult to guarantee under conditions in which the speaker is distant from the microphone due to the distortions caused by acoustic phenomena, such as noise and reverberation. Among the different research directions pursued around this problem, the adoption of multi-channel approaches is of great interest to the community given the potential of taking advantage of information diversity. In this thesis we elaborate on approaches that exploit different instances of a sound source, captured by various largely spaced microphones, in order to extract a Distant Speech Recognition hypothesis. Two original solutions are presented, based on information fusion approaches at different levels of the recognition system, one at front-end stage and one at post-decoding stage, namely for the problems of channel selection (CS) and hypothesis combination. First, a new CS framework is proposed. Cepstral distance (CD), which is effectively applied in other acoustic processing fields, is the basis of the CS method developed. Experimental results confirmed the advantages of a CD-based selection schema under different scenarios. The second contribution concerns the combination of information extracted from the individual decoding processes performed over the multiple captured signals. It is shown how temporal cues can be identified in the hypothesis space, and be beneficial for the elaboration of a multi-microphone confusion network, from which the final speech transcription is derived. The proposed methods are applicable in a setting equipped with synchronized distributed microphones, independently of the proximity between the sensors. Analysis of the novel concepts were performed over synthetic and real-captured data. Both approaches achieved positive results at the different assessment tasks they were exposed to.

Degree

thesis:*
Grantor dc:publisher
Università degli studi di Trento
Year dc:date
2016

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Guerrero Flores, Cristina Maritza
Contributors dc:contributor
  • Omologo, Maurizio

Subjects

dc:subject × 1

Rights

dc:rights
Statement dc:rights
  • info:eu-repo/semantics/closedAccess
  • license:Tutti i diritti riservati (All rights reserved)
Language dc:language
eng

Identifiers

dc:identifier.*
OAI identifier oai:identifier
oai:iris.unitn.it:11572/368955

Chain of custody

source
Harvested from
Università degli Studi di Trento
Base URL
iris.unitn.it/oai/request
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Guerrero Flores, Cristina Maritza. Information Fusion Approaches for Distant Speech Recognition in a Multi-microphone Setting. Università degli studi di Trento, 2016. https://hdl.handle.net/11572/368955