{"id":{"repo_id":"birmingham","oai_identifier":"oai:etheses.bham.ac.uk:16"},"canonical_url":"https://search.dev.ndltd.org/etd/birmingham/oai:etheses.bham.ac.uk:16","repository":{"repo_id":"birmingham","name":"University of Birmingham","base_url":"https://etheses.bham.ac.uk/cgi/oai2"},"display":{"title":"Speech recognition in programmable logic","abstract":"Speech recognition is a computationally demanding task, especially the decoding part, which converts pre-processed speech data into words or sub-word units, and which incorporates Viterbi decoding and Gaussian distribution calculations. In this thesis, this part of the recognition process is implemented in programmable logic, specifically, on a field-programmable gate array (FPGA). Relevant background material about speech recognition is presented, along with a critical review of previous hardware implementations. Designs for a decoder suitable for implementation in hardware are then described. These include details of how multiple speech files can be processed in parallel, and an original implementation of an algorithm for summing Gaussian mixture components in the log domain. These designs are then implemented on an FPGA. An assessment is made as to how appropriate it is to use hardware for speech recognition. It is concluded that while certain parts of the recognition algorithm are not well suited to this medium, much of it is, and so an efficient implementation is possible. Also presented is an original analysis of the requirements of speech recognition for hardware and software, which relates the parameters that dictate the complexity of the system to processing speed and bandwidth. The FPGA implementations are compared to equivalent software, written for that purpose. For a contemporary FPGA and processor, the FPGA outperforms the software by an order of magnitude.","abstract_html":"Speech recognition is a computationally demanding task, especially the decoding part, which converts pre-processed speech data into words or sub-word units, and which incorporates Viterbi decoding and Gaussian distribution calculations. In this thesis, this part of the recognition process is implemented in programmable logic, specifically, on a field-programmable gate array (FPGA). Relevant background material about speech recognition is presented, along with a critical review of previous hardware implementations. Designs for a decoder suitable for implementation in hardware are then described. These include details of how multiple speech files can be processed in parallel, and an original implementation of an algorithm for summing Gaussian mixture components in the log domain. These designs are then implemented on an FPGA. An assessment is made as to how appropriate it is to use hardware for speech recognition. It is concluded that while certain parts of the recognition algorithm are not well suited to this medium, much of it is, and so an efficient implementation is possible. Also presented is an original analysis of the requirements of speech recognition for hardware and software, which relates the parameters that dictate the complexity of the system to processing speed and bandwidth. The FPGA implementations are compared to equivalent software, written for that purpose. For a contemporary FPGA and processor, the FPGA outperforms the software by an order of magnitude.","abstract_has_math":false,"creators":["Melnikoff, Stephen Jonathan"],"institution":"University of Birmingham","degree_name":"d_ph","degree_level":"d_ph","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2003,"date_issued":"2003-12","date_published":"2003-12","updated_at":"2026-07-24T01:10:41Z","subjects":["TK Electrical engineering. Electronics Nuclear engineering","QA75 Electronic computers. Computer science"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://etheses.bham.ac.uk//id/eprint/16/1/Decl_IS_Melnikoff03PhD.jpg","outbound_label":"Repository record","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.sponsor","label":"Sponsor","values":["na"]},{"key":"dc:creator","label":"Author","values":["Melnikoff, Stephen Jonathan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2003-12"]},{"key":"dc:date.issued","label":"Date","values":["2003-12"]},{"key":"dc:publisher.department","label":"Dc Publisher Department","values":["School of Engineering","School of Engineering, Department of Electronic, Electrical and Systems Engineering"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["University of Birmingham"]},{"key":"dc:relation.isreferencedby","label":"Dc Relation Isreferencedby","values":["http://etheses.bham.ac.uk//id/eprint/16/"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["d_ph"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["d_ph"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["TK Electrical engineering. Electronics Nuclear engineering","QA75 Electronic computers. Computer science"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["http://etheses.bham.ac.uk//id/eprint/16/1/Decl_IS_Melnikoff03PhD.jpg","http://etheses.bham.ac.uk//id/eprint/16/2/Melnikoff03PhD.pdf"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Speech recognition is a computationally demanding task, especially the decoding part, which converts pre-processed speech data into words or sub-word units, and which incorporates Viterbi decoding and Gaussian distribution calculations. In this thesis, this part of the recognition process is implemented in programmable logic, specifically, on a field-programmable gate array (FPGA). Relevant background material about speech recognition is presented, along with a critical review of previous hardware implementations. Designs for a decoder suitable for implementation in hardware are then described. These include details of how multiple speech files can be processed in parallel, and an original implementation of an algorithm for summing Gaussian mixture components in the log domain. These designs are then implemented on an FPGA. An assessment is made as to how appropriate it is to use hardware for speech recognition. It is concluded that while certain parts of the recognition algorithm are not well suited to this medium, much of it is, and so an efficient implementation is possible. Also presented is an original analysis of the requirements of speech recognition for hardware and software, which relates the parameters that dictate the complexity of the system to processing speed and bandwidth. The FPGA implementations are compared to equivalent software, written for that purpose. For a contemporary FPGA and processor, the FPGA outperforms the software by an order of magnitude."]},{"key":"dc:format","label":"Dc Format","values":["image/jpeg","application/pdf"]},{"key":"dc:title","label":"Title","values":["Speech recognition in programmable logic"]}]}],"canonical_facts":{"dc:contributor.sponsor":["na"],"dc:creator":["Melnikoff, Stephen Jonathan"],"dc:date":["2003-12"],"dc:date.issued":["2003-12"],"dc:description.abstract":["Speech recognition is a computationally demanding task, especially the decoding part, which converts pre-processed speech data into words or sub-word units, and which incorporates Viterbi decoding and Gaussian distribution calculations. In this thesis, this part of the recognition process is implemented in programmable logic, specifically, on a field-programmable gate array (FPGA). Relevant background material about speech recognition is presented, along with a critical review of previous hardware implementations. Designs for a decoder suitable for implementation in hardware are then described. These include details of how multiple speech files can be processed in parallel, and an original implementation of an algorithm for summing Gaussian mixture components in the log domain. These designs are then implemented on an FPGA. An assessment is made as to how appropriate it is to use hardware for speech recognition. It is concluded that while certain parts of the recognition algorithm are not well suited to this medium, much of it is, and so an efficient implementation is possible. Also presented is an original analysis of the requirements of speech recognition for hardware and software, which relates the parameters that dictate the complexity of the system to processing speed and bandwidth. The FPGA implementations are compared to equivalent software, written for that purpose. For a contemporary FPGA and processor, the FPGA outperforms the software by an order of magnitude."],"dc:format":["image/jpeg","application/pdf"],"dc:identifier.uri":["http://etheses.bham.ac.uk//id/eprint/16/1/Decl_IS_Melnikoff03PhD.jpg","http://etheses.bham.ac.uk//id/eprint/16/2/Melnikoff03PhD.pdf"],"dc:publisher.department":["School of Engineering","School of Engineering, Department of Electronic, Electrical and Systems Engineering"],"dc:publisher.institution":["University of Birmingham"],"dc:relation.isreferencedby":["http://etheses.bham.ac.uk//id/eprint/16/"],"dc:subject":["TK Electrical engineering. Electronics Nuclear engineering","QA75 Electronic computers. Computer science"],"dc:title":["Speech recognition in programmable logic"],"dc:type":["Thesis"],"dc:type.qualificationlevel":["d_ph"],"dc:type.qualificationname":["d_ph"]},"updated_at":"2026-07-24T01:10:41Z"}