{"id":{"repo_id":"cent-lancashire","oai_identifier":"oai:clok.uclan.ac.uk:20360"},"canonical_url":"https://search.dev.ndltd.org/etd/cent-lancashire/oai:clok.uclan.ac.uk:20360","repository":{"repo_id":"cent-lancashire","name":"University of Central Lancashire","base_url":"https://clok.uclan.ac.uk/cgi/oai2"},"display":{"title":"Lossy compression of speech using perceptual criteria","abstract":"The research contained in this thesis provides an investigation into a new method of minimising the perceptual differences when encoding digitised speech. An application of the perceptual criteria is described in the context of a codebook encoding methodology Some of the background studies covered aspects of psychoacoustics, in particular the effects of the human outer, middle and inner ear. Models approximating each region of the ear are utilised and concatenated into a single overall auditory response path model. As the objective of the research is to encode and decode speech waveforms, some study into how speech is produced and the classification of speech sounds is required. From this there is a description of a basic speech production model which is modelled as a digital filter. A review of the main categories for coding schemes that are currently employed is presented along with commonly used coding methods. In particular the codebook coding method is reviewed in sufficient detail to contrast with the new coding method. The development of a new perceptual minimisation criterion which relies on dual application of the auditory response path model on the original and reconstructed speech waveforms is described. In this the ordering of eodebook searches, the frequency spectrum used as the search target, windowing functions with durations and placement are all analysed to determine the optimum encoder design. Also described are a number of prospective gain algorithms which cover both time and frequency domain implementations. A new encoder is constructed which fully integrates the new perceptual criterion into the minimisation of the original and reconstructed speech waveforms. In the minimisation no part of the traditional encoder method is used, however both methods use a similar technique for determining gain factors. Speech derived from both encoders was subjectively assessed by a number of untrained, independent listeners. The results presented show that both methods are comparable but there is a slight preference towards the traditional encoder. A measure of the complexity indicated that the new minimisation method is also more complex than the traditional encoder.","abstract_html":"The research contained in this thesis provides an investigation into a new method of minimising the perceptual differences when encoding digitised speech. An application of the perceptual criteria is described in the context of a codebook encoding methodology Some of the background studies covered aspects of psychoacoustics, in particular the effects of the human outer, middle and inner ear. Models approximating each region of the ear are utilised and concatenated into a single overall auditory response path model. As the objective of the research is to encode and decode speech waveforms, some study into how speech is produced and the classification of speech sounds is required. From this there is a description of a basic speech production model which is modelled as a digital filter. A review of the main categories for coding schemes that are currently employed is presented along with commonly used coding methods. In particular the codebook coding method is reviewed in sufficient detail to contrast with the new coding method. The development of a new perceptual minimisation criterion which relies on dual application of the auditory response path model on the original and reconstructed speech waveforms is described. In this the ordering of eodebook searches, the frequency spectrum used as the search target, windowing functions with durations and placement are all analysed to determine the optimum encoder design. Also described are a number of prospective gain algorithms which cover both time and frequency domain implementations. A new encoder is constructed which fully integrates the new perceptual criterion into the minimisation of the original and reconstructed speech waveforms. In the minimisation no part of the traditional encoder method is used, however both methods use a similar technique for determining gain factors. Speech derived from both encoders was subjectively assessed by a number of untrained, independent listeners. The results presented show that both methods are comparable but there is a slight preference towards the traditional encoder. A measure of the complexity indicated that the new minimisation method is also more complex than the traditional encoder.","abstract_has_math":false,"creators":["O'Donnell, Michael"],"institution":"University of Central Lancashire","degree_name":"phd","degree_level":"doctoral","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":1998,"date_issued":"1998-07","date_published":"1998-07","updated_at":"2026-07-24T01:36:10Z","subjects":["H150 - Engineering design","W240 - Industrial/product design"],"languages":["en"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":null,"outbound_label":null,"outbound_source":null},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["O'Donnell, Michael"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["1998-07-01"]},{"key":"dc:date.issued","label":"Date","values":["1998-07"]},{"key":"dc:publisher.department","label":"Dc Publisher Department","values":["Engineering and Product Design"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["University of Central Lancashire"]},{"key":"dc:relation.isreferencedby","label":"Dc Relation Isreferencedby","values":["https://knowledge.lancashire.ac.uk/id/eprint/20360/"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["doctoral"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["phd"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["H150 - Engineering design","W240 - Industrial/product design"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://knowledge.lancashire.ac.uk/id/eprint/20360/1/20360%20Michael%20O%20Donell%20July98%20Lossy%20Compression%20of%20Speech%20Using%20Perceptual%20Criteria%20Doctor%20of%20Philosophy%20unpublished%20July98%20University%20of%20Central%20Lancashire%20engineering%20and%20product%20design%20175.pdf"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["The research contained in this thesis provides an investigation into a new method of minimising the perceptual differences when encoding digitised speech. An application of the perceptual criteria is described in the context of a codebook encoding methodology Some of the background studies covered aspects of psychoacoustics, in particular the effects of the human outer, middle and inner ear. Models approximating each region of the ear are utilised and concatenated into a single overall auditory response path model. As the objective of the research is to encode and decode speech waveforms, some study into how speech is produced and the classification of speech sounds is required. From this there is a description of a basic speech production model which is modelled as a digital filter. A review of the main categories for coding schemes that are currently employed is presented along with commonly used coding methods. In particular the codebook coding method is reviewed in sufficient detail to contrast with the new coding method. The development of a new perceptual minimisation criterion which relies on dual application of the auditory response path model on the original and reconstructed speech waveforms is described. In this the ordering of eodebook searches, the frequency spectrum used as the search target, windowing functions with durations and placement are all analysed to determine the optimum encoder design. Also described are a number of prospective gain algorithms which cover both time and frequency domain implementations. A new encoder is constructed which fully integrates the new perceptual criterion into the minimisation of the original and reconstructed speech waveforms. In the minimisation no part of the traditional encoder method is used, however both methods use a similar technique for determining gain factors. Speech derived from both encoders was subjectively assessed by a number of untrained, independent listeners. The results presented show that both methods are comparable but there is a slight preference towards the traditional encoder. A measure of the complexity indicated that the new minimisation method is also more complex than the traditional encoder."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Lossy compression of speech using perceptual criteria"]}]}],"canonical_facts":{"dc:creator":["O'Donnell, Michael"],"dc:date":["1998-07-01"],"dc:date.issued":["1998-07"],"dc:description.abstract":["The research contained in this thesis provides an investigation into a new method of minimising the perceptual differences when encoding digitised speech. An application of the perceptual criteria is described in the context of a codebook encoding methodology Some of the background studies covered aspects of psychoacoustics, in particular the effects of the human outer, middle and inner ear. Models approximating each region of the ear are utilised and concatenated into a single overall auditory response path model. As the objective of the research is to encode and decode speech waveforms, some study into how speech is produced and the classification of speech sounds is required. From this there is a description of a basic speech production model which is modelled as a digital filter. A review of the main categories for coding schemes that are currently employed is presented along with commonly used coding methods. In particular the codebook coding method is reviewed in sufficient detail to contrast with the new coding method. The development of a new perceptual minimisation criterion which relies on dual application of the auditory response path model on the original and reconstructed speech waveforms is described. In this the ordering of eodebook searches, the frequency spectrum used as the search target, windowing functions with durations and placement are all analysed to determine the optimum encoder design. Also described are a number of prospective gain algorithms which cover both time and frequency domain implementations. A new encoder is constructed which fully integrates the new perceptual criterion into the minimisation of the original and reconstructed speech waveforms. In the minimisation no part of the traditional encoder method is used, however both methods use a similar technique for determining gain factors. Speech derived from both encoders was subjectively assessed by a number of untrained, independent listeners. The results presented show that both methods are comparable but there is a slight preference towards the traditional encoder. A measure of the complexity indicated that the new minimisation method is also more complex than the traditional encoder."],"dc:format":["application/pdf"],"dc:identifier.uri":["https://knowledge.lancashire.ac.uk/id/eprint/20360/1/20360%20Michael%20O%20Donell%20July98%20Lossy%20Compression%20of%20Speech%20Using%20Perceptual%20Criteria%20Doctor%20of%20Philosophy%20unpublished%20July98%20University%20of%20Central%20Lancashire%20engineering%20and%20product%20design%20175.pdf"],"dc:language":["en"],"dc:publisher.department":["Engineering and Product Design"],"dc:publisher.institution":["University of Central Lancashire"],"dc:relation.isreferencedby":["https://knowledge.lancashire.ac.uk/id/eprint/20360/"],"dc:subject":["H150 - Engineering design","W240 - Industrial/product design"],"dc:title":["Lossy compression of speech using perceptual criteria"],"dc:type":["Thesis"],"dc:type.qualificationlevel":["doctoral"],"dc:type.qualificationname":["phd"]},"updated_at":"2026-07-24T01:36:10Z"}