{"id":{"repo_id":"cambridge","oai_identifier":"oai:www.repository.cam.ac.uk:1810/379891"},"canonical_url":"https://search.dev.ndltd.org/etd/cambridge/oai:www.repository.cam.ac.uk:1810/379891","repository":{"repo_id":"cambridge","name":"Cambridge University","base_url":"https://api.repository.cam.ac.uk/server/oai/request"},"display":{"title":"Automatic assessment of voice similarity and its implications for forensic applications","abstract":"Speakers of varying degrees of similarity are required for forensic purposes, for instance officer or witness protection, relevant populations for forensic speaker recognition, or voice parades and earwitness assessments. In the past decade, the collection of speech has become much easier, offering the opportunity to assess the suitability of large quantities of speakers for such purposes. However, manual assessment of voices by a forensic phonetician or the appointment of lay listeners to judge voice similarity is time-consuming and costly, thus limiting the number of voices that can be processed in a case. Automatic similarity assessment of speakers could radically reduce processing time while expanding the voice search pool before additional verification by a trained phonetician. This thesis explores the selection of similar-sounding speakers based on perceptual judgements and automatically measured features for such forensic applications. The primary research questions addressed are whether automatic speaker recognition may be used to assess perceived voice similarity and how large databases may be filtered according to speaker similarity. The study employs various combinations of features, speaker modelling approaches, and distance measures and uses correlation analyses, clustering methods, and ranking to find subgroups of similar-sounding speakers. A listener experiment is conducted to gain a deeper understanding of the perception of extreme similarity among unrelated speakers. The research highlights that it is indeed possible to filter similar-sounding speakers from large databases in a semi-automatic manner to the level of hard-to-distinguish unrelated speakers. Applications in likelihood ratio-based forensic automatic speaker recognition, voice parades, and speech synthesis are explored to varying degrees with key findings including that the perceived similarity of the relevant population to the questioned speaker may bias the strength of evidence, and that similarity of synthetic speech to a target speaker may not be assessed in the same way as natural speech. The findings of this dissertation have implications for future research in the field of speaker similarity and have practical applications in forensics.","abstract_html":"Speakers of varying degrees of similarity are required for forensic purposes, for instance officer or witness protection, relevant populations for forensic speaker recognition, or voice parades and earwitness assessments. In the past decade, the collection of speech has become much easier, offering the opportunity to assess the suitability of large quantities of speakers for such purposes. However, manual assessment of voices by a forensic phonetician or the appointment of lay listeners to judge voice similarity is time-consuming and costly, thus limiting the number of voices that can be processed in a case. Automatic similarity assessment of speakers could radically reduce processing time while expanding the voice search pool before additional verification by a trained phonetician. This thesis explores the selection of similar-sounding speakers based on perceptual judgements and automatically measured features for such forensic applications. The primary research questions addressed are whether automatic speaker recognition may be used to assess perceived voice similarity and how large databases may be filtered according to speaker similarity. The study employs various combinations of features, speaker modelling approaches, and distance measures and uses correlation analyses, clustering methods, and ranking to find subgroups of similar-sounding speakers. A listener experiment is conducted to gain a deeper understanding of the perception of extreme similarity among unrelated speakers. The research highlights that it is indeed possible to filter similar-sounding speakers from large databases in a semi-automatic manner to the level of hard-to-distinguish unrelated speakers. Applications in likelihood ratio-based forensic automatic speaker recognition, voice parades, and speech synthesis are explored to varying degrees with key findings including that the perceived similarity of the relevant population to the questioned speaker may bias the strength of evidence, and that similarity of synthetic speech to a target speaker may not be assessed in the same way as natural speech. The findings of this dissertation have implications for future research in the field of speaker similarity and have practical applications in forensics.","abstract_has_math":false,"creators":["Gerlach, Linda"],"institution":"University of Cambridge","degree_name":"Doctor of Philosophy (PhD)","degree_level":"Doctoral","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["McDougall, Kirsty"],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-05-27","date_published":"2024-05-27","updated_at":"2026-07-22T22:24:28Z","subjects":["automatic speaker recognition","forensic phonetics","forensic voice comparison","likelihood ratio","relevant population","speaker clustering","speaker similarity","voice parades","voice perception","voice similarity","voice synthesis","voice twins"],"languages":["eng"],"rights":[],"rights_urls":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/5b773739-89aa-470c-9a3b-2e6d821bca90/download","http://purl.org/NET/rdflicense/allrightsreserved"],"identifier_entries":[{"key":"dc:creator.authoridentifier","label":"Author Identifier","values":["0000000256560803"],"render_values":[{"text":"0000-0002-5656-0803","href":"https://orcid.org/0000-0002-5656-0803","code":true}]}]},"links":{"outbound_url":"https://doi.org/10.17863/CAM.115862","outbound_label":"DOI","outbound_source":"dc:identifier.doi"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["McDougall, Kirsty"]},{"key":"dc:contributor.sponsor","label":"Sponsor","values":["Cambridge European & Selwyn Oxford Wave Research Studentship"]},{"key":"dc:creator","label":"Author","values":["Gerlach, Linda"]},{"key":"dc:creator.authoridentifier","label":"Author Identifier","values":["0000000256560803"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2024-05-27"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["University of Cambridge"]},{"key":"dc:relation.isreferencedby.uri","label":"Dc Relation Isreferencedby URI","values":["https://www.repository.cam.ac.uk/handle/1810/379891"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["Doctoral"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["Doctor of Philosophy (PhD)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["automatic speaker recognition","forensic phonetics","forensic voice comparison","likelihood ratio","relevant population","speaker clustering","speaker similarity","voice parades","voice perception","voice similarity","voice synthesis","voice twins"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/5b773739-89aa-470c-9a3b-2e6d821bca90/download","http://purl.org/NET/rdflicense/allrightsreserved"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.17863/CAM.115862"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/1afd3b48-0e39-408c-949a-c450e29de835/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Speakers of varying degrees of similarity are required for forensic purposes, for instance officer or witness protection, relevant populations for forensic speaker recognition, or voice parades and earwitness assessments. In the past decade, the collection of speech has become much easier, offering the opportunity to assess the suitability of large quantities of speakers for such purposes. However, manual assessment of voices by a forensic phonetician or the appointment of lay listeners to judge voice similarity is time-consuming and costly, thus limiting the number of voices that can be processed in a case. Automatic similarity assessment of speakers could radically reduce processing time while expanding the voice search pool before additional verification by a trained phonetician. This thesis explores the selection of similar-sounding speakers based on perceptual judgements and automatically measured features for such forensic applications. The primary research questions addressed are whether automatic speaker recognition may be used to assess perceived voice similarity and how large databases may be filtered according to speaker similarity. The study employs various combinations of features, speaker modelling approaches, and distance measures and uses correlation analyses, clustering methods, and ranking to find subgroups of similar-sounding speakers. A listener experiment is conducted to gain a deeper understanding of the perception of extreme similarity among unrelated speakers. The research highlights that it is indeed possible to filter similar-sounding speakers from large databases in a semi-automatic manner to the level of hard-to-distinguish unrelated speakers. Applications in likelihood ratio-based forensic automatic speaker recognition, voice parades, and speech synthesis are explored to varying degrees with key findings including that the perceived similarity of the relevant population to the questioned speaker may bias the strength of evidence, and that similarity of synthetic speech to a target speaker may not be assessed in the same way as natural speech. The findings of this dissertation have implications for future research in the field of speaker similarity and have practical applications in forensics."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["78c3829ab259e4a0f6bfab6929626323","87eda9de84448d1f82354d60eee3eb5f"]},{"key":"dc:title","label":"Title","values":["Automatic assessment of voice similarity and its implications for forensic applications"]}]}],"canonical_facts":{"dc:contributor.advisor":["McDougall, Kirsty"],"dc:contributor.sponsor":["Cambridge European & Selwyn Oxford Wave Research Studentship"],"dc:creator":["Gerlach, Linda"],"dc:creator.authoridentifier":["0000000256560803"],"dc:date.issued":["2024-05-27"],"dc:description.abstract":["Speakers of varying degrees of similarity are required for forensic purposes, for instance officer or witness protection, relevant populations for forensic speaker recognition, or voice parades and earwitness assessments. In the past decade, the collection of speech has become much easier, offering the opportunity to assess the suitability of large quantities of speakers for such purposes. However, manual assessment of voices by a forensic phonetician or the appointment of lay listeners to judge voice similarity is time-consuming and costly, thus limiting the number of voices that can be processed in a case. Automatic similarity assessment of speakers could radically reduce processing time while expanding the voice search pool before additional verification by a trained phonetician. This thesis explores the selection of similar-sounding speakers based on perceptual judgements and automatically measured features for such forensic applications. The primary research questions addressed are whether automatic speaker recognition may be used to assess perceived voice similarity and how large databases may be filtered according to speaker similarity. The study employs various combinations of features, speaker modelling approaches, and distance measures and uses correlation analyses, clustering methods, and ranking to find subgroups of similar-sounding speakers. A listener experiment is conducted to gain a deeper understanding of the perception of extreme similarity among unrelated speakers. The research highlights that it is indeed possible to filter similar-sounding speakers from large databases in a semi-automatic manner to the level of hard-to-distinguish unrelated speakers. Applications in likelihood ratio-based forensic automatic speaker recognition, voice parades, and speech synthesis are explored to varying degrees with key findings including that the perceived similarity of the relevant population to the questioned speaker may bias the strength of evidence, and that similarity of synthetic speech to a target speaker may not be assessed in the same way as natural speech. The findings of this dissertation have implications for future research in the field of speaker similarity and have practical applications in forensics."],"dc:format.checksum.md5":["78c3829ab259e4a0f6bfab6929626323","87eda9de84448d1f82354d60eee3eb5f"],"dc:identifier.doi":["https://doi.org/10.17863/CAM.115862"],"dc:identifier.uri":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/1afd3b48-0e39-408c-949a-c450e29de835/download"],"dc:language":["eng"],"dc:publisher.institution":["University of Cambridge"],"dc:relation.isreferencedby.uri":["https://www.repository.cam.ac.uk/handle/1810/379891"],"dc:rights":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/5b773739-89aa-470c-9a3b-2e6d821bca90/download","http://purl.org/NET/rdflicense/allrightsreserved"],"dc:subject":["automatic speaker recognition","forensic phonetics","forensic voice comparison","likelihood ratio","relevant population","speaker clustering","speaker similarity","voice parades","voice perception","voice similarity","voice synthesis","voice twins"],"dc:title":["Automatic assessment of voice similarity and its implications for forensic applications"],"dc:type":["Thesis"],"dc:type.qualificationlevel":["Doctoral"],"dc:type.qualificationname":["Doctor of Philosophy (PhD)"]},"updated_at":"2026-07-22T22:24:28Z"}