{"id":{"repo_id":"unsw","oai_identifier":"oai:unsworks.library.unsw.edu.au:1959.4/70882"},"canonical_url":"https://search.dev.ndltd.org/etd/unsw/oai:unsworks.library.unsw.edu.au:1959.4/70882","repository":{"repo_id":"unsw","name":"University of New South Wales","base_url":"https://unsworks.unsw.edu.au/oai/provider"},"display":{"title":"The development of reference standards for genomics.","abstract":"Despite decades from the publication of the first draft, the reference human genome remains incomplete, with many unsolved difficult regions. Next-generation sequencing (NGS) has become a central tool for the detection of genetic variation in biological research and clinical diagnosis. However, despite its advantages, NGS suffers from errors and biases that can confound the detection of true variants. Reference standards are control materials, with known properties, against which to test performance. In recent years, the decreasing costs of DNA synthesis have enabled the creative design of synthetic reference controls for genomics. This thesis describes the development of synthetic DNA controls to qualitatively and quantitatively analyse genome features. First, a synthetic ladder, consisting of a single DNA molecule with artificial sequence elements at known copy-numbers, to accurately measure sequence abundance in NGS libraries. The synthetic ladder provides a universal reference to quantify and mitigate the impact of technical variation, independent of the human genome, improving quantitative comparisons within and between samples. Secondly, a synthetic chromosome containing mirrored representations of diverse clinically relevant variants and human genome features, such as HLA alleles and immune receptors. The synthetic chromosome provides a ground-truth reference, with unambiguous representation of low-confidence and difficult-to-sequence genome regions. Therefore, I used the synthetic chromosome to benchmark different experimental and analytical pipelines, highlighting weaknesses and strengths, and providing best-practices guidelines. Finally, the COVID-19 pandemic revealed the importance of controls for accurate and standardised diagnosis of SARS-CoV-2. Therefore, I made a pair of chimeric A/B standards for SARS-CoV-2 diagnostic RT-PCR testing. Each standard contains multiple target sequences joined in tandem, where targets present in standard A are absent in B, and vice-versa. This enables control cross-validation, unambiguously distinguishing control and test failures. In summary, the rapid development of genome technologies often results in a diverse, but fragmented landscape in genomics that hinders data compatibility and inter-operability. To bridge that gap my research provides a standardised quantitative reference, more diverse representations of challenging human genome features and alternative design principles for diagnostic test controls.","abstract_html":"Despite decades from the publication of the first draft, the reference human genome remains incomplete, with many unsolved difficult regions. Next-generation sequencing (NGS) has become a central tool for the detection of genetic variation in biological research and clinical diagnosis. However, despite its advantages, NGS suffers from errors and biases that can confound the detection of true variants. Reference standards are control materials, with known properties, against which to test performance. In recent years, the decreasing costs of DNA synthesis have enabled the creative design of synthetic reference controls for genomics. This thesis describes the development of synthetic DNA controls to qualitatively and quantitatively analyse genome features. First, a synthetic ladder, consisting of a single DNA molecule with artificial sequence elements at known copy-numbers, to accurately measure sequence abundance in NGS libraries. The synthetic ladder provides a universal reference to quantify and mitigate the impact of technical variation, independent of the human genome, improving quantitative comparisons within and between samples. Secondly, a synthetic chromosome containing mirrored representations of diverse clinically relevant variants and human genome features, such as HLA alleles and immune receptors. The synthetic chromosome provides a ground-truth reference, with unambiguous representation of low-confidence and difficult-to-sequence genome regions. Therefore, I used the synthetic chromosome to benchmark different experimental and analytical pipelines, highlighting weaknesses and strengths, and providing best-practices guidelines. Finally, the COVID-19 pandemic revealed the importance of controls for accurate and standardised diagnosis of SARS-CoV-2. Therefore, I made a pair of chimeric A/B standards for SARS-CoV-2 diagnostic RT-PCR testing. Each standard contains multiple target sequences joined in tandem, where targets present in standard A are absent in B, and vice-versa. This enables control cross-validation, unambiguously distinguishing control and test failures. In summary, the rapid development of genome technologies often results in a diverse, but fragmented landscape in genomics that hinders data compatibility and inter-operability. To bridge that gap my research provides a standardised quantitative reference, more diverse representations of challenging human genome features and alternative design principles for diagnostic test controls.","abstract_has_math":false,"creators":["Martins Reis, Andre Luiz"],"institution":"UNSW, Sydney","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2021,"date_issued":"2021","date_published":"2021","updated_at":"2026-07-24T05:33:43Z","subjects":["Reference standards","Genomics"],"languages":["EN"],"rights":["open access","CC BY-NC-ND 3.0","free_to_read"],"rights_urls":["https://purl.org/coar/access_right/c_abf2","https://creativecommons.org/licenses/by-nc-nd/3.0/au/"],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["https://doi.org/10.26190/unsworks/2287"],"render_values":[{"text":"https://doi.org/10.26190/unsworks/2287","href":"https://doi.org/10.26190/unsworks/2287","code":true}]}]},"links":{"outbound_url":"http://hdl.handle.net/1959.4/70882","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["Martins Reis, Andre Luiz"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2021"]},{"key":"dc:publisher","label":"Institution","values":["UNSW, Sydney"]},{"key":"dc:type","label":"Dc Type","values":["doctoral thesis","http://purl.org/coar/resource_type/c_db06"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Reference standards","Genomics"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["EN"]},{"key":"dc:rights","label":"Dc Rights","values":["open access","https://purl.org/coar/access_right/c_abf2","CC BY-NC-ND 3.0","https://creativecommons.org/licenses/by-nc-nd/3.0/au/","free_to_read"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/1959.4/70882","https://unsworks.unsw.edu.au/bitstreams/bda36bba-432c-4e04-ad3f-e4aa837ea186/download","https://doi.org/10.26190/unsworks/2287"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Despite decades from the publication of the first draft, the reference human genome remains incomplete, with many unsolved difficult regions. Next-generation sequencing (NGS) has become a central tool for the detection of genetic variation in biological research and clinical diagnosis. However, despite its advantages, NGS suffers from errors and biases that can confound the detection of true variants. Reference standards are control materials, with known properties, against which to test performance. In recent years, the decreasing costs of DNA synthesis have enabled the creative design of synthetic reference controls for genomics. This thesis describes the development of synthetic DNA controls to qualitatively and quantitatively analyse genome features. First, a synthetic ladder, consisting of a single DNA molecule with artificial sequence elements at known copy-numbers, to accurately measure sequence abundance in NGS libraries. The synthetic ladder provides a universal reference to quantify and mitigate the impact of technical variation, independent of the human genome, improving quantitative comparisons within and between samples. Secondly, a synthetic chromosome containing mirrored representations of diverse clinically relevant variants and human genome features, such as HLA alleles and immune receptors. The synthetic chromosome provides a ground-truth reference, with unambiguous representation of low-confidence and difficult-to-sequence genome regions. Therefore, I used the synthetic chromosome to benchmark different experimental and analytical pipelines, highlighting weaknesses and strengths, and providing best-practices guidelines. Finally, the COVID-19 pandemic revealed the importance of controls for accurate and standardised diagnosis of SARS-CoV-2. Therefore, I made a pair of chimeric A/B standards for SARS-CoV-2 diagnostic RT-PCR testing. Each standard contains multiple target sequences joined in tandem, where targets present in standard A are absent in B, and vice-versa. This enables control cross-validation, unambiguously distinguishing control and test failures. In summary, the rapid development of genome technologies often results in a diverse, but fragmented landscape in genomics that hinders data compatibility and inter-operability. To bridge that gap my research provides a standardised quantitative reference, more diverse representations of challenging human genome features and alternative design principles for diagnostic test controls."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["The development of reference standards for genomics."]}]}],"canonical_facts":{"dc:creator":["Martins Reis, Andre Luiz"],"dc:date":["2021"],"dc:description":["Despite decades from the publication of the first draft, the reference human genome remains incomplete, with many unsolved difficult regions. Next-generation sequencing (NGS) has become a central tool for the detection of genetic variation in biological research and clinical diagnosis. However, despite its advantages, NGS suffers from errors and biases that can confound the detection of true variants. Reference standards are control materials, with known properties, against which to test performance. In recent years, the decreasing costs of DNA synthesis have enabled the creative design of synthetic reference controls for genomics. This thesis describes the development of synthetic DNA controls to qualitatively and quantitatively analyse genome features. First, a synthetic ladder, consisting of a single DNA molecule with artificial sequence elements at known copy-numbers, to accurately measure sequence abundance in NGS libraries. The synthetic ladder provides a universal reference to quantify and mitigate the impact of technical variation, independent of the human genome, improving quantitative comparisons within and between samples. Secondly, a synthetic chromosome containing mirrored representations of diverse clinically relevant variants and human genome features, such as HLA alleles and immune receptors. The synthetic chromosome provides a ground-truth reference, with unambiguous representation of low-confidence and difficult-to-sequence genome regions. Therefore, I used the synthetic chromosome to benchmark different experimental and analytical pipelines, highlighting weaknesses and strengths, and providing best-practices guidelines. Finally, the COVID-19 pandemic revealed the importance of controls for accurate and standardised diagnosis of SARS-CoV-2. Therefore, I made a pair of chimeric A/B standards for SARS-CoV-2 diagnostic RT-PCR testing. Each standard contains multiple target sequences joined in tandem, where targets present in standard A are absent in B, and vice-versa. This enables control cross-validation, unambiguously distinguishing control and test failures. In summary, the rapid development of genome technologies often results in a diverse, but fragmented landscape in genomics that hinders data compatibility and inter-operability. To bridge that gap my research provides a standardised quantitative reference, more diverse representations of challenging human genome features and alternative design principles for diagnostic test controls."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/1959.4/70882","https://unsworks.unsw.edu.au/bitstreams/bda36bba-432c-4e04-ad3f-e4aa837ea186/download","https://doi.org/10.26190/unsworks/2287"],"dc:language":["EN"],"dc:publisher":["UNSW, Sydney"],"dc:rights":["open access","https://purl.org/coar/access_right/c_abf2","CC BY-NC-ND 3.0","https://creativecommons.org/licenses/by-nc-nd/3.0/au/","free_to_read"],"dc:subject":["Reference standards","Genomics"],"dc:title":["The development of reference standards for genomics."],"dc:type":["doctoral thesis","http://purl.org/coar/resource_type/c_db06"]},"updated_at":"2026-07-24T05:33:43Z"}