{"id":{"repo_id":"mit","oai_identifier":"oai:dspace.mit.edu:1721.1/33086"},"canonical_url":"https://search.dev.ndltd.org/etd/mit/oai:dspace.mit.edu:1721.1/33086","repository":{"repo_id":"mit","name":"MIT","base_url":"https://dspace.mit.edu/oai/request"},"display":{"title":"Hidden Markov Model inference copy number change in array-CGH data","abstract":"Cancer development and progression typically features genomic instability frequently resulting in genomic changes involving DNA copy number gains or losses. Identifying the genomic location of these regional alterations provides important opportunities for the discovery of potential novel oncogenes and tumor suppressors. Recently, array based competitive genomic hybridization (array-CGH) has become available as a powerful approach for genome-wide detection of DNA copy number changes. Array-CGH assesses DNA copy number in tumor samples through competitive hybridization on microarrays containing probes for thousands of genes. The datasets generated are complex and require statistical methods to accurately define discrete and uniform copy number from the data and to identify transitions between genomic regions with altered copy number. Several approaches based on different statistical frameworks have been developed. However, a fundamental informatic issue in array-CGH analysis remains unsolved by these methods. In particular, sample-specific data compression, a result of tumor cells being commonly admixed with normal cells in many tumor types, must be accounted for in each sample analyzed. Additionally, in order to accurately assess deviations from normal copy number, the copy number readout must be shifted to faithfully represent the baseline copy number in each tumor sample. Failure to appropriately address these issues reduces the accuracy of the data in hard-threshold based high-level analysis.","abstract_html":"Cancer development and progression typically features genomic instability frequently resulting in genomic changes involving DNA copy number gains or losses. Identifying the genomic location of these regional alterations provides important opportunities for the discovery of potential novel oncogenes and tumor suppressors. Recently, array based competitive genomic hybridization (array-CGH) has become available as a powerful approach for genome-wide detection of DNA copy number changes. Array-CGH assesses DNA copy number in tumor samples through competitive hybridization on microarrays containing probes for thousands of genes. The datasets generated are complex and require statistical methods to accurately define discrete and uniform copy number from the data and to identify transitions between genomic regions with altered copy number. Several approaches based on different statistical frameworks have been developed. However, a fundamental informatic issue in array-CGH analysis remains unsolved by these methods. In particular, sample-specific data compression, a result of tumor cells being commonly admixed with normal cells in many tumor types, must be accounted for in each sample analyzed. Additionally, in order to accurately assess deviations from normal copy number, the copy number readout must be shifted to faithfully represent the baseline copy number in each tumor sample. Failure to appropriately address these issues reduces the accuracy of the data in hard-threshold based high-level analysis.","abstract_has_math":false,"creators":["Zhang, Yunyu"],"institution":"Massachusetts Institute of Technology","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":"Harvard University--MIT Division of Health Sciences and Technology.","school":null,"contributors":[],"advisors":["Lynda Chin, Cheng Li and Cameron W. Brennan."],"committee_chairs":[],"committee_members":[],"year":2005,"date_issued":"2005","date_published":"2005","updated_at":"2026-07-22T22:22:24Z","subjects":["Harvard University--MIT Division of Health Sciences and Technology."],"languages":["eng"],"rights":["M.I.T. theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. See provided URL for inquiries about permission."],"rights_urls":["http://dspace.mit.edu/handle/1721.1/7582"],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/1721.1/33086","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Lynda Chin, Cheng Li and Cameron W. Brennan."]},{"key":"dc:contributor.department","label":"Department","values":["Harvard University--MIT Division of Health Sciences and Technology."]},{"key":"dc:contributor.other","label":"Dc Contributor Other","values":["Harvard University--MIT Division of Health Sciences and Technology."]},{"key":"dc:creator","label":"Author","values":["Zhang, Yunyu"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2006-06-19T17:39:13Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2006-06-19T17:39:13Z"]},{"key":"dc:date.issued","label":"Date","values":["2005"]},{"key":"dc:publisher","label":"Institution","values":["Massachusetts Institute of Technology"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Harvard University--MIT Division of Health Sciences and Technology."]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["M.I.T. theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. See provided URL for inquiries about permission."]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://dspace.mit.edu/handle/1721.1/7582"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["http://hdl.handle.net/1721.1/33086"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Thesis (S.M.)--Harvard-MIT Division of Health Sciences and Technology, 2005.","Includes bibliographical references (p. 56-57)."]},{"key":"dc:description.abstract","label":"Abstract","values":["Cancer development and progression typically features genomic instability frequently resulting in genomic changes involving DNA copy number gains or losses. Identifying the genomic location of these regional alterations provides important opportunities for the discovery of potential novel oncogenes and tumor suppressors. Recently, array based competitive genomic hybridization (array-CGH) has become available as a powerful approach for genome-wide detection of DNA copy number changes. Array-CGH assesses DNA copy number in tumor samples through competitive hybridization on microarrays containing probes for thousands of genes. The datasets generated are complex and require statistical methods to accurately define discrete and uniform copy number from the data and to identify transitions between genomic regions with altered copy number. Several approaches based on different statistical frameworks have been developed. However, a fundamental informatic issue in array-CGH analysis remains unsolved by these methods. In particular, sample-specific data compression, a result of tumor cells being commonly admixed with normal cells in many tumor types, must be accounted for in each sample analyzed. Additionally, in order to accurately assess deviations from normal copy number, the copy number readout must be shifted to faithfully represent the baseline copy number in each tumor sample. Failure to appropriately address these issues reduces the accuracy of the data in hard-threshold based high-level analysis.","(cont.) By using the natural framework Hidden Markov Models (HMM) to model the distribution of array-CGH signals, a method infer the absolute copy number and identify change points has been developed to address the above problems. This method has been validated on independent dataset and its utility in inference on array-CGH data is demonstrated here."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["S.M."]},{"key":"dc:format.mimetype","label":"Dc Format Mimetype","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Hidden Markov Model inference copy number change in array-CGH data"]}]}],"canonical_facts":{"dc:contributor.advisor":["Lynda Chin, Cheng Li and Cameron W. Brennan."],"dc:contributor.department":["Harvard University--MIT Division of Health Sciences and Technology."],"dc:contributor.other":["Harvard University--MIT Division of Health Sciences and Technology."],"dc:creator":["Zhang, Yunyu"],"dc:date.accessioned":["2006-06-19T17:39:13Z"],"dc:date.available":["2006-06-19T17:39:13Z"],"dc:date.issued":["2005"],"dc:description":["Thesis (S.M.)--Harvard-MIT Division of Health Sciences and Technology, 2005.","Includes bibliographical references (p. 56-57)."],"dc:description.abstract":["Cancer development and progression typically features genomic instability frequently resulting in genomic changes involving DNA copy number gains or losses. Identifying the genomic location of these regional alterations provides important opportunities for the discovery of potential novel oncogenes and tumor suppressors. Recently, array based competitive genomic hybridization (array-CGH) has become available as a powerful approach for genome-wide detection of DNA copy number changes. Array-CGH assesses DNA copy number in tumor samples through competitive hybridization on microarrays containing probes for thousands of genes. The datasets generated are complex and require statistical methods to accurately define discrete and uniform copy number from the data and to identify transitions between genomic regions with altered copy number. Several approaches based on different statistical frameworks have been developed. However, a fundamental informatic issue in array-CGH analysis remains unsolved by these methods. In particular, sample-specific data compression, a result of tumor cells being commonly admixed with normal cells in many tumor types, must be accounted for in each sample analyzed. Additionally, in order to accurately assess deviations from normal copy number, the copy number readout must be shifted to faithfully represent the baseline copy number in each tumor sample. Failure to appropriately address these issues reduces the accuracy of the data in hard-threshold based high-level analysis.","(cont.) By using the natural framework Hidden Markov Models (HMM) to model the distribution of array-CGH signals, a method infer the absolute copy number and identify change points has been developed to address the above problems. This method has been validated on independent dataset and its utility in inference on array-CGH data is demonstrated here."],"dc:description.degree":["S.M."],"dc:format.mimetype":["application/pdf"],"dc:identifier.uri":["http://hdl.handle.net/1721.1/33086"],"dc:language.iso":["eng"],"dc:publisher":["Massachusetts Institute of Technology"],"dc:rights":["M.I.T. theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. See provided URL for inquiries about permission."],"dc:rights.uri":["http://dspace.mit.edu/handle/1721.1/7582"],"dc:subject":["Harvard University--MIT Division of Health Sciences and Technology."],"dc:title":["Hidden Markov Model inference copy number change in array-CGH data"],"dc:type":["Thesis"]},"updated_at":"2026-07-22T22:22:24Z"}