{"id":{"repo_id":"mit","oai_identifier":"oai:dspace.mit.edu:1721.1/104556"},"canonical_url":"https://search.dev.ndltd.org/etd/mit/oai:dspace.mit.edu:1721.1/104556","repository":{"repo_id":"mit","name":"MIT","base_url":"https://dspace.mit.edu/oai/request"},"display":{"title":"Return on investment and library complexity analysis for DNA sequencing","abstract":"Understanding the profiles of information acquisition during DNA sequencing experiments is critical to the design and implementation of large-scale studies in medical and population genetics. One known technical challenge and cost driver in next-generation sequencing data is the occurrence of non-independent observations that are created from sequencing artifacts and duplication events from polymerase chain reaction (PCR). The current study demonstrates improved return on investment (ROI) modeling strategies to better anticipate the impact of non-independent observations in multiple forms of next-generation sequencing data. Here, a physical modeling approach based on Pó1ya urn was evaluated using both multi-point estimation and duplicate set occupancy vectors. The results of this study can be used to reduce sequencing costs by improving aspects of experimental design including sample pooling strategies, top-up events, and termination of non-informative samples.","abstract_html":"Understanding the profiles of information acquisition during DNA sequencing experiments is critical to the design and implementation of large-scale studies in medical and population genetics. One known technical challenge and cost driver in next-generation sequencing data is the occurrence of non-independent observations that are created from sequencing artifacts and duplication events from polymerase chain reaction (PCR). The current study demonstrates improved return on investment (ROI) modeling strategies to better anticipate the impact of non-independent observations in multiple forms of next-generation sequencing data. Here, a physical modeling approach based on Pó1ya urn was evaluated using both multi-point estimation and duplicate set occupancy vectors. The results of this study can be used to reduce sequencing costs by improving aspects of experimental design including sample pooling strategies, top-up events, and termination of non-informative samples.","abstract_has_math":false,"creators":["Hogstrom, Larson J"],"institution":"Massachusetts Institute of Technology","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":"Massachusetts Institute of Technology. Computation for Design and Optimization Program.","school":null,"contributors":[],"advisors":["Yossi Farjoun and Bonnie Berger."],"committee_chairs":[],"committee_members":[],"year":2016,"date_issued":"2016","date_published":"2016","updated_at":"2026-07-22T22:21:26Z","subjects":["Computation for Design and Optimization Program."],"languages":["eng"],"rights":["M.I.T. theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. See provided URL for inquiries about permission."],"rights_urls":["http://dspace.mit.edu/handle/1721.1/7582"],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/1721.1/104556","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Yossi Farjoun and Bonnie Berger."]},{"key":"dc:contributor.department","label":"Department","values":["Massachusetts Institute of Technology. Computation for Design and Optimization Program."]},{"key":"dc:contributor.other","label":"Dc Contributor Other","values":["Massachusetts Institute of Technology. Computation for Design and Optimization Program."]},{"key":"dc:creator","label":"Author","values":["Hogstrom, Larson J"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2016-09-30T19:35:26Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2016-09-30T19:35:26Z"]},{"key":"dc:date.issued","label":"Date","values":["2016"]},{"key":"dc:publisher","label":"Institution","values":["Massachusetts Institute of Technology"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computation for Design and Optimization Program."]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["M.I.T. theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. See provided URL for inquiries about permission."]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://dspace.mit.edu/handle/1721.1/7582"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["http://hdl.handle.net/1721.1/104556"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Thesis: S.M., Massachusetts Institute of Technology, Computation for Design and Optimization Program, 2016.","Cataloged from PDF version of thesis.","Includes bibliographical references (page 49)."]},{"key":"dc:description.abstract","label":"Abstract","values":["Understanding the profiles of information acquisition during DNA sequencing experiments is critical to the design and implementation of large-scale studies in medical and population genetics. One known technical challenge and cost driver in next-generation sequencing data is the occurrence of non-independent observations that are created from sequencing artifacts and duplication events from polymerase chain reaction (PCR). The current study demonstrates improved return on investment (ROI) modeling strategies to better anticipate the impact of non-independent observations in multiple forms of next-generation sequencing data. Here, a physical modeling approach based on Pó1ya urn was evaluated using both multi-point estimation and duplicate set occupancy vectors. The results of this study can be used to reduce sequencing costs by improving aspects of experimental design including sample pooling strategies, top-up events, and termination of non-informative samples."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["S.M."]},{"key":"dc:title","label":"Title","values":["Return on investment and library complexity analysis for DNA sequencing"]}]}],"canonical_facts":{"dc:contributor.advisor":["Yossi Farjoun and Bonnie Berger."],"dc:contributor.department":["Massachusetts Institute of Technology. Computation for Design and Optimization Program."],"dc:contributor.other":["Massachusetts Institute of Technology. Computation for Design and Optimization Program."],"dc:creator":["Hogstrom, Larson J"],"dc:date.accessioned":["2016-09-30T19:35:26Z"],"dc:date.available":["2016-09-30T19:35:26Z"],"dc:date.issued":["2016"],"dc:description":["Thesis: S.M., Massachusetts Institute of Technology, Computation for Design and Optimization Program, 2016.","Cataloged from PDF version of thesis.","Includes bibliographical references (page 49)."],"dc:description.abstract":["Understanding the profiles of information acquisition during DNA sequencing experiments is critical to the design and implementation of large-scale studies in medical and population genetics. One known technical challenge and cost driver in next-generation sequencing data is the occurrence of non-independent observations that are created from sequencing artifacts and duplication events from polymerase chain reaction (PCR). The current study demonstrates improved return on investment (ROI) modeling strategies to better anticipate the impact of non-independent observations in multiple forms of next-generation sequencing data. Here, a physical modeling approach based on Pó1ya urn was evaluated using both multi-point estimation and duplicate set occupancy vectors. The results of this study can be used to reduce sequencing costs by improving aspects of experimental design including sample pooling strategies, top-up events, and termination of non-informative samples."],"dc:description.degree":["S.M."],"dc:identifier.uri":["http://hdl.handle.net/1721.1/104556"],"dc:language.iso":["eng"],"dc:publisher":["Massachusetts Institute of Technology"],"dc:rights":["M.I.T. theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. See provided URL for inquiries about permission."],"dc:rights.uri":["http://dspace.mit.edu/handle/1721.1/7582"],"dc:subject":["Computation for Design and Optimization Program."],"dc:title":["Return on investment and library complexity analysis for DNA sequencing"],"dc:type":["Thesis"]},"updated_at":"2026-07-22T22:21:26Z"}