{"id":{"repo_id":"houston","oai_identifier":"oai:uh-ir.tdl.org:10657/5687"},"canonical_url":"https://search.dev.ndltd.org/etd/houston/oai:uh-ir.tdl.org:10657/5687","repository":{"repo_id":"houston","name":"University of Houston","base_url":"https://uh-ir.tdl.org/server/oai/request"},"display":{"title":"Performance models for parallel applications under failures","abstract":"Due to the growing size of compute clusters, large scale parallel applications increasingly have to deal with hardware malfunctions and other failure scenarios during execution. The overall goal of this research is to get good performance of parallel applications despite failures. This dissertation introduces two mathematical models to improve resilience of parallel applications on two different frameworks. The first one is a mathematical model to minimize job completion time for inter-dependent parallel processes running in a volunteer environment by finding the optimal checkpoint interval. Validation is performed with a sample real world application running on a pool of distributed volunteer nodes. The results shows that the predicted checkpoint interval gives performance closed to optimal checkpoint interval determined empirically after extensive experimentation. The second part of the dissertation evaluates the performance of Hadoop MapReduce applications, with different execution parameters and under different failure scenarios. The dissertation introduces performance models for Hadoop MapReduce applications considering node and process failures. Having a performance model allows to determine optimal settings for some of the parameters, such as split size. Validation of the model is done by running two MapReduce applications with different parameter settings. The results show that different applications require different settings for the same MapReduce parameters and the proposed model can predict the performance very well.","abstract_html":"Due to the growing size of compute clusters, large scale parallel applications increasingly have to deal with hardware malfunctions and other failure scenarios during execution. The overall goal of this research is to get good performance of parallel applications despite failures. This dissertation introduces two mathematical models to improve resilience of parallel applications on two different frameworks. The first one is a mathematical model to minimize job completion time for inter-dependent parallel processes running in a volunteer environment by finding the optimal checkpoint interval. Validation is performed with a sample real world application running on a pool of distributed volunteer nodes. The results shows that the predicted checkpoint interval gives performance closed to optimal checkpoint interval determined empirically after extensive experimentation. The second part of the dissertation evaluates the performance of Hadoop MapReduce applications, with different execution parameters and under different failure scenarios. The dissertation introduces performance models for Hadoop MapReduce applications considering node and process failures. Having a performance model allows to determine optimal settings for some of the parameters, such as split size. Validation of the model is done by running two MapReduce applications with different parameter settings. The results show that different applications require different settings for the same MapReduce parameters and the proposed model can predict the performance very well.","abstract_has_math":false,"creators":["Rahman, Mohammad Tanvir 1983-"],"institution":"University of Houston","degree_name":"Doctor of Philosophy","degree_level":"Doctoral","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":[],"advisors":["Gabriel, Edgar"],"committee_chairs":[],"committee_members":["Subhlok, Jaspal","Pandurangan, Gopal","Cheung, Margaret S."],"year":2017,"date_issued":"2017-12","date_published":"2017-12","updated_at":"2026-07-24T02:32:58Z","subjects":["Performance models","Parallel Execution","Fault tolerance","Volunteer computing","Checkpointing","Replication","Host Selection","MapReduce","Hadoop"],"languages":["eng"],"rights":["The author of this work is the copyright owner. UH Libraries and the Texas Digital Library have their permission to store and provide access to this work. UH Libraries has secured permission to reproduce any and all previously published materials contained in the work. Further transmission, reproduction, or presentation of this work is prohibited except with permission of the author(s)."],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/10657/5687","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Gabriel, Edgar"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Subhlok, Jaspal","Pandurangan, Gopal","Cheung, Margaret S."]},{"key":"dc:creator","label":"Author","values":["Rahman, Mohammad Tanvir 1983-"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2020-01-03T21:50:38Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2020-01-03T21:50:38Z"]},{"key":"dc:date.issued","label":"Date","values":["2017-12"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Doctoral"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctor of Philosophy"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Houston"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Performance models","Parallel Execution","Fault tolerance","Volunteer computing","Checkpointing","Replication","Host Selection","MapReduce","Hadoop"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["The author of this work is the copyright owner. UH Libraries and the Texas Digital Library have their permission to store and provide access to this work. UH Libraries has secured permission to reproduce any and all previously published materials contained in the work. Further transmission, reproduction, or presentation of this work is prohibited except with permission of the author(s)."]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10657/5687"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Due to the growing size of compute clusters, large scale parallel applications increasingly have to deal with hardware malfunctions and other failure scenarios during execution. The overall goal of this research is to get good performance of parallel applications despite failures. This dissertation introduces two mathematical models to improve resilience of parallel applications on two different frameworks. The first one is a mathematical model to minimize job completion time for inter-dependent parallel processes running in a volunteer environment by finding the optimal checkpoint interval. Validation is performed with a sample real world application running on a pool of distributed volunteer nodes. The results shows that the predicted checkpoint interval gives performance closed to optimal checkpoint interval determined empirically after extensive experimentation. The second part of the dissertation evaluates the performance of Hadoop MapReduce applications, with different execution parameters and under different failure scenarios. The dissertation introduces performance models for Hadoop MapReduce applications considering node and process failures. Having a performance model allows to determine optimal settings for some of the parameters, such as split size. Validation of the model is done by running two MapReduce applications with different parameter settings. The results show that different applications require different settings for the same MapReduce parameters and the proposed model can predict the performance very well."]},{"key":"dc:format.mimetype","label":"Dc Format Mimetype","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Performance models for parallel applications under failures"]}]}],"canonical_facts":{"dc:contributor.advisor":["Gabriel, Edgar"],"dc:contributor.committeemember":["Subhlok, Jaspal","Pandurangan, Gopal","Cheung, Margaret S."],"dc:creator":["Rahman, Mohammad Tanvir 1983-"],"dc:date.accessioned":["2020-01-03T21:50:38Z"],"dc:date.available":["2020-01-03T21:50:38Z"],"dc:date.issued":["2017-12"],"dc:description.abstract":["Due to the growing size of compute clusters, large scale parallel applications increasingly have to deal with hardware malfunctions and other failure scenarios during execution. The overall goal of this research is to get good performance of parallel applications despite failures. This dissertation introduces two mathematical models to improve resilience of parallel applications on two different frameworks. The first one is a mathematical model to minimize job completion time for inter-dependent parallel processes running in a volunteer environment by finding the optimal checkpoint interval. Validation is performed with a sample real world application running on a pool of distributed volunteer nodes. The results shows that the predicted checkpoint interval gives performance closed to optimal checkpoint interval determined empirically after extensive experimentation. The second part of the dissertation evaluates the performance of Hadoop MapReduce applications, with different execution parameters and under different failure scenarios. The dissertation introduces performance models for Hadoop MapReduce applications considering node and process failures. Having a performance model allows to determine optimal settings for some of the parameters, such as split size. Validation of the model is done by running two MapReduce applications with different parameter settings. The results show that different applications require different settings for the same MapReduce parameters and the proposed model can predict the performance very well."],"dc:format.mimetype":["application/pdf"],"dc:identifier.uri":["https://hdl.handle.net/10657/5687"],"dc:language.iso":["eng"],"dc:rights":["The author of this work is the copyright owner. UH Libraries and the Texas Digital Library have their permission to store and provide access to this work. UH Libraries has secured permission to reproduce any and all previously published materials contained in the work. Further transmission, reproduction, or presentation of this work is prohibited except with permission of the author(s)."],"dc:subject":["Performance models","Parallel Execution","Fault tolerance","Volunteer computing","Checkpointing","Replication","Host Selection","MapReduce","Hadoop"],"dc:title":["Performance models for parallel applications under failures"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Doctoral"],"thesis:degree_name":["Doctor of Philosophy"],"thesis:institution_name":["University of Houston"]},"updated_at":"2026-07-24T02:32:58Z"}