{"id":{"repo_id":"columbus-state","oai_identifier":"oai:csuepress.columbusstate.edu:theses_dissertations-1008"},"canonical_url":"https://search.dev.ndltd.org/etd/columbus-state/oai:csuepress.columbusstate.edu:theses_dissertations-1008","repository":{"repo_id":"columbus-state","name":"Columbus State University","base_url":"https://csuepress.columbusstate.edu/do/oai/"},"display":{"title":"Advanced I/O Techniques for Efficient and Highly Available Process Crash Recovery Protocols","abstract":"<p>As the number of CPU cores in high-performance computing platforms continues to grow, the availability and reliability of these systems become a primary concern. As such, some solutions are physical (ie. power backup) and some are software driven. Lawrence Berkeley National Laboratory has created a system-level fault-tolerant checkpoint/restart implementation for Linux Clusters. This allows processes to restart computations at the last known checkpoint in the event the system crashes. The checkpoint data creation is highly dependent on system input and output operations. This paper proposes: (i) a technique to improve the efficiency of these I/O operations and (ii) an alternative checkpoint creation method to increase availability and reliability of checkpointing data.</p>","abstract_html":"&lt;p&gt;As the number of CPU cores in high-performance computing platforms continues to grow, the availability and reliability of these systems become a primary concern. As such, some solutions are physical (ie. power backup) and some are software driven. Lawrence Berkeley National Laboratory has created a system-level fault-tolerant checkpoint/restart implementation for Linux Clusters. This allows processes to restart computations at the last known checkpoint in the event the system crashes. The checkpoint data creation is highly dependent on system input and output operations. This paper proposes: (i) a technique to improve the efficiency of these I/O operations and (ii) an alternative checkpoint creation method to increase availability and reliability of checkpointing data.&lt;/p&gt;","abstract_has_math":false,"creators":["Cornwell, Jason Warren"],"institution":null,"degree_name":"Computer Science - Applied Computing Track","degree_level":"Thesis","degree_discipline":"TSYS School of Computer Science","degree_department":null,"school":null,"contributors":["Dr. Angkul Kongmunvattana","Dr. Eugen Ionascu","Dr. Wayne Summers"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2011,"date_issued":"2011-03-01T08:00:00Z","date_published":"2011-03-01T08:00:00Z","updated_at":"2026-07-24T01:44:48Z","subjects":["Caching","Checkpoint","Crash Recovery","Fault-Tolerant System","High Performance Computing","Network File System","Computer Sciences","OS and Networks","Programming Languages and Compilers"],"languages":["English"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://csuepress.columbusstate.edu/theses_dissertations/8","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Dr. Angkul Kongmunvattana","Dr. Eugen Ionascu","Dr. Wayne Summers"]},{"key":"dc:creator","label":"Author","values":["Cornwell, Jason Warren"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2015-09-29T07:00:00Z"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["TSYS School of Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Computer Science - Applied Computing Track"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Caching","Checkpoint","Crash Recovery","Fault-Tolerant System","High Performance Computing","Network File System","Computer Sciences","OS and Networks","Programming Languages and Compilers"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["English"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://csuepress.columbusstate.edu/theses_dissertations/8"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>As the number of CPU cores in high-performance computing platforms continues to grow, the availability and reliability of these systems become a primary concern. As such, some solutions are physical (ie. power backup) and some are software driven. Lawrence Berkeley National Laboratory has created a system-level fault-tolerant checkpoint/restart implementation for Linux Clusters. This allows processes to restart computations at the last known checkpoint in the event the system crashes. The checkpoint data creation is highly dependent on system input and output operations. This paper proposes: (i) a technique to improve the efficiency of these I/O operations and (ii) an alternative checkpoint creation method to increase availability and reliability of checkpointing data.</p>"]},{"key":"dc:title","label":"Title","values":["Advanced I/O Techniques for Efficient and Highly Available Process Crash Recovery Protocols"]}]}],"canonical_facts":{"dc:contributor":["Dr. Angkul Kongmunvattana","Dr. Eugen Ionascu","Dr. Wayne Summers"],"dc:creator":["Cornwell, Jason Warren"],"dc:date.available":["2015-09-29T07:00:00Z"],"dc:description.abstract":["<p>As the number of CPU cores in high-performance computing platforms continues to grow, the availability and reliability of these systems become a primary concern. As such, some solutions are physical (ie. power backup) and some are software driven. Lawrence Berkeley National Laboratory has created a system-level fault-tolerant checkpoint/restart implementation for Linux Clusters. This allows processes to restart computations at the last known checkpoint in the event the system crashes. The checkpoint data creation is highly dependent on system input and output operations. This paper proposes: (i) a technique to improve the efficiency of these I/O operations and (ii) an alternative checkpoint creation method to increase availability and reliability of checkpointing data.</p>"],"dc:identifier":["https://csuepress.columbusstate.edu/theses_dissertations/8"],"dc:language":["English"],"dc:subject":["Caching","Checkpoint","Crash Recovery","Fault-Tolerant System","High Performance Computing","Network File System","Computer Sciences","OS and Networks","Programming Languages and Compilers"],"dc:title":["Advanced I/O Techniques for Efficient and Highly Available Process Crash Recovery Protocols"],"thesis:degree_discipline":["TSYS School of Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["Computer Science - Applied Computing Track"]},"updated_at":"2026-07-24T01:44:48Z"}