{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/49497"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/49497","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Big data storage workload characterization, modeling and synthetic generation","abstract":"A huge increase in data storage and processing requirements has lead to Big Data, for which next generation storage systems are being designed and implemented. As Big Data stresses the storage layer in new ways, a better understanding of these workloads and the availability of flexible workload generators are increasingly important to facilitate the proper design and performance tuning of storage subsystems like data replication, metadata management, and caching. Our hypothesis is that the autonomic modeling of Big Data storage system workloads through a combination of measurement, and statistical and machine learning techniques is feasible, novel, and useful. We consider the case of one common type of Big Data storage cluster: A cluster dedicated to supporting a mix of MapReduce jobs. We analyze 6-month traces from two large clusters at Yahoo and identify interesting properties of the workloads. We present a novel model for capturing popularity and short-term temporal correlations in object request streams, and show how unsupervised statistical clustering can be used to enable autonomic type-aware workload generation that is suitable for emerging workloads. We extend this model to include other relevant properties of storage systems (file creation and deletion, pre-existing namespaces and hierarchical namespaces) and use the extended model to implement MimesisBench, a realistic namespace metadata benchmark for next-generation storage systems. Finally, we demonstrate the usefulness of MimesisBench through a study of the scalability and performance of the Hadoop Distributed File System name node.","abstract_html":"A huge increase in data storage and processing requirements has lead to Big Data, for which next generation storage systems are being designed and implemented. As Big Data stresses the storage layer in new ways, a better understanding of these workloads and the availability of flexible workload generators are increasingly important to facilitate the proper design and performance tuning of storage subsystems like data replication, metadata management, and caching. Our hypothesis is that the autonomic modeling of Big Data storage system workloads through a combination of measurement, and statistical and machine learning techniques is feasible, novel, and useful. We consider the case of one common type of Big Data storage cluster: A cluster dedicated to supporting a mix of MapReduce jobs. We analyze 6-month traces from two large clusters at Yahoo and identify interesting properties of the workloads. We present a novel model for capturing popularity and short-term temporal correlations in object request streams, and show how unsupervised statistical clustering can be used to enable autonomic type-aware workload generation that is suitable for emerging workloads. We extend this model to include other relevant properties of storage systems (file creation and deletion, pre-existing namespaces and hierarchical namespaces) and use the extended model to implement MimesisBench, a realistic namespace metadata benchmark for next-generation storage systems. Finally, we demonstrate the usefulness of MimesisBench through a study of the scalability and performance of the Hadoop Distributed File System name node.","abstract_has_math":false,"creators":["Abad, Cristina L."],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Campbell, Roy H.","Nahrstedt, Klara","Gupta, Indranil","Lu, Yi","Cherkasova, Ludmila"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2014,"date_issued":"2014-05-30T16:47:05Z","date_published":"2014-05-30T16:47:05Z","updated_at":"2026-07-22T22:25:38Z","subjects":["Big Data","Hadoop","MapReduce","workload","Mimesis","MimesisBench","Hadoop Distributed File System (HDFS)","storage","locality","popularity"],"languages":["en"],"rights":["Copyright 2014 Cristina Abad"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/49497","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Campbell, Roy H.","Nahrstedt, Klara","Gupta, Indranil","Lu, Yi","Cherkasova, Ludmila"]},{"key":"dc:creator","label":"Author","values":["Abad, Cristina L."]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2014-05-30T16:47:05Z","2014-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Big Data","Hadoop","MapReduce","workload","Mimesis","MimesisBench","Hadoop Distributed File System (HDFS)","storage","locality","popularity"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2014 Cristina Abad"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/49497"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["A huge increase in data storage and processing requirements has lead to Big Data, for which next generation storage systems are being designed and implemented. As Big Data stresses the storage layer in new ways, a better understanding of these workloads and the availability of flexible workload generators are increasingly important to facilitate the proper design and performance tuning of storage subsystems like data replication, metadata management, and caching. Our hypothesis is that the autonomic modeling of Big Data storage system workloads through a combination of measurement, and statistical and machine learning techniques is feasible, novel, and useful. We consider the case of one common type of Big Data storage cluster: A cluster dedicated to supporting a mix of MapReduce jobs. We analyze 6-month traces from two large clusters at Yahoo and identify interesting properties of the workloads. We present a novel model for capturing popularity and short-term temporal correlations in object request streams, and show how unsupervised statistical clustering can be used to enable autonomic type-aware workload generation that is suitable for emerging workloads. We extend this model to include other relevant properties of storage systems (file creation and deletion, pre-existing namespaces and hierarchical namespaces) and use the extended model to implement MimesisBench, a realistic namespace metadata benchmark for next-generation storage systems. Finally, we demonstrate the usefulness of MimesisBench through a study of the scalability and performance of the Hadoop Distributed File System name node.","Item withdrawn by Laura Spradlin (lspradl2@illinois.edu) on 2014-02-24T22:51:25Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Abad_Cristina.pdf: 3312494 bytes, checksum: 055576c3dd33a17c5c7c71ad8c671159 (MD5)","Made available in DSpace on 2014-05-30T16:47:05Z (GMT). No. of bitstreams: 2 Cristina_Abad.pdf: 3312494 bytes, checksum: 055576c3dd33a17c5c7c71ad8c671159 (MD5) license.txt: 4060 bytes, checksum: 12069ec5062dd6f80a09563fe1a20411 (MD5)","Updated contributor Cristina Abad to Cristina L. Abad. Metadata cleaned by kappleg2@illinois.edu 2015-5-4."]},{"key":"dc:title","label":"Title","values":["Big data storage workload characterization, modeling and synthetic generation"]}]}],"canonical_facts":{"dc:contributor":["Campbell, Roy H.","Nahrstedt, Klara","Gupta, Indranil","Lu, Yi","Cherkasova, Ludmila"],"dc:creator":["Abad, Cristina L."],"dc:date":["2014-05-30T16:47:05Z","2014-05"],"dc:description":["A huge increase in data storage and processing requirements has lead to Big Data, for which next generation storage systems are being designed and implemented. As Big Data stresses the storage layer in new ways, a better understanding of these workloads and the availability of flexible workload generators are increasingly important to facilitate the proper design and performance tuning of storage subsystems like data replication, metadata management, and caching. Our hypothesis is that the autonomic modeling of Big Data storage system workloads through a combination of measurement, and statistical and machine learning techniques is feasible, novel, and useful. We consider the case of one common type of Big Data storage cluster: A cluster dedicated to supporting a mix of MapReduce jobs. We analyze 6-month traces from two large clusters at Yahoo and identify interesting properties of the workloads. We present a novel model for capturing popularity and short-term temporal correlations in object request streams, and show how unsupervised statistical clustering can be used to enable autonomic type-aware workload generation that is suitable for emerging workloads. We extend this model to include other relevant properties of storage systems (file creation and deletion, pre-existing namespaces and hierarchical namespaces) and use the extended model to implement MimesisBench, a realistic namespace metadata benchmark for next-generation storage systems. Finally, we demonstrate the usefulness of MimesisBench through a study of the scalability and performance of the Hadoop Distributed File System name node.","Item withdrawn by Laura Spradlin (lspradl2@illinois.edu) on 2014-02-24T22:51:25Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Abad_Cristina.pdf: 3312494 bytes, checksum: 055576c3dd33a17c5c7c71ad8c671159 (MD5)","Made available in DSpace on 2014-05-30T16:47:05Z (GMT). No. of bitstreams: 2 Cristina_Abad.pdf: 3312494 bytes, checksum: 055576c3dd33a17c5c7c71ad8c671159 (MD5) license.txt: 4060 bytes, checksum: 12069ec5062dd6f80a09563fe1a20411 (MD5)","Updated contributor Cristina Abad to Cristina L. Abad. Metadata cleaned by kappleg2@illinois.edu 2015-5-4."],"dc:identifier":["http://hdl.handle.net/2142/49497"],"dc:language":["en"],"dc:rights":["Copyright 2014 Cristina Abad"],"dc:subject":["Big Data","Hadoop","MapReduce","workload","Mimesis","MimesisBench","Hadoop Distributed File System (HDFS)","storage","locality","popularity"],"dc:title":["Big data storage workload characterization, modeling and synthetic generation"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:38Z"}