{"id":{"repo_id":"buffalo","oai_identifier":"oai:ubir.buffalo.edu:10477/79447"},"canonical_url":"https://search.dev.ndltd.org/etd/buffalo/oai:ubir.buffalo.edu:10477/79447","repository":{"repo_id":"buffalo","name":"Buffalo","base_url":"https://ubir.buffalo.edu/oai/request"},"display":{"title":"Efficient and Scalable Metadata Access for Distributed Applications from Edge to the Cloud","abstract":"Ph.D.","abstract_html":"Ph.D.","abstract_has_math":false,"creators":["Zhang, Bing; 0000-0002-5590-0001"],"institution":"State University of New York at Buffalo","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":["Kosar, Tevfik","Computer Science and Engineering"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019-04-04T20:33:19Z","date_published":"2019-04-04T20:33:19Z","updated_at":"2026-07-27T19:05:19Z","subjects":["computer science"],"languages":["eng"],"rights":["Users of works found in University at Buffalo Institutional Repository (UBIR) are responsible for identifying and contacting the copyright owner for permission to reuse. University at Buffalo Libraries do not manage rights for copyright-protected works and cannot assist with permissions.","Copyright retained by author."],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/10477/79447","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Kosar, Tevfik","Computer Science and Engineering"]},{"key":"dc:creator","label":"Author","values":["Zhang, Bing; 0000-0002-5590-0001"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2019-04-04T20:33:19Z","2019","2019-01-31 12:04:16"]},{"key":"dc:publisher","label":"Institution","values":["State University of New York at Buffalo"]},{"key":"dc:type","label":"Dc Type","values":["Text","Dissertation"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["computer science"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Users of works found in University at Buffalo Institutional Repository (UBIR) are responsible for identifying and contacting the copyright owner for permission to reuse. University at Buffalo Libraries do not manage rights for copyright-protected works and cannot assist with permissions.","Copyright retained by author."]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/10477/79447"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Ph.D.","We are witnessing a new era that offers new opportunities to conduct data-intensive scientific research with the help of recent advancements in computational, storage, and network technologies. With the rapid deployment of distributed infrastructures and the collaborations between different organizations, it is feasible and promising to run scientific applications on large-scale geo-distributed infrastructures. In many application domains including environmental and coastal hazard prediction, climate modeling, high-energy physics, astronomy, and genome mapping, the volume of data generated already exceeds petabytes, while the corresponding metadata amounts to terabytes or even more. Even though most studies have been conducted on physical data transferring, there has been little work focusing on remotely accessing and transferring large-scale metadata in wide-area networks. Considering wide-area network latency, the frequency of revalidation of metadata, and rapid growth of Internet of Things (IoT), a novel metadata access and transferring mechanism is demanded and becomes a cornerstone of modern distributed IT infrastructures. In this dissertation, we propose a novel solution for efficient and scalable metadata access for distributed applications across wide-area networks. Our solution combines novel pipelining and concurrent transfer mechanisms with reliability, provides distributed continuum caching and prefetching strategies to sidestep fetching latency, and achieves scalable and high-performance stateless fetch/prefetch services in the Cloud. Besides optimizing the metadata transfer performance, we also study the phenomenon of semantic locality in real trace logs which is not well utilized in metadata access prediction. We implement our predictor based on this observation and compare it with three existing state-of-art prefetch schemes (NEXUS, AMP, FARMER) on Yahoo! Hadoop audit traces. By effectively caching and prefetching metadata based on the access pattern, our continuum caching and prefetching mechanism greatly improves local cache hit rate and reduces the average fetching latency. We replayed approximately 20 million metadata access operations from real audit traces, in which our system achieved 80% accuracy during prefetch prediction and reduced the average fetch latency 50% compared to the state-of-the-art mechanisms."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Efficient and Scalable Metadata Access for Distributed Applications from Edge to the Cloud"]}]}],"canonical_facts":{"dc:contributor":["Kosar, Tevfik","Computer Science and Engineering"],"dc:creator":["Zhang, Bing; 0000-0002-5590-0001"],"dc:date":["2019-04-04T20:33:19Z","2019","2019-01-31 12:04:16"],"dc:description":["Ph.D.","We are witnessing a new era that offers new opportunities to conduct data-intensive scientific research with the help of recent advancements in computational, storage, and network technologies. With the rapid deployment of distributed infrastructures and the collaborations between different organizations, it is feasible and promising to run scientific applications on large-scale geo-distributed infrastructures. In many application domains including environmental and coastal hazard prediction, climate modeling, high-energy physics, astronomy, and genome mapping, the volume of data generated already exceeds petabytes, while the corresponding metadata amounts to terabytes or even more. Even though most studies have been conducted on physical data transferring, there has been little work focusing on remotely accessing and transferring large-scale metadata in wide-area networks. Considering wide-area network latency, the frequency of revalidation of metadata, and rapid growth of Internet of Things (IoT), a novel metadata access and transferring mechanism is demanded and becomes a cornerstone of modern distributed IT infrastructures. In this dissertation, we propose a novel solution for efficient and scalable metadata access for distributed applications across wide-area networks. Our solution combines novel pipelining and concurrent transfer mechanisms with reliability, provides distributed continuum caching and prefetching strategies to sidestep fetching latency, and achieves scalable and high-performance stateless fetch/prefetch services in the Cloud. Besides optimizing the metadata transfer performance, we also study the phenomenon of semantic locality in real trace logs which is not well utilized in metadata access prediction. We implement our predictor based on this observation and compare it with three existing state-of-art prefetch schemes (NEXUS, AMP, FARMER) on Yahoo! Hadoop audit traces. By effectively caching and prefetching metadata based on the access pattern, our continuum caching and prefetching mechanism greatly improves local cache hit rate and reduces the average fetching latency. We replayed approximately 20 million metadata access operations from real audit traces, in which our system achieved 80% accuracy during prefetch prediction and reduced the average fetch latency 50% compared to the state-of-the-art mechanisms."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/10477/79447"],"dc:language":["eng"],"dc:publisher":["State University of New York at Buffalo"],"dc:rights":["Users of works found in University at Buffalo Institutional Repository (UBIR) are responsible for identifying and contacting the copyright owner for permission to reuse. University at Buffalo Libraries do not manage rights for copyright-protected works and cannot assist with permissions.","Copyright retained by author."],"dc:subject":["computer science"],"dc:title":["Efficient and Scalable Metadata Access for Distributed Applications from Edge to the Cloud"],"dc:type":["Text","Dissertation"]},"updated_at":"2026-07-27T19:05:19Z"}