{"id":{"repo_id":"mississippi","oai_identifier":"oai:egrove.olemiss.edu:etd-1944"},"canonical_url":"https://search.dev.ndltd.org/etd/mississippi/oai:egrove.olemiss.edu:etd-1944","repository":{"repo_id":"mississippi","name":"University of Mississippi","base_url":"https://egrove.olemiss.edu/do/oai/"},"display":{"title":"A Work-Stealing For Dynamic Workload Balancing On CPU-GPU Heterogeneous Computing Platforms","abstract":"<p>Although many general purpose workloads have been accelerated on graphical processing units (gpus) over the last decade, other applications whose runtime behaviors are dynamic and irregular such as ones based on trees and graphs have suffered from serious workload imbalance problem caused by architectural differences between cpu and gpu processors. In this thesis, we propose a work-stealing framework to overcome such problems. Our proposed framework allows cpu and gpu threads to steal tasks from each other as well as within the same device by leveraging fine-grained data sharing and thread communication feature available on modern cpu-gpu heterogeneous systems. The implementation of bfs application on the top of our framework achieves a minimum of 8.5% performance improvement over the one with coarse-grained task partitioning scheme. It also achieves 16% performance improvement on average over its non-stealing counterpart.</p>","abstract_html":"&lt;p&gt;Although many general purpose workloads have been accelerated on graphical processing units (gpus) over the last decade, other applications whose runtime behaviors are dynamic and irregular such as ones based on trees and graphs have suffered from serious workload imbalance problem caused by architectural differences between cpu and gpu processors. In this thesis, we propose a work-stealing framework to overcome such problems. Our proposed framework allows cpu and gpu threads to steal tasks from each other as well as within the same device by leveraging fine-grained data sharing and thread communication feature available on modern cpu-gpu heterogeneous systems. The implementation of bfs application on the top of our framework achieves a minimum of 8.5% performance improvement over the one with coarse-grained task partitioning scheme. It also achieves 16% performance improvement on average over its non-stealing counterpart.&lt;/p&gt;","abstract_has_math":false,"creators":["Gad, Esraa A."],"institution":null,"degree_name":"M.S. in Engineering Science","degree_level":"Thesis","degree_discipline":"Computer and Information Science","degree_department":null,"school":null,"contributors":["Byunghyun Jang","Feng Wang","Philip J. Rhodes"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2017,"date_issued":"2017-01-01T08:00:00Z","date_published":"2017-01-01T08:00:00Z","updated_at":"2026-07-24T03:06:07Z","subjects":["Gpgpu","Graph Algorithms","Heterogeneous Computing","Work-Stealing","Computer Sciences"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://egrove.olemiss.edu/etd/945","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Byunghyun Jang","Feng Wang","Philip J. Rhodes"]},{"key":"dc:creator","label":"Author","values":["Gad, Esraa A."]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2019-06-20T07:00:00Z"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer and Information Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S. in Engineering Science"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Gpgpu","Graph Algorithms","Heterogeneous Computing","Work-Stealing","Computer Sciences"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://egrove.olemiss.edu/etd/945"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>Although many general purpose workloads have been accelerated on graphical processing units (gpus) over the last decade, other applications whose runtime behaviors are dynamic and irregular such as ones based on trees and graphs have suffered from serious workload imbalance problem caused by architectural differences between cpu and gpu processors. In this thesis, we propose a work-stealing framework to overcome such problems. Our proposed framework allows cpu and gpu threads to steal tasks from each other as well as within the same device by leveraging fine-grained data sharing and thread communication feature available on modern cpu-gpu heterogeneous systems. The implementation of bfs application on the top of our framework achieves a minimum of 8.5% performance improvement over the one with coarse-grained task partitioning scheme. It also achieves 16% performance improvement on average over its non-stealing counterpart.</p>"]},{"key":"dc:title","label":"Title","values":["A Work-Stealing For Dynamic Workload Balancing On CPU-GPU Heterogeneous Computing Platforms"]}]}],"canonical_facts":{"dc:contributor":["Byunghyun Jang","Feng Wang","Philip J. Rhodes"],"dc:creator":["Gad, Esraa A."],"dc:date.available":["2019-06-20T07:00:00Z"],"dc:description.abstract":["<p>Although many general purpose workloads have been accelerated on graphical processing units (gpus) over the last decade, other applications whose runtime behaviors are dynamic and irregular such as ones based on trees and graphs have suffered from serious workload imbalance problem caused by architectural differences between cpu and gpu processors. In this thesis, we propose a work-stealing framework to overcome such problems. Our proposed framework allows cpu and gpu threads to steal tasks from each other as well as within the same device by leveraging fine-grained data sharing and thread communication feature available on modern cpu-gpu heterogeneous systems. The implementation of bfs application on the top of our framework achieves a minimum of 8.5% performance improvement over the one with coarse-grained task partitioning scheme. It also achieves 16% performance improvement on average over its non-stealing counterpart.</p>"],"dc:identifier":["https://egrove.olemiss.edu/etd/945"],"dc:subject":["Gpgpu","Graph Algorithms","Heterogeneous Computing","Work-Stealing","Computer Sciences"],"dc:title":["A Work-Stealing For Dynamic Workload Balancing On CPU-GPU Heterogeneous Computing Platforms"],"thesis:degree_discipline":["Computer and Information Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S. in Engineering Science"]},"updated_at":"2026-07-24T03:06:07Z"}