{"id":{"repo_id":"eastern-wash","oai_identifier":"oai:dc.ewu.edu:theses-1382"},"canonical_url":"https://search.dev.ndltd.org/etd/eastern-wash/oai:dc.ewu.edu:theses-1382","repository":{"repo_id":"eastern-wash","name":"Eastern Washington University","base_url":"https://dc.ewu.edu/do/oai/"},"display":{"title":"Dynamically parallel CAMSHIFT: GPU accelerated object tracking in digital video","abstract":"<p>\"The CAMSHIFT algorithm is widely used for tracking dynamically sized and positioned objects in real-time applications. In spite of its extensive study on the platform of sequential CPU, its research on massively parallel Graphical Processing Unit (GPU) platform is quite limited. In this work, we designed and implemented two different parallel algorithms for CAMSHIFT using CUDA. The first design performs calculations on the GPU, but requires iterative data transfers back to the host CPU for condition checking, which bottlenecks the entire program. In the second design, we propose an enhanced parallel reduction-based CAMSHIFT using dynamic parallelism to reduce overhead of data transfers between the CPU and GPU. Test results for a 400 by 400 search window show that the second design is up to five times faster than the first design and nine times faster than a pure CPU implementation. We also investigate the deployment of dynamic parallelism for multiple object tracking using CAMSHIFT\"--Leaf iv.</p>","abstract_html":"&lt;p&gt;&quot;The CAMSHIFT algorithm is widely used for tracking dynamically sized and positioned objects in real-time applications. In spite of its extensive study on the platform of sequential CPU, its research on massively parallel Graphical Processing Unit (GPU) platform is quite limited. In this work, we designed and implemented two different parallel algorithms for CAMSHIFT using CUDA. The first design performs calculations on the GPU, but requires iterative data transfers back to the host CPU for condition checking, which bottlenecks the entire program. In the second design, we propose an enhanced parallel reduction-based CAMSHIFT using dynamic parallelism to reduce overhead of data transfers between the CPU and GPU. Test results for a 400 by 400 search window show that the second design is up to five times faster than the first design and nine times faster than a pure CPU implementation. We also investigate the deployment of dynamic parallelism for multiple object tracking using CAMSHIFT&quot;--Leaf iv.&lt;/p&gt;","abstract_has_math":false,"creators":["Perry, Matthew J."],"institution":null,"degree_name":"Master of Science (MS) in Computer Science","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2016,"date_issued":"2016-01-01T08:00:00Z","date_published":"2016-01-01T08:00:00Z","updated_at":"2026-07-24T02:13:35Z","subjects":["Digital video--Mathematical models","Graphics processing units","Parallel algorithms","Coding theory","CUDA (Computer architecture)","Graphics and Human Computer Interfaces","Theory and Algorithms"],"languages":[],"rights":["Access is available to all users"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://dc.ewu.edu/theses/382","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["Perry, Matthew J."]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science (MS) in Computer Science"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Digital video--Mathematical models","Graphics processing units","Parallel algorithms","Coding theory","CUDA (Computer architecture)","Graphics and Human Computer Interfaces","Theory and Algorithms"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["Access is available to all users"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://dc.ewu.edu/theses/382"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>\"The CAMSHIFT algorithm is widely used for tracking dynamically sized and positioned objects in real-time applications. In spite of its extensive study on the platform of sequential CPU, its research on massively parallel Graphical Processing Unit (GPU) platform is quite limited. In this work, we designed and implemented two different parallel algorithms for CAMSHIFT using CUDA. The first design performs calculations on the GPU, but requires iterative data transfers back to the host CPU for condition checking, which bottlenecks the entire program. In the second design, we propose an enhanced parallel reduction-based CAMSHIFT using dynamic parallelism to reduce overhead of data transfers between the CPU and GPU. Test results for a 400 by 400 search window show that the second design is up to five times faster than the first design and nine times faster than a pure CPU implementation. We also investigate the deployment of dynamic parallelism for multiple object tracking using CAMSHIFT\"--Leaf iv.</p>"]},{"key":"dc:title","label":"Title","values":["Dynamically parallel CAMSHIFT: GPU accelerated object tracking in digital video"]}]}],"canonical_facts":{"dc:creator":["Perry, Matthew J."],"dc:description.abstract":["<p>\"The CAMSHIFT algorithm is widely used for tracking dynamically sized and positioned objects in real-time applications. In spite of its extensive study on the platform of sequential CPU, its research on massively parallel Graphical Processing Unit (GPU) platform is quite limited. In this work, we designed and implemented two different parallel algorithms for CAMSHIFT using CUDA. The first design performs calculations on the GPU, but requires iterative data transfers back to the host CPU for condition checking, which bottlenecks the entire program. In the second design, we propose an enhanced parallel reduction-based CAMSHIFT using dynamic parallelism to reduce overhead of data transfers between the CPU and GPU. Test results for a 400 by 400 search window show that the second design is up to five times faster than the first design and nine times faster than a pure CPU implementation. We also investigate the deployment of dynamic parallelism for multiple object tracking using CAMSHIFT\"--Leaf iv.</p>"],"dc:identifier":["https://dc.ewu.edu/theses/382"],"dc:rights":["Access is available to all users"],"dc:subject":["Digital video--Mathematical models","Graphics processing units","Parallel algorithms","Coding theory","CUDA (Computer architecture)","Graphics and Human Computer Interfaces","Theory and Algorithms"],"dc:title":["Dynamically parallel CAMSHIFT: GPU accelerated object tracking in digital video"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["Master of Science (MS) in Computer Science"]},"updated_at":"2026-07-24T02:13:35Z"}