{"id":{"repo_id":"buffalo","oai_identifier":"oai:ubir.buffalo.edu:10477/86730"},"canonical_url":"https://search.dev.ndltd.org/etd/buffalo/oai:ubir.buffalo.edu:10477/86730","repository":{"repo_id":"buffalo","name":"Buffalo","base_url":"https://ubir.buffalo.edu/oai/request"},"display":{"title":"Multi-Sensor Fusion for Fast and Robust Computer Vision Applications","abstract":"Ph.D.","abstract_html":"Ph.D.","abstract_has_math":false,"creators":["Dasari, Radhakrishna"],"institution":"State University of New York at Buffalo","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":["Chen, Chang Wen","Computer Science and Engineering"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-02-21T21:44:30Z","date_published":"2025-02-21T21:44:30Z","updated_at":"2026-07-27T19:05:34Z","subjects":["computer science","computer engineering"],"languages":["eng"],"rights":["Users of works found in University at Buffalo Institutional Repository (UBIR) are responsible for identifying and contacting the copyright owner for permission to reuse. University at Buffalo Libraries do not manage rights for copyright-protected works and cannot assist with permissions.","Copyright retained by author."],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/10477/86730","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Chen, Chang Wen","Computer Science and Engineering"]},{"key":"dc:creator","label":"Author","values":["Dasari, Radhakrishna"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-02-21T21:44:30Z","2020"]},{"key":"dc:publisher","label":"Institution","values":["State University of New York at Buffalo"]},{"key":"dc:type","label":"Dc Type","values":["Text","Dissertation"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["computer science","computer engineering"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Users of works found in University at Buffalo Institutional Repository (UBIR) are responsible for identifying and contacting the copyright owner for permission to reuse. University at Buffalo Libraries do not manage rights for copyright-protected works and cannot assist with permissions.","Copyright retained by author."]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/10477/86730"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Ph.D.","From sensation to perception, human brain processes multi-sensory data by leveraging memory and cognition. Multi-sensory integration is the study of how information from the different sensory modalities may be combined to achieve robustness. Extending the applicability of multi-sensory integration to solve computer vision problems is an interesting area of research. Considering the ubiquity of smartphones, which in our context are multi-sensory vision systems, this research has great level of relevance and impact. It can be applied for fast and robust practical computer vision implementations and help overcome power, memory and computational constraints on mobile platform. Smartphone cameras which were intended to replace dedicated point and shoot cameras are beginning to host sophisticated vision applications. They are equipped with multiple cameras, GPS, inertial sensors - accelerometer, gyroscope and a wide variety of sensors. Hence, it is preferable to use multi-modal data wherever applicable, to enhance the performance of mobile vision applications. Multi-Sensor fusion for vision encompasses all the systems that use more than a single imaging sensor. Stereo and Multi-view vision, Multi-camera arrays, visual-inertial systems, vision with depth sensing are few prominent multi-sensory vision systems. The focus of this thesis is in identifying and solving multiple vision problems in a fast and robust way, which otherwise could not be possible without adding a new sensing system. A thorough research on the capabilities and limitations of multi-sensor fusion in addressing major vision problems is the objective of this work. My thesis can be summarized as four major studies as follows: In the first study, the role of sensor fusion is thoroughly explored in the case of color correction for multi-camera arrays. Maintaining color consistency across panoramic video captured by multi-camera array is a challenging problem. Selection of one of the images as reference for color correction may yield poor results because no individual camera image may represent the color palette of entire scene. We address this issue by capturing a separate low resolution color image sensor with a field of view that encompasses entire scene. The color statistics of the reference image are used to bring each camera array image into a uniform radiance and color palette, which gives a real-time performance for color correction. We then estimate optimal color correction parameters using a joint pairwise optimization using Levenberg-Marquardt algorithm for robustness. In the second study, sensor fusion based robust image alignment for mobile HDR Imaging is proposed. large camera motion poses significant challenge to Mobile High Dynamic Range (HDR) Imaging due to hand-held capture of input images and limited computational resources. Aligning images only by detecting and matching image features is computationally expensive and can also be erratic. A robust multi-sensory method is proposed for aligning exposure bracketed images on mobile cameras. Inertial sensor based camera pose estimate is used to pre-warp images and iteratively align them by minimizing alignment error. Local motion masks are simultaneously estimated, which can be used to eliminate ghosting artifacts in the final HDR image. In the third study, a thorough exploration of the theoretical and practical limitations of the key use cases for multi-sensor fusion on smartphones are explored on a multi-sensor video dataset with seven actions from the same subjects captured in three different camera motion scenarios - minimal, pure rotational and both rotational and translational. Accurate estimate of camera pose change, position, orientation, movement and location context are the areas of focus. Multi-sensor solutions for vision problems hold a lot of promise for computer vision applications implemented on smartphones and wearable cameras. In the fourth and final study, a multi-sensor fusion based human action recognition in mobile videos using MPEG-7 standard CDVS (Compact Descriptors for Visual Search) Features is evaluated. When implemented at the hardware level on mobile phones, CDVS features are available for all the images and video frames captured in near real-time. A mobile action recognition framework is proposed, which leverages the strength of visual-inertial auto-calibration of mobile camera and global-local motion estimation, deep learning based human detection and GPS based context.","**To request an accessible version of the file(s) associated with this item, contact library@buffalo.edu. Please include the item's persistent URL [http://hdl.handle.net/. . .] in your request.**"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Multi-Sensor Fusion for Fast and Robust Computer Vision Applications"]}]}],"canonical_facts":{"dc:contributor":["Chen, Chang Wen","Computer Science and Engineering"],"dc:creator":["Dasari, Radhakrishna"],"dc:date":["2025-02-21T21:44:30Z","2020"],"dc:description":["Ph.D.","From sensation to perception, human brain processes multi-sensory data by leveraging memory and cognition. Multi-sensory integration is the study of how information from the different sensory modalities may be combined to achieve robustness. Extending the applicability of multi-sensory integration to solve computer vision problems is an interesting area of research. Considering the ubiquity of smartphones, which in our context are multi-sensory vision systems, this research has great level of relevance and impact. It can be applied for fast and robust practical computer vision implementations and help overcome power, memory and computational constraints on mobile platform. Smartphone cameras which were intended to replace dedicated point and shoot cameras are beginning to host sophisticated vision applications. They are equipped with multiple cameras, GPS, inertial sensors - accelerometer, gyroscope and a wide variety of sensors. Hence, it is preferable to use multi-modal data wherever applicable, to enhance the performance of mobile vision applications. Multi-Sensor fusion for vision encompasses all the systems that use more than a single imaging sensor. Stereo and Multi-view vision, Multi-camera arrays, visual-inertial systems, vision with depth sensing are few prominent multi-sensory vision systems. The focus of this thesis is in identifying and solving multiple vision problems in a fast and robust way, which otherwise could not be possible without adding a new sensing system. A thorough research on the capabilities and limitations of multi-sensor fusion in addressing major vision problems is the objective of this work. My thesis can be summarized as four major studies as follows: In the first study, the role of sensor fusion is thoroughly explored in the case of color correction for multi-camera arrays. Maintaining color consistency across panoramic video captured by multi-camera array is a challenging problem. Selection of one of the images as reference for color correction may yield poor results because no individual camera image may represent the color palette of entire scene. We address this issue by capturing a separate low resolution color image sensor with a field of view that encompasses entire scene. The color statistics of the reference image are used to bring each camera array image into a uniform radiance and color palette, which gives a real-time performance for color correction. We then estimate optimal color correction parameters using a joint pairwise optimization using Levenberg-Marquardt algorithm for robustness. In the second study, sensor fusion based robust image alignment for mobile HDR Imaging is proposed. large camera motion poses significant challenge to Mobile High Dynamic Range (HDR) Imaging due to hand-held capture of input images and limited computational resources. Aligning images only by detecting and matching image features is computationally expensive and can also be erratic. A robust multi-sensory method is proposed for aligning exposure bracketed images on mobile cameras. Inertial sensor based camera pose estimate is used to pre-warp images and iteratively align them by minimizing alignment error. Local motion masks are simultaneously estimated, which can be used to eliminate ghosting artifacts in the final HDR image. In the third study, a thorough exploration of the theoretical and practical limitations of the key use cases for multi-sensor fusion on smartphones are explored on a multi-sensor video dataset with seven actions from the same subjects captured in three different camera motion scenarios - minimal, pure rotational and both rotational and translational. Accurate estimate of camera pose change, position, orientation, movement and location context are the areas of focus. Multi-sensor solutions for vision problems hold a lot of promise for computer vision applications implemented on smartphones and wearable cameras. In the fourth and final study, a multi-sensor fusion based human action recognition in mobile videos using MPEG-7 standard CDVS (Compact Descriptors for Visual Search) Features is evaluated. When implemented at the hardware level on mobile phones, CDVS features are available for all the images and video frames captured in near real-time. A mobile action recognition framework is proposed, which leverages the strength of visual-inertial auto-calibration of mobile camera and global-local motion estimation, deep learning based human detection and GPS based context.","**To request an accessible version of the file(s) associated with this item, contact library@buffalo.edu. Please include the item's persistent URL [http://hdl.handle.net/. . .] in your request.**"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/10477/86730"],"dc:language":["eng"],"dc:publisher":["State University of New York at Buffalo"],"dc:rights":["Users of works found in University at Buffalo Institutional Repository (UBIR) are responsible for identifying and contacting the copyright owner for permission to reuse. University at Buffalo Libraries do not manage rights for copyright-protected works and cannot assist with permissions.","Copyright retained by author."],"dc:subject":["computer science","computer engineering"],"dc:title":["Multi-Sensor Fusion for Fast and Robust Computer Vision Applications"],"dc:type":["Text","Dissertation"]},"updated_at":"2026-07-27T19:05:34Z"}