{"id":{"repo_id":"must-thes","oai_identifier":"oai:scholarsmine.mst.edu:doctoral_dissertations-4137"},"canonical_url":"https://search.dev.ndltd.org/etd/must-thes/oai:scholarsmine.mst.edu:doctoral_dissertations-4137","repository":{"repo_id":"must-thes","name":"Missouri University of Science and Technology","base_url":"https://scholarsmine.mst.edu/do/oai/"},"display":{"title":"Attention mechanism in deep neural networks for computer vision tasks","abstract":"<p>“Attention mechanism, which is one of the most important algorithms in the deep Learning community, was initially designed in the natural language processing for enhancing the feature representation of key sentence fragments over the context. In recent years, the attention mechanism has been widely adopted in solving computer vision tasks by guiding deep neural networks (DNNs) to focus on specific image features for better understanding the semantic information of the image. However, the attention mechanism is not only capable of helping DNNs understand semantics, but also useful for the feature fusion, visual cue discovering, and temporal information selection, which are seldom researched. In this study, we take the classic attention mechanism a step further by proposing the <em>Semantic Attention Guidance Unit</em> (SAGU) for multi-level feature fusion to tackle the challenging Biomedical Image Segmentation task. Furthermore, we propose a novel framework that consists of (1) <em>Semantic Attention Unit</em> (SAU), which is an advanced version of SAGU for adaptively bringing high-level semantics to mid-level features, (2) <em>Two-level Spatial Attention Module</em> (TSPAM) for discovering multiple visual cues within the image, and (3) <em>Temporal Attention Module</em> (TAM) for temporal information selection to solve the Videobased Person Re-identification task. To validate our newly proposed attention mechanisms, extensive experiments are conducted on challenging datasets. Our methods obtain competitive performance and outperform state-of-the-art methods. Selective publications are also presented in the Appendix”--Abstract, page iii.</p>","abstract_html":"&lt;p&gt;“Attention mechanism, which is one of the most important algorithms in the deep Learning community, was initially designed in the natural language processing for enhancing the feature representation of key sentence fragments over the context. In recent years, the attention mechanism has been widely adopted in solving computer vision tasks by guiding deep neural networks (DNNs) to focus on specific image features for better understanding the semantic information of the image. However, the attention mechanism is not only capable of helping DNNs understand semantics, but also useful for the feature fusion, visual cue discovering, and temporal information selection, which are seldom researched. In this study, we take the classic attention mechanism a step further by proposing the &lt;em&gt;Semantic Attention Guidance Unit&lt;/em&gt; (SAGU) for multi-level feature fusion to tackle the challenging Biomedical Image Segmentation task. Furthermore, we propose a novel framework that consists of (1) &lt;em&gt;Semantic Attention Unit&lt;/em&gt; (SAU), which is an advanced version of SAGU for adaptively bringing high-level semantics to mid-level features, (2) &lt;em&gt;Two-level Spatial Attention Module&lt;/em&gt; (TSPAM) for discovering multiple visual cues within the image, and (3) &lt;em&gt;Temporal Attention Module&lt;/em&gt; (TAM) for temporal information selection to solve the Videobased Person Re-identification task. To validate our newly proposed attention mechanisms, extensive experiments are conducted on challenging datasets. Our methods obtain competitive performance and outperform state-of-the-art methods. Selective publications are also presented in the Appendix”--Abstract, page iii.&lt;/p&gt;","abstract_has_math":false,"creators":["Li, Haohan"],"institution":"Missouri University of Science and Technology","degree_name":"Ph. D. in Computer Science","degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":null,"date_issued":"","date_published":null,"updated_at":"2026-07-24T03:18:09Z","subjects":["Active Learning","Attention Mechanism","Biomedical Image Analysis","Computer Vision","Deep Learning","Person Re-Identification","Computer Sciences"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://scholarsmine.mst.edu/doctoral_dissertations/3132","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["Li, Haohan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:type","label":"Dc Type","values":["Dissertation - Open Access"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph. D. in Computer Science"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Missouri University of Science and Technology"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Active Learning","Attention Mechanism","Biomedical Image Analysis","Computer Vision","Deep Learning","Person Re-Identification","Computer Sciences"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://scholarsmine.mst.edu/doctoral_dissertations/3132"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>“Attention mechanism, which is one of the most important algorithms in the deep Learning community, was initially designed in the natural language processing for enhancing the feature representation of key sentence fragments over the context. In recent years, the attention mechanism has been widely adopted in solving computer vision tasks by guiding deep neural networks (DNNs) to focus on specific image features for better understanding the semantic information of the image. However, the attention mechanism is not only capable of helping DNNs understand semantics, but also useful for the feature fusion, visual cue discovering, and temporal information selection, which are seldom researched. In this study, we take the classic attention mechanism a step further by proposing the <em>Semantic Attention Guidance Unit</em> (SAGU) for multi-level feature fusion to tackle the challenging Biomedical Image Segmentation task. Furthermore, we propose a novel framework that consists of (1) <em>Semantic Attention Unit</em> (SAU), which is an advanced version of SAGU for adaptively bringing high-level semantics to mid-level features, (2) <em>Two-level Spatial Attention Module</em> (TSPAM) for discovering multiple visual cues within the image, and (3) <em>Temporal Attention Module</em> (TAM) for temporal information selection to solve the Videobased Person Re-identification task. To validate our newly proposed attention mechanisms, extensive experiments are conducted on challenging datasets. Our methods obtain competitive performance and outperform state-of-the-art methods. Selective publications are also presented in the Appendix”--Abstract, page iii.</p>"]},{"key":"dc:title","label":"Title","values":["Attention mechanism in deep neural networks for computer vision tasks"]}]}],"canonical_facts":{"dc:creator":["Li, Haohan"],"dc:description.abstract":["<p>“Attention mechanism, which is one of the most important algorithms in the deep Learning community, was initially designed in the natural language processing for enhancing the feature representation of key sentence fragments over the context. In recent years, the attention mechanism has been widely adopted in solving computer vision tasks by guiding deep neural networks (DNNs) to focus on specific image features for better understanding the semantic information of the image. However, the attention mechanism is not only capable of helping DNNs understand semantics, but also useful for the feature fusion, visual cue discovering, and temporal information selection, which are seldom researched. In this study, we take the classic attention mechanism a step further by proposing the <em>Semantic Attention Guidance Unit</em> (SAGU) for multi-level feature fusion to tackle the challenging Biomedical Image Segmentation task. Furthermore, we propose a novel framework that consists of (1) <em>Semantic Attention Unit</em> (SAU), which is an advanced version of SAGU for adaptively bringing high-level semantics to mid-level features, (2) <em>Two-level Spatial Attention Module</em> (TSPAM) for discovering multiple visual cues within the image, and (3) <em>Temporal Attention Module</em> (TAM) for temporal information selection to solve the Videobased Person Re-identification task. To validate our newly proposed attention mechanisms, extensive experiments are conducted on challenging datasets. Our methods obtain competitive performance and outperform state-of-the-art methods. Selective publications are also presented in the Appendix”--Abstract, page iii.</p>"],"dc:identifier":["https://scholarsmine.mst.edu/doctoral_dissertations/3132"],"dc:subject":["Active Learning","Attention Mechanism","Biomedical Image Analysis","Computer Vision","Deep Learning","Person Re-Identification","Computer Sciences"],"dc:title":["Attention mechanism in deep neural networks for computer vision tasks"],"dc:type":["Dissertation - Open Access"],"thesis:degree_name":["Ph. D. in Computer Science"],"thesis:institution_name":["Missouri University of Science and Technology"]},"updated_at":"2026-07-24T03:18:09Z"}