{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/125729"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/125729","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Towards a foundation model for multi-modal and hyperspectral geospatial data","abstract":"Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2026-08-01","abstract_html":"Submission published under a 24 month embargo labeled &#x27;U of I Access&#x27;, the embargo will last until 2026-08-01","abstract_has_math":false,"creators":["Si, Haozhe"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Zhao, Han"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-07-19","date_published":"2024-07-19","updated_at":"2026-07-22T22:25:02Z","subjects":["Computer Vision","Machine Learning"],"languages":["en","eng"],"rights":["Copyright 2024 Haozhe Si"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/125729","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Zhao, Han"]},{"key":"dc:creator","label":"Author","values":["Si, Haozhe"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2024-07-19","2024-08"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer Vision","Machine Learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2024 Haozhe Si"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/125729"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2026-08-01","The student, Haozhe Si, accepted the attached license on 2024-07-15 at 17:23.","The student, Haozhe Si, submitted this Thesis for approval on 2024-07-15 at 17:27.","This Thesis was approved for publication on 2024-07-19 at 10:23.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21125 on 2025-02-04 at 21:17:12","Geospatial imagery data, such as that collected by different satellite-based sensing systems at different times, holds immense potential for enabling a wide range of high-impact applications. Such potential comes from the rich and contextualized information provided by geospatial imagery across multiple dimensions, channels, and sensing modalities. To unlock the insights from geospatial data, recent work has adapted existing self-supervised learning (SSL) approaches; however, they fall short of tailored training objects and model architectures, leading to inflexibility and computational inefficiencies especially when facing an increasing number of channels and modalities. In light of existing limitations, we introduce a novel framework consisting of three key components: i) a Multi-Modal Masked Autoencoder (MM-MAE) that fuses features from different modalities; ii) a Masked-Channel Reconstruction objective that exploits interchannel relationships in hyperspectral data; and iii) a Spatial-Spectral Vision Transformer (S2ViT), incorporating novel Low-Rank Spatial-Spectral Attention Blocks, which flexibly assigns attention to different dimensions. Experimental results demonstrate that our proposed method surpasses current state-of-the-art multi-modal geospatial foundation models, achieving superior performance with less computation and fewer parameters. The flexibility and extensibility of our framework make it a promising solution for future geospatial data analysis tasks that involve a wide range of modalities and dimensions. Consequently, our pretrained model can be effectively applied to various downstream tasks, such as land-cover classification, land functionality management, and marine debris detection, eventually supporting informed decision-making for sustainable development and environmental conservation."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Towards a foundation model for multi-modal and hyperspectral geospatial data"]}]}],"canonical_facts":{"dc:contributor":["Zhao, Han"],"dc:creator":["Si, Haozhe"],"dc:date":["2024-07-19","2024-08"],"dc:description":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2026-08-01","The student, Haozhe Si, accepted the attached license on 2024-07-15 at 17:23.","The student, Haozhe Si, submitted this Thesis for approval on 2024-07-15 at 17:27.","This Thesis was approved for publication on 2024-07-19 at 10:23.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21125 on 2025-02-04 at 21:17:12","Geospatial imagery data, such as that collected by different satellite-based sensing systems at different times, holds immense potential for enabling a wide range of high-impact applications. Such potential comes from the rich and contextualized information provided by geospatial imagery across multiple dimensions, channels, and sensing modalities. To unlock the insights from geospatial data, recent work has adapted existing self-supervised learning (SSL) approaches; however, they fall short of tailored training objects and model architectures, leading to inflexibility and computational inefficiencies especially when facing an increasing number of channels and modalities. In light of existing limitations, we introduce a novel framework consisting of three key components: i) a Multi-Modal Masked Autoencoder (MM-MAE) that fuses features from different modalities; ii) a Masked-Channel Reconstruction objective that exploits interchannel relationships in hyperspectral data; and iii) a Spatial-Spectral Vision Transformer (S2ViT), incorporating novel Low-Rank Spatial-Spectral Attention Blocks, which flexibly assigns attention to different dimensions. Experimental results demonstrate that our proposed method surpasses current state-of-the-art multi-modal geospatial foundation models, achieving superior performance with less computation and fewer parameters. The flexibility and extensibility of our framework make it a promising solution for future geospatial data analysis tasks that involve a wide range of modalities and dimensions. Consequently, our pretrained model can be effectively applied to various downstream tasks, such as land-cover classification, land functionality management, and marine debris detection, eventually supporting informed decision-making for sustainable development and environmental conservation."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/125729"],"dc:language":["en","eng"],"dc:rights":["Copyright 2024 Haozhe Si"],"dc:subject":["Computer Vision","Machine Learning"],"dc:title":["Towards a foundation model for multi-modal and hyperspectral geospatial data"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:02Z"}