{"id":{"repo_id":"colo-mines","oai_identifier":"oai:repository.mines.edu:11124/178455"},"canonical_url":"https://search.dev.ndltd.org/etd/colo-mines/oai:repository.mines.edu:11124/178455","repository":{"repo_id":"colo-mines","name":"Colorado School of Mines","base_url":"https://repository.mines.edu/server/oai/request"},"display":{"title":"Novel topography-based classification for mountain basins: utility and limitations of unsupervised machine learning, A","abstract":"Unsupervised machine learning algorithms are commonly used data analytic methods with applications spanning many disciplines. In the field of hydrology, K-means and similar clustering methods have been shown to be useful discerning differences in hydrologic signatures between catchments. Specifically, these approaches apply a clustering algorithm to hydrologic data, including hydroclimatic data, land cover data, and topographic data. Topographic attributes alone have also been clustered to establish distinct distributions of mountain ranges to understand species response to climate change and habitat availability, but similar approaches haven’t been applied to hydrology to help understand categories of hydrologic response to climate change. Here, I show that calculated distributions of elevation, slope, and aspect for 604 delineated catchments in the USGS GAGES II dataset along with elevation data originating from the USGS National Elevation Dataset can be used to partition catchments into two clusters. Two clusters were determined to be optimal for this dataset according to the elbow method and average silhouette method. The notion that two clusters is optimal suggests there is minimal differences between clusters when simplifying the topographies of catchments to only distributions of elevation, slope, and aspect. Modeling experiments were conducted using HEC-HMS to assess hydrologic response of clustered catchments to the same hydroclimatic forcing parameters. A sensitivity analysis of models run for the basins closest to the center of Cluster 1 and Cluster 2 indicated Cluster 1 Center shows more variation in the timing of peak flow while Cluster 2 Center shows a shorter time to peak flow (TP), time to baseflow from peak flow (TB), and a faster recession. A Kolmogorov-Smirnov (KS) test indicated that the summary statistics TP, TB, and R2 are statistically significant and do differentiate Cluster 1 Center and Cluster 2 Center. However, when modeling randomly selected basins in each cluster, the lack of visual difference between hydrographs and a KS test indicated no statistical significance for TP, TB, and slope of recession rate (b). These results together suggest no hydrological difference between Cluster 1 and Cluster 2. The linear regression fit of recession rate against discharge performed well for Cluster 1 Center as indicated by a high R2 (0.86) but performed poorly for Cluster 2 Center and for the randomly selected basins. These results demonstrate pitfalls of the K-means clustering algorithm that must be considered. This study is somewhat unique in that I use an external validation method to determine the extent to which the K-means clustering produced meaningful categories. This external validation method indicated no differences between clusters and that the K-means algorithm did not create meaningful classifications. I propose a validation method is needed such as hydrologic modeling of the clustered catchments to assess the integrity of the K-means results and to ensure classifications are distinct and meaningful. Without proper validation techniques, the results of K-means algorithms and broader unsupervised clustering methods should not be assumed to be correctly partitioned.","abstract_html":"Unsupervised machine learning algorithms are commonly used data analytic methods with applications spanning many disciplines. In the field of hydrology, K-means and similar clustering methods have been shown to be useful discerning differences in hydrologic signatures between catchments. Specifically, these approaches apply a clustering algorithm to hydrologic data, including hydroclimatic data, land cover data, and topographic data. Topographic attributes alone have also been clustered to establish distinct distributions of mountain ranges to understand species response to climate change and habitat availability, but similar approaches haven’t been applied to hydrology to help understand categories of hydrologic response to climate change. Here, I show that calculated distributions of elevation, slope, and aspect for 604 delineated catchments in the USGS GAGES II dataset along with elevation data originating from the USGS National Elevation Dataset can be used to partition catchments into two clusters. Two clusters were determined to be optimal for this dataset according to the elbow method and average silhouette method. The notion that two clusters is optimal suggests there is minimal differences between clusters when simplifying the topographies of catchments to only distributions of elevation, slope, and aspect. Modeling experiments were conducted using HEC-HMS to assess hydrologic response of clustered catchments to the same hydroclimatic forcing parameters. A sensitivity analysis of models run for the basins closest to the center of Cluster 1 and Cluster 2 indicated Cluster 1 Center shows more variation in the timing of peak flow while Cluster 2 Center shows a shorter time to peak flow (TP), time to baseflow from peak flow (TB), and a faster recession. A Kolmogorov-Smirnov (KS) test indicated that the summary statistics TP, TB, and R2 are statistically significant and do differentiate Cluster 1 Center and Cluster 2 Center. However, when modeling randomly selected basins in each cluster, the lack of visual difference between hydrographs and a KS test indicated no statistical significance for TP, TB, and slope of recession rate (b). These results together suggest no hydrological difference between Cluster 1 and Cluster 2. The linear regression fit of recession rate against discharge performed well for Cluster 1 Center as indicated by a high R2 (0.86) but performed poorly for Cluster 2 Center and for the randomly selected basins. These results demonstrate pitfalls of the K-means clustering algorithm that must be considered. This study is somewhat unique in that I use an external validation method to determine the extent to which the K-means clustering produced meaningful categories. This external validation method indicated no differences between clusters and that the K-means algorithm did not create meaningful classifications. I propose a validation method is needed such as hydrologic modeling of the clustered catchments to assess the integrity of the K-means results and to ensure classifications are distinct and meaningful. Without proper validation techniques, the results of K-means algorithms and broader unsupervised clustering methods should not be assumed to be correctly partitioned.","abstract_has_math":false,"creators":["Pfaff, Brian"],"institution":"Colorado School of Mines. Arthur Lakes Library","degree_name":"Master of Science (M.S.)","degree_level":"Masters","degree_discipline":"Geology and Geological Engineering","degree_department":null,"school":null,"contributors":[],"advisors":["Marshall, Adrienne M."],"committee_chairs":[],"committee_members":["Singha, Kamini","Roth, Danica"],"year":2023,"date_issued":"2023","date_published":"2023","updated_at":"2026-07-24T01:43:54Z","subjects":["experimental hydrologic modeling","K-means clustering","surface water hydrology","topography","unsupervised machine learning validation"],"languages":["eng","English"],"rights":["Copyright of the original work is retained by the author."],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["T 9516"],"render_values":[{"text":"T 9516","href":null,"code":true}]}]},"links":{"outbound_url":"https://hdl.handle.net/11124/178455","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Marshall, Adrienne M."]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Singha, Kamini","Roth, Danica"]},{"key":"dc:creator","label":"Author","values":["Pfaff, Brian"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2023-10-23T22:04:36Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2023-10-23T22:04:36Z"]},{"key":"dc:date.issued","label":"Date","values":["2023"]},{"key":"dc:publisher","label":"Institution","values":["Colorado School of Mines. Arthur Lakes Library"]},{"key":"dc:type","label":"Dc Type","values":["Text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Geology and Geological Engineering"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Masters"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science (M.S.)"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Colorado School of Mines"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["experimental hydrologic modeling","K-means clustering","surface water hydrology","topography","unsupervised machine learning validation"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["English"]},{"key":"dc:language.iso","label":"Language (ISO)","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright of the original work is retained by the author."]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["Pfaff_mines_0052N_12586.pdf","T 9516"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/11124/178455"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Includes bibliographical references.","2023 Spring."]},{"key":"dc:description.abstract","label":"Abstract","values":["Unsupervised machine learning algorithms are commonly used data analytic methods with applications spanning many disciplines. In the field of hydrology, K-means and similar clustering methods have been shown to be useful discerning differences in hydrologic signatures between catchments. Specifically, these approaches apply a clustering algorithm to hydrologic data, including hydroclimatic data, land cover data, and topographic data. Topographic attributes alone have also been clustered to establish distinct distributions of mountain ranges to understand species response to climate change and habitat availability, but similar approaches haven’t been applied to hydrology to help understand categories of hydrologic response to climate change. Here, I show that calculated distributions of elevation, slope, and aspect for 604 delineated catchments in the USGS GAGES II dataset along with elevation data originating from the USGS National Elevation Dataset can be used to partition catchments into two clusters. Two clusters were determined to be optimal for this dataset according to the elbow method and average silhouette method. The notion that two clusters is optimal suggests there is minimal differences between clusters when simplifying the topographies of catchments to only distributions of elevation, slope, and aspect. Modeling experiments were conducted using HEC-HMS to assess hydrologic response of clustered catchments to the same hydroclimatic forcing parameters. A sensitivity analysis of models run for the basins closest to the center of Cluster 1 and Cluster 2 indicated Cluster 1 Center shows more variation in the timing of peak flow while Cluster 2 Center shows a shorter time to peak flow (TP), time to baseflow from peak flow (TB), and a faster recession. A Kolmogorov-Smirnov (KS) test indicated that the summary statistics TP, TB, and R2 are statistically significant and do differentiate Cluster 1 Center and Cluster 2 Center. However, when modeling randomly selected basins in each cluster, the lack of visual difference between hydrographs and a KS test indicated no statistical significance for TP, TB, and slope of recession rate (b). These results together suggest no hydrological difference between Cluster 1 and Cluster 2. The linear regression fit of recession rate against discharge performed well for Cluster 1 Center as indicated by a high R2 (0.86) but performed poorly for Cluster 2 Center and for the randomly selected basins. These results demonstrate pitfalls of the K-means clustering algorithm that must be considered. This study is somewhat unique in that I use an external validation method to determine the extent to which the K-means clustering produced meaningful categories. This external validation method indicated no differences between clusters and that the K-means algorithm did not create meaningful classifications. I propose a validation method is needed such as hydrologic modeling of the clustered catchments to assess the integrity of the K-means results and to ensure classifications are distinct and meaningful. Without proper validation techniques, the results of K-means algorithms and broader unsupervised clustering methods should not be assumed to be correctly partitioned."]},{"key":"dc:format.medium","label":"Dc Format Medium","values":["born digital","masters theses"]},{"key":"dc:title","label":"Title","values":["Novel topography-based classification for mountain basins: utility and limitations of unsupervised machine learning, A"]}]}],"canonical_facts":{"dc:contributor.advisor":["Marshall, Adrienne M."],"dc:contributor.committeemember":["Singha, Kamini","Roth, Danica"],"dc:creator":["Pfaff, Brian"],"dc:date.accessioned":["2023-10-23T22:04:36Z"],"dc:date.available":["2023-10-23T22:04:36Z"],"dc:date.issued":["2023"],"dc:description":["Includes bibliographical references.","2023 Spring."],"dc:description.abstract":["Unsupervised machine learning algorithms are commonly used data analytic methods with applications spanning many disciplines. In the field of hydrology, K-means and similar clustering methods have been shown to be useful discerning differences in hydrologic signatures between catchments. Specifically, these approaches apply a clustering algorithm to hydrologic data, including hydroclimatic data, land cover data, and topographic data. Topographic attributes alone have also been clustered to establish distinct distributions of mountain ranges to understand species response to climate change and habitat availability, but similar approaches haven’t been applied to hydrology to help understand categories of hydrologic response to climate change. Here, I show that calculated distributions of elevation, slope, and aspect for 604 delineated catchments in the USGS GAGES II dataset along with elevation data originating from the USGS National Elevation Dataset can be used to partition catchments into two clusters. Two clusters were determined to be optimal for this dataset according to the elbow method and average silhouette method. The notion that two clusters is optimal suggests there is minimal differences between clusters when simplifying the topographies of catchments to only distributions of elevation, slope, and aspect. Modeling experiments were conducted using HEC-HMS to assess hydrologic response of clustered catchments to the same hydroclimatic forcing parameters. A sensitivity analysis of models run for the basins closest to the center of Cluster 1 and Cluster 2 indicated Cluster 1 Center shows more variation in the timing of peak flow while Cluster 2 Center shows a shorter time to peak flow (TP), time to baseflow from peak flow (TB), and a faster recession. A Kolmogorov-Smirnov (KS) test indicated that the summary statistics TP, TB, and R2 are statistically significant and do differentiate Cluster 1 Center and Cluster 2 Center. However, when modeling randomly selected basins in each cluster, the lack of visual difference between hydrographs and a KS test indicated no statistical significance for TP, TB, and slope of recession rate (b). These results together suggest no hydrological difference between Cluster 1 and Cluster 2. The linear regression fit of recession rate against discharge performed well for Cluster 1 Center as indicated by a high R2 (0.86) but performed poorly for Cluster 2 Center and for the randomly selected basins. These results demonstrate pitfalls of the K-means clustering algorithm that must be considered. This study is somewhat unique in that I use an external validation method to determine the extent to which the K-means clustering produced meaningful categories. This external validation method indicated no differences between clusters and that the K-means algorithm did not create meaningful classifications. I propose a validation method is needed such as hydrologic modeling of the clustered catchments to assess the integrity of the K-means results and to ensure classifications are distinct and meaningful. Without proper validation techniques, the results of K-means algorithms and broader unsupervised clustering methods should not be assumed to be correctly partitioned."],"dc:format.medium":["born digital","masters theses"],"dc:identifier":["Pfaff_mines_0052N_12586.pdf","T 9516"],"dc:identifier.uri":["https://hdl.handle.net/11124/178455"],"dc:language":["English"],"dc:language.iso":["eng"],"dc:publisher":["Colorado School of Mines. Arthur Lakes Library"],"dc:rights":["Copyright of the original work is retained by the author."],"dc:subject":["experimental hydrologic modeling","K-means clustering","surface water hydrology","topography","unsupervised machine learning validation"],"dc:title":["Novel topography-based classification for mountain basins: utility and limitations of unsupervised machine learning, A"],"dc:type":["Text"],"thesis:degree_discipline":["Geology and Geological Engineering"],"thesis:degree_level":["Masters"],"thesis:degree_name":["Master of Science (M.S.)"],"thesis:institution_name":["Colorado School of Mines"]},"updated_at":"2026-07-24T01:43:54Z"}