{"id":{"repo_id":"cuny","oai_identifier":"oai:academicworks.cuny.edu:cc_etds_theses-2172"},"canonical_url":"https://search.dev.ndltd.org/etd/cuny/oai:academicworks.cuny.edu:cc_etds_theses-2172","repository":{"repo_id":"cuny","name":"City University of New York - City College","base_url":"https://academicworks.cuny.edu/do/oai/"},"display":{"title":"Analysis of Chemical Elements in Basalts using Mislabeled Data, a Machine Learning Approach","abstract":"<p>Scientists use basalt chemistry to discriminate among different tectonic settings. There are well-known chemical elements used to classify tectonic settings. An exploration of new features is done using Logistic Regression and Random Forest to discover any new elements of interest. The models were used with other tools, such as recursive feature elimination and permutations, to increase reliability. Among the scarcely explored chemical elements are Terbium (Tb), Holmium (Ho), <a href=\"https://en.wikipedia.org/wiki/Samarium\" target=\"_blank\">Samarium</a> (Sm), and <a href=\"https://en.wikipedia.org/wiki/Erbium\" target=\"_blank\">Erbium</a> (Er). The data used for the exploration contained many outliers. Therefore, an ensemble model was created to explore the location and composition of such outliers. The ensemble was tested with synthetic data to measure performance. The synthetic data with the same distribution as the underlying data showed an accuracy of 73%, while other distributions of synthetic data reached up to 98% accuracy.</p>","abstract_html":"&lt;p&gt;Scientists use basalt chemistry to discriminate among different tectonic settings. There are well-known chemical elements used to classify tectonic settings. An exploration of new features is done using Logistic Regression and Random Forest to discover any new elements of interest. The models were used with other tools, such as recursive feature elimination and permutations, to increase reliability. Among the scarcely explored chemical elements are Terbium (Tb), Holmium (Ho), &lt;a href=&quot;https://en.wikipedia.org/wiki/Samarium&quot; target=&quot;_blank&quot;&gt;Samarium&lt;/a&gt; (Sm), and &lt;a href=&quot;https://en.wikipedia.org/wiki/Erbium&quot; target=&quot;_blank&quot;&gt;Erbium&lt;/a&gt; (Er). The data used for the exploration contained many outliers. Therefore, an ensemble model was created to explore the location and composition of such outliers. The ensemble was tested with synthetic data to measure performance. The synthetic data with the same distribution as the underlying data showed an accuracy of 73%, while other distributions of synthetic data reached up to 98% accuracy.&lt;/p&gt;","abstract_has_math":false,"creators":["Vivar, Jenifer"],"institution":null,"degree_name":"Master of Science (M.S.)","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Karin Block","Michael Grossberg"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2023,"date_issued":"2023-01-01T08:00:00Z","date_published":"2023-01-01T08:00:00Z","updated_at":"2026-07-24T01:58:07Z","subjects":["Basalts","Chemical Features","Basalt Chemistry","Random Forest","Machine Learning","Outlier Detection","Ternary Plots","RFE","Model Interpretation","Outlier Ensemble Model","Data Science","Geochemistry","Geology"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://academicworks.cuny.edu/cc_etds_theses/1146","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Karin Block","Michael Grossberg"]},{"key":"dc:creator","label":"Author","values":["Vivar, Jenifer"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2023-05-25T07:00:00Z"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science (M.S.)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Basalts","Chemical Features","Basalt Chemistry","Random Forest","Machine Learning","Outlier Detection","Ternary Plots","RFE","Model Interpretation","Outlier Ensemble Model","Data Science","Geochemistry","Geology"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://academicworks.cuny.edu/cc_etds_theses/1146"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>Scientists use basalt chemistry to discriminate among different tectonic settings. There are well-known chemical elements used to classify tectonic settings. An exploration of new features is done using Logistic Regression and Random Forest to discover any new elements of interest. The models were used with other tools, such as recursive feature elimination and permutations, to increase reliability. Among the scarcely explored chemical elements are Terbium (Tb), Holmium (Ho), <a href=\"https://en.wikipedia.org/wiki/Samarium\" target=\"_blank\">Samarium</a> (Sm), and <a href=\"https://en.wikipedia.org/wiki/Erbium\" target=\"_blank\">Erbium</a> (Er). The data used for the exploration contained many outliers. Therefore, an ensemble model was created to explore the location and composition of such outliers. The ensemble was tested with synthetic data to measure performance. The synthetic data with the same distribution as the underlying data showed an accuracy of 73%, while other distributions of synthetic data reached up to 98% accuracy.</p>"]},{"key":"dc:title","label":"Title","values":["Analysis of Chemical Elements in Basalts using Mislabeled Data, a Machine Learning Approach"]}]}],"canonical_facts":{"dc:contributor":["Karin Block","Michael Grossberg"],"dc:creator":["Vivar, Jenifer"],"dc:date.available":["2023-05-25T07:00:00Z"],"dc:description.abstract":["<p>Scientists use basalt chemistry to discriminate among different tectonic settings. There are well-known chemical elements used to classify tectonic settings. An exploration of new features is done using Logistic Regression and Random Forest to discover any new elements of interest. The models were used with other tools, such as recursive feature elimination and permutations, to increase reliability. Among the scarcely explored chemical elements are Terbium (Tb), Holmium (Ho), <a href=\"https://en.wikipedia.org/wiki/Samarium\" target=\"_blank\">Samarium</a> (Sm), and <a href=\"https://en.wikipedia.org/wiki/Erbium\" target=\"_blank\">Erbium</a> (Er). The data used for the exploration contained many outliers. Therefore, an ensemble model was created to explore the location and composition of such outliers. The ensemble was tested with synthetic data to measure performance. The synthetic data with the same distribution as the underlying data showed an accuracy of 73%, while other distributions of synthetic data reached up to 98% accuracy.</p>"],"dc:identifier":["https://academicworks.cuny.edu/cc_etds_theses/1146"],"dc:subject":["Basalts","Chemical Features","Basalt Chemistry","Random Forest","Machine Learning","Outlier Detection","Ternary Plots","RFE","Model Interpretation","Outlier Ensemble Model","Data Science","Geochemistry","Geology"],"dc:title":["Analysis of Chemical Elements in Basalts using Mislabeled Data, a Machine Learning Approach"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["Master of Science (M.S.)"]},"updated_at":"2026-07-24T01:58:07Z"}