{"id":{"repo_id":"vt","oai_identifier":"oai:vtechworks.lib.vt.edu:10919/141131"},"canonical_url":"https://search.dev.ndltd.org/etd/vt/oai:vtechworks.lib.vt.edu:10919/141131","repository":{"repo_id":"vt","name":"Virginia Tech","base_url":"https://vtechworks.lib.vt.edu/oai/request"},"display":{"title":"Machine Learning for Structure-Agnostic Chemical Analysis from Chromatographic Data","abstract":"Environmental monitoring relies heavily on gas chromatography (GC) to measure airborne contaminants such as volatile organic compounds (VOCs), yet many detected compounds lack structural or spectral references, limiting identification, property estimation, and quantitative analysis. This thesis investigates how machine learning (ML) can extract chemically meaningful information directly from chromatographic data to overcome these limitations. First, ML models are developed to establish a bidirectional relationship between chromatographic retention behavior on orthogonal GC phases and key physicochemical properties (vapor pressure, Henry's law constant, and solubility). Using XGBoost regression models trained on the NIST retention index database, a structure-agnostic \"Index-to-Property\" model predicts physicochemical properties from paired retention indices, while a complementary \"Property-to-Index\" model predicts retention behavior from known properties, achieving predictive performance up to R^2=0.98. Second, this work demonstrates that compound identity and concentration can be inferred directly from chromatographic peak shape, bypassing manual peak integration. ML classification and regression models trained on peaks from ambient atmospheric samples achieve 89% identification accuracy and a mean absolute error of 0.085 ppbv in concentration prediction. Together, these results show that machine learning can address key identification and data reduction challenges in environmental GC, enabling faster, structure-independent interpretation of complex mixtures.","abstract_html":"Environmental monitoring relies heavily on gas chromatography (GC) to measure airborne contaminants such as volatile organic compounds (VOCs), yet many detected compounds lack structural or spectral references, limiting identification, property estimation, and quantitative analysis. This thesis investigates how machine learning (ML) can extract chemically meaningful information directly from chromatographic data to overcome these limitations. First, ML models are developed to establish a bidirectional relationship between chromatographic retention behavior on orthogonal GC phases and key physicochemical properties (vapor pressure, Henry&#x27;s law constant, and solubility). Using XGBoost regression models trained on the NIST retention index database, a structure-agnostic &quot;Index-to-Property&quot; model predicts physicochemical properties from paired retention indices, while a complementary &quot;Property-to-Index&quot; model predicts retention behavior from known properties, achieving predictive performance up to R^2=0.98. Second, this work demonstrates that compound identity and concentration can be inferred directly from chromatographic peak shape, bypassing manual peak integration. ML classification and regression models trained on peaks from ambient atmospheric samples achieve 89% identification accuracy and a mean absolute error of 0.085 ppbv in concentration prediction. Together, these results show that machine learning can address key identification and data reduction challenges in environmental GC, enabling faster, structure-independent interpretation of complex mixtures.","abstract_has_math":false,"creators":["Lahouar, Adam"],"institution":"Virginia Tech","degree_name":"Master of Science","degree_level":"masters","degree_discipline":"Computer Science & Applications","degree_department":"Computer Science and#38; Applications","school":null,"contributors":[],"advisors":[],"committee_chairs":["Eldardiry, Hoda Mohamed"],"committee_members":["Isaacman-VanWertz, Gabriel","Yanardag Delul, Pinar"],"year":2026,"date_issued":"2026-02-03","date_published":"2026-02-03","updated_at":"2026-07-22T22:19:10Z","subjects":["Gas Chromatography","Machine Learning","Compound Classification"],"languages":["en"],"rights":["In Copyright"],"rights_urls":["http://rightsstatements.org/vocab/InC/1.0/"],"identifier_entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:45632"],"render_values":[{"text":"vt_gsexam:45632","href":null,"code":true}]}]},"links":{"outbound_url":"https://hdl.handle.net/10919/141131","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.committeechair","label":"Committee Chair","values":["Eldardiry, Hoda Mohamed"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Isaacman-VanWertz, Gabriel","Yanardag Delul, Pinar"]},{"key":"dc:contributor.department","label":"Department","values":["Computer Science and#38; Applications"]},{"key":"dc:creator","label":"Author","values":["Lahouar, Adam"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2026-02-04T09:00:30Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2026-02-04T09:00:30Z"]},{"key":"dc:date.issued","label":"Date","values":["2026-02-03"]},{"key":"dc:publisher","label":"Institution","values":["Virginia Tech"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science & Applications"]},{"key":"thesis:degree_level","label":"Degree Level","values":["masters"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Virginia Polytechnic Institute and State University"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Gas Chromatography","Machine Learning","Compound Classification"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["In Copyright"]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://rightsstatements.org/vocab/InC/1.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:45632"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10919/141131"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Environmental monitoring relies heavily on gas chromatography (GC) to measure airborne contaminants such as volatile organic compounds (VOCs), yet many detected compounds lack structural or spectral references, limiting identification, property estimation, and quantitative analysis. This thesis investigates how machine learning (ML) can extract chemically meaningful information directly from chromatographic data to overcome these limitations. First, ML models are developed to establish a bidirectional relationship between chromatographic retention behavior on orthogonal GC phases and key physicochemical properties (vapor pressure, Henry's law constant, and solubility). Using XGBoost regression models trained on the NIST retention index database, a structure-agnostic \"Index-to-Property\" model predicts physicochemical properties from paired retention indices, while a complementary \"Property-to-Index\" model predicts retention behavior from known properties, achieving predictive performance up to R^2=0.98. Second, this work demonstrates that compound identity and concentration can be inferred directly from chromatographic peak shape, bypassing manual peak integration. ML classification and regression models trained on peaks from ambient atmospheric samples achieve 89% identification accuracy and a mean absolute error of 0.085 ppbv in concentration prediction. Together, these results show that machine learning can address key identification and data reduction challenges in environmental GC, enabling faster, structure-independent interpretation of complex mixtures."]},{"key":"dc:description.abstractgeneral","label":"General Abstract","values":["Gas chromatography is an important method for monitoring air pollution, but many detected chemicals cannot be fully identified because reference information is missing or incomplete. This makes it difficult to understand what these compounds are, how they behave in the environment, and how much of them are present. This thesis explores how machine learning can help extract useful chemical information directly from chromatographic data. First, machine learning is used to relate chemical behavior in a gas chromatograph to important physical properties, allowing unknown compounds to be characterized without knowing their chemical structures. Second, machine learning is used to analyze the shape of chromatographic signals to identify compounds and estimate their concentrations automatically, reducing the need for time-consuming manual data processing. Overall, this research shows how machine learning can expand the capabilities of gas chromatography for environmental monitoring, improving both the speed and depth of chemical analysis over traditional methods."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["Master of Science"]},{"key":"dc:format.medium","label":"Dc Format Medium","values":["ETD"]},{"key":"dc:title","label":"Title","values":["Machine Learning for Structure-Agnostic Chemical Analysis from Chromatographic Data"]}]}],"canonical_facts":{"dc:contributor.committeechair":["Eldardiry, Hoda Mohamed"],"dc:contributor.committeemember":["Isaacman-VanWertz, Gabriel","Yanardag Delul, Pinar"],"dc:contributor.department":["Computer Science and#38; Applications"],"dc:creator":["Lahouar, Adam"],"dc:date.accessioned":["2026-02-04T09:00:30Z"],"dc:date.available":["2026-02-04T09:00:30Z"],"dc:date.issued":["2026-02-03"],"dc:description.abstract":["Environmental monitoring relies heavily on gas chromatography (GC) to measure airborne contaminants such as volatile organic compounds (VOCs), yet many detected compounds lack structural or spectral references, limiting identification, property estimation, and quantitative analysis. This thesis investigates how machine learning (ML) can extract chemically meaningful information directly from chromatographic data to overcome these limitations. First, ML models are developed to establish a bidirectional relationship between chromatographic retention behavior on orthogonal GC phases and key physicochemical properties (vapor pressure, Henry's law constant, and solubility). Using XGBoost regression models trained on the NIST retention index database, a structure-agnostic \"Index-to-Property\" model predicts physicochemical properties from paired retention indices, while a complementary \"Property-to-Index\" model predicts retention behavior from known properties, achieving predictive performance up to R^2=0.98. Second, this work demonstrates that compound identity and concentration can be inferred directly from chromatographic peak shape, bypassing manual peak integration. ML classification and regression models trained on peaks from ambient atmospheric samples achieve 89% identification accuracy and a mean absolute error of 0.085 ppbv in concentration prediction. Together, these results show that machine learning can address key identification and data reduction challenges in environmental GC, enabling faster, structure-independent interpretation of complex mixtures."],"dc:description.abstractgeneral":["Gas chromatography is an important method for monitoring air pollution, but many detected chemicals cannot be fully identified because reference information is missing or incomplete. This makes it difficult to understand what these compounds are, how they behave in the environment, and how much of them are present. This thesis explores how machine learning can help extract useful chemical information directly from chromatographic data. First, machine learning is used to relate chemical behavior in a gas chromatograph to important physical properties, allowing unknown compounds to be characterized without knowing their chemical structures. Second, machine learning is used to analyze the shape of chromatographic signals to identify compounds and estimate their concentrations automatically, reducing the need for time-consuming manual data processing. Overall, this research shows how machine learning can expand the capabilities of gas chromatography for environmental monitoring, improving both the speed and depth of chemical analysis over traditional methods."],"dc:description.degree":["Master of Science"],"dc:format.medium":["ETD"],"dc:identifier.other":["vt_gsexam:45632"],"dc:identifier.uri":["https://hdl.handle.net/10919/141131"],"dc:language.iso":["en"],"dc:publisher":["Virginia Tech"],"dc:rights":["In Copyright"],"dc:rights.uri":["http://rightsstatements.org/vocab/InC/1.0/"],"dc:subject":["Gas Chromatography","Machine Learning","Compound Classification"],"dc:title":["Machine Learning for Structure-Agnostic Chemical Analysis from Chromatographic Data"],"dc:type":["Thesis"],"thesis:degree_discipline":["Computer Science & Applications"],"thesis:degree_level":["masters"],"thesis:degree_name":["Master of Science"],"thesis:institution_name":["Virginia Polytechnic Institute and State University"]},"updated_at":"2026-07-22T22:19:10Z"}