University of Kansas
Explainable Machine Learning Prediction of Antimicrobial Peptide Targeting Streptococcus mutans
Abstract
dc:description.abstractStreptococcus mutans is mainly recognized for its role in dental caries, yet it is also increasingly recognized as a significant contributor to systemic diseases. When it gains access to the bloodstream during daily oral hygiene or a dental procedure, S. mutans can become a dangerous systemic pathogen. In dental plaque, within the oral cavity, interactions between bacteria and dietary sugars lead to acid production and selection of acid-tolerant species. S. mutans thrives on the tooth enamel through adhesion and biofilm formation, where continuous acid production drives demineralization and leads to dental caries. This thesis developed an explainable machine learning pipeline to design antimicrobial peptides targeting S. mutans in responsiveness to the acidic conditions of the cariogenic microenvironments. Our explainable ML pipeline approach transforms the targeted peptide design into a data-driven methodology and bridges the explainability layers to uncouple and optimize features independently. The study first assembled a curated dataset of 100 unique anti-S. mutans peptides with harmonized minimum inhibitory concentration (MIC) values, then trained and benchmarked classification and regression models across physicochemical descriptors, protein language model embeddings, and combined feature representations. Descriptor-based classical models performed best for binary potency classification, whereas combined representations performed best for quantitative potency prediction. Feature-importance analysis identified molecular weight, net charge, estimated isoelectric point, length, and sequence entropy as the main predictors of potency, supporting charge-related behavior as a useful design axis under cariogenic conditions. We used a robust peptide discovery by integrating explicit generation constraints into a pipeline that uses variational autoencoder (VAE) sampling, genetic algorithm (GA) optimization, and guided decoding to generate new peptide candidates. Our ML design combines the distinct phases through VAE sampling which creates a continuous design space, GA optimization enables to search by soft constraints and guided decoding enforcing hard rules. The integrated workflow enables a closed-loop and highly controllable generative pipeline. The resulting peptide candidates were subsequently screened in silico for predicted hemolysis, toxicity, solubility, structural plausibility, pH-dependent behavior, and novelty relative to the training set. The thesis generated a pipeline that yielded 20 prioritized Tier 1 candidates predicted to be non-hemolytic, non-toxic, soluble, and structurally plausible, providing lead sequences for future experimental testing.
Degree
thesis:*- Name thesis:degree_name
- M.S.
- Discipline thesis:degree_discipline
- Bioengineering
- Grantor dc:publisher
- University of Kansas
- Year dc:date.issued
- 2026
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Riad, Md Aliul Islam
- Advisor dc:contributor.advisor
-
- Tamerler, Candan
Subjects
dc:subject × 6Rights
- Language dc:language.iso
- en
Identifiers
dc:identifier.*- Dc Identifier Other
- https://www.proquest.com/LegacyDocView/DISSNUM/32703470
- OAI identifier oai:identifier
- oai:kuscholarworks.ku.edu:1808/39521