Back to results

University of Illinois at Urbana-Champaign

Data-rich experimentation and computer-aided strategies in reaction discovery, selectivity optimization, and structure elucidation

Abstract

dc:description

Hypothesis-driven empirical research is prevalent in synthetic organic chemistry. Traditional workflows rely heavily on the intuition of experimentalists and their ability to and trends or key results in sparse and deficient data to formulate testable hypotheses. An emergent paradigm in chemistry is an approach, in which the data is collected en masse, and the computer-aided analysis then allows to reduce the data and identify meaningful trends. High-Throughput Experimentation and the approach of data rich experimentation allows to experimentally navigate large chemical spaces during reaction optimizations. Despite the success of HTE in industrial applications, it is not nearly as popular in fundamental research, as it is often more important to demonstrate novel findings, rather than a perfectly optimized process. The greatest intellectual challenge of data rich chemical science is to provide smart solutions towards the experiment design, data extraction and modeling to aid synthetic efforts. This document is a compilation of four different projects, which tell a story about the power of synergy between experimental and computational science in chemistry. Chapter 1 of this document contains the description of the data-rich platform for the reaction discovery. A structure-agnostic reaction product identification by means of Liquid Chromatography-Mass Spectrometry approach was realized through a stable isotope fingerprinting strategy. The introduction of traceable isotopic multiplets into the mass spectra by means of selective deuteration of the starting materials allowed to identify the presence of homocoupling or heterocoupling products up to 3 components. Paired with an unsupervised machine learning approach towards decomposing complicated mass spectral datasets, and an automated pattern identification, we were able to achieve a robust discovery platform. Its utility was then demonstrated in the discovery of several previously unknown reactions. Chapter 2 of this document describes the chemoinformatics-guided approach to selection of the phosphoramidite ligand universal training set. A large ligand library was enumerated from its 2D depictions using molli package, described later in this document. The library was then clustered using a k-means clustering approach on the reduced average steric occupancy + average electronic indicator field descriptors to provide a set of compounds termed the phosphoramidite univeral training set. This set was used in a high-throughput screening campaign in the borylative Heck reaction sequence, which allowed to identify highly selective reaction conditions. Chapter 3 of this document describes the development and application of the molli software package. Since the adoption of the rich in silico data oriented strategy, the necessity to handle large molecule collections programmatically increased, as did the need for a modern, consistent and fast interface to those functions. In this section, the advances towards a novel approach to parsing of ChemDraw™ .CDXML les via z-coordinate hinting are described. Benchmarking of the improved GBCA descriptor calculations, as well as algorithmic improvements are listed. Example end-to-end calculation workflows are discussed. Chapter 4 of this document describes the computational approach towards the structural investigation of an unprecedented Pd–Pd dimeric structure. It illustrates the power and the pitfalls of Density Functional Theory study of a complex, which cannot be isolated in its pure form. By means of density functional theory and multireference self-consistent field approaches we were able to establish an openshell nature of one of the key metallic intermediates in the anhydrous Suzuki–Miyaura cross-coupling reaction. So far unprecedented avoided Pd–Pd metallic bonding is described, and the likely hypotheses for the emergence of this phenomenon were formulated.

Degree

thesis:*
Name thesis:degree_name
Ph.D.
Level thesis:degree_level
Dissertation
Discipline thesis:degree_discipline
Chemistry
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2024

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Shved, Alexander S
Contributors dc:contributor
  • Denmark, Scott E
  • Sarlah, David
  • Fataftah, Majed S
  • Pogorelov, Taras V

Subjects

dc:subject × 3

Rights

dc:rights
Statement dc:rights
  • © 2024 Alexander S. Shved
Language dc:language
en, eng

Identifiers

dc:identifier.*
Handle dc:identifier
https://hdl.handle.net/2142/127502

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Shved, Alexander S. Data-rich experimentation and computer-aided strategies in reaction discovery, selectivity optimization, and structure elucidation. Dissertation thesis, University of Illinois at Urbana-Champaign, 2024. https://hdl.handle.net/2142/127502