<?xml version="1.0"?>
<metadata xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:dc="http://purl.org/dc/elements/1.1/"><dc:title>Benchmarking study of reaction representation for selectivity modeling in asymmetric catalysis</dc:title><dc:creator>Godec,	Nejc	(Avtor)
	</dc:creator><dc:creator>Podlipnik,	Črtomir	(Mentor)
	</dc:creator><dc:subject>Reaction representation</dc:subject><dc:subject>selectivity modeling</dc:subject><dc:subject>machine learning</dc:subject><dc:description>The prediction of enantioselectivity in asymmetric catalysis is a challenging task requiring the consideration of multiple molecular components and reaction conditions. Traditionally, the selection of catalysts and reaction conditions relies on chemical intuition and experimental screening. The application of machine learning methods offers a systematic alternative for modeling and predicting reaction selectivity. The primary challenge of this method is encoding a chemical reaction in a computer-readable descriptor vector. The purpose of this work was to benchmark different modeling strategies for predicting enantioselectivity in asymmetric catalysis. Benchmarking was performed on two datasets taken from the literature. The first dataset consists of asymmetric additions to imines catalyzed by BINOL-derived phosphoric acids, while the second dataset describes asymmetric N,S-acetal formation and includes external test sets for model validation. Various combinations of descriptor types, machine learning algorithms, and reaction representations were systematically compared. Descriptor types include fragment-based descriptors (ChyLine and CircuS), molecular fingerprints (AtomPairs, Avalon, Morgan, Morgan features, RDKFP, RDKFP layered, and Torsion), and, where available, physicochemical descriptors. Descriptors and fingerprints were calculated using the DOPtools library v.1.2 and its command line interface. Models were trained using Random Forest Regression, Support Vector Regression, and eXtreme Gradient Boosting Regression machine learning algorithms. In addition to the commonly used component-wise reaction representation, an alternative mixture representation, in which descriptors are calculated directly from a SMILES string representing the whole chemical reaction, was proposed and evaluated. Model performance was evaluated using cross-validation, leave-one-reaction-out cross-validation for the imine dataset, and external test set validation for the N,S-acetal dataset. The results show that fragment-based descriptors, particularly ChyLine fragments, resulted in better performance than fingerprints in most cases, although certain fingerprint types performed comparably in the N,S-acetal dataset. Random Forest Regression consistently produced the most reliable models across both datasets. Component-wise reaction representation generally outperformed mixture representation, the latter being computationally more efficient but showed reduced discriminative power. In conclusion, the component-wise reaction representation combined with fragment-based ChyLine descriptors and Random Forest Regression machine learning algorithm is the most suitable approach for modeling enantioselectivity in the studied asymmetric catalytic reactions. The proposed mixture representation does not provide a general replacement for the component-wise approach.</dc:description><dc:date>2026</dc:date><dc:date>2026-05-04 10:20:17</dc:date><dc:type>Magistrsko delo/naloga</dc:type><dc:identifier>182218</dc:identifier><dc:identifier>VisID: 25647</dc:identifier><dc:identifier>COBISS_ID: 278285571</dc:identifier><dc:language>sl</dc:language></metadata>
