Recommended Free Tools
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Active learning can help researchers decide which chemical reactions to run next, while geometric deep learning can help predict which atom in a molecule will react. A 2026 study by Mason Minot, Yannick Stenzhorn, Jens Wolfard and colleagues combined those approaches for C–H borylation, a reaction used to add a chemical handle that can support later molecule diversification. The work addresses a practical challenge in drug discovery: making useful predictions when relevant experimental data are limited.
Why reaction prediction is difficult in drug discovery
A molecule’s usefulness depends not only on its properties but also on whether chemists can make and modify it. Drug-like molecules can contain several similar carbon–hydrogen (C–H) sites, and a reaction may occur at one site rather than another. Predicting that position is called regioselectivity prediction.
| # | Preview | Product | Price | |
|---|---|---|---|---|
| 1 |
|
Basic Principles of Drug Discovery and Development | $132.53 | Buy on Amazon |
| 2 |
|
Drugs: From Discovery to Approval | $59.12 | Buy on Amazon |
| 3 |
|
Chemistry and Pharmacology of Drug Discovery | $130.46 | Buy on Amazon |
| 4 |
|
Molecular Targeted Drug Discovery: A Guide to How Modern Medicines are Created | $145.00 | Buy on Amazon |
| 5 |
|
Textbook of Drug Design and Discovery | $81.34 | Buy on Amazon |
The study focuses on C–H borylation. It replaces a C–H bond with a boronate ester, which can serve as a handle for later cross-coupling and other diversification. That makes the reaction relevant to late-stage functionalization: changing a complex molecule near the end of a synthesis rather than rebuilding it from scratch.
Prediction is harder when the available data are sparse or unrepresentative. A model trained on familiar molecular series may perform less reliably on a new scaffold, and a useful experiment must measure more than whether a reaction succeeds: for regioselectivity, the chemist also needs to know where it happens.
#1 Best Overall
How the closed-loop workflow works
The authors connect experiment selection with model training in a repeated cycle. Instead of treating data collection as a separate preliminary task, they use results from one stage to guide which experiments to perform next.
- Start with experimental data. The initial active-learning selection set covered 518 substrates. The authors benchmarked Random Forest, CatBoost and XGBoost ensembles, selecting XGBoost as a practical choice based on experimental-set classification performance and uncertainty calibration.
- Prioritize experiments. The XGBoost ensemble served as an oracle to help choose candidate substrates and reaction configurations. In this workflow, its reported scoring time was approximately 0.1 seconds for 22,253 candidate molecules; that is a result for this candidate pool, not a general speed guarantee.
- Run prospective experiments. Chemists tested selected candidates and added the outcomes to the dataset. Three rounds tested 50 previously unreported substrates in total.
- Train outcome and position models. With the expanded data, the researchers trained geometric graph neural networks (GNNs) to predict reaction feasibility and atom-level regioselectivity. A GNN represents a molecule as a graph and can also use its three-dimensional geometry.
The acquisition model and the final atom-level task are connected indirectly: the active-learning oracle prioritizes binary reaction feasibility, while regioselectivity asks which particular atom reacts. The workflow therefore does not use one objective as a direct substitute for the other.
Rank #2
What data the study collected
The reported datasets grew through prospective work and additional Roche borylation experiments. The article describes an initial 518-substrate selection set and a final combined feasibility dataset covering 568 unique substrates.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
| Dataset or activity | Reported scale | What it represents |
|---|---|---|
| Yield and binary reaction outcomes | 6,865 reaction records across 568 unique substrates | Data used for yield-related and binary feasibility modeling. A reaction was labeled positive under the binary rule when yield was at least 5%. |
| Prospective active-learning rounds | 4,821 reactions from approximately 96 reaction configurations | Results across three rounds on 50 previously unreported substrates; the rounds screened 30 substrates, then 10, then 10. |
| Regioselectivity | 812 starting materials and 920 borylated products | A final set expanded with products from active-learning rounds and other Roche borylation experiments. |
In the reported binary outcome data, 24% of reactions were positive under the at-least-5%-yield labeling rule. This proportion describes the study’s dataset, not the general success rate of C–H borylation. The earlier binary data were more imbalanced, with 7% negative reactions; the authors identify class imbalance as one source of performance variability.
Rank #3
How the models compare
The paper compares ten geometric GNNs with XGBoost models using three kinds of data split. A random split can put closely related molecules in both training and test data. Butina clustering and Bemis–Murcko scaffold splits provide harder tests of whether a model generalizes to molecular clusters or scaffolds that were held out from training.
| Evaluation split | What it tests | Reported result |
|---|---|---|
| Random | Performance when train and test sets may contain related molecular examples | GNN and condition-aware XGBoost results were more comparable. |
| Butina-clustered | Generalization across molecular clusters | The authors report that GNNs outperformed XGBoost in this more challenging setting. |
| Bemis–Murcko scaffold | Generalization to held-out molecular scaffolds | On the final active-learning round’s scaffold split, GNNs achieved mean Matthews correlation coefficient (MCC) values of 0.43–0.50. The strongest condition-aware XGBoost comparator scored 0.35 ± 0.12. |
MCC is a measure of classification quality that accounts for both positive and negative predictions; it is useful when class counts are uneven. The reported scaffold results are specific to this paper’s data and evaluation setup, not a universal performance benchmark.
The input representation also mattered. Fingerprint-only XGBoost performed poorly across the splits, while the condition-aware XGBoost model was stronger. Geometric GNNs use molecular structure and three-dimensional information, which may help when predicting outcomes depends on the arrangement of atoms and reaction sites. Still, the authors found no single GNN architecture that was consistently best across their tests.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →What the prospective regioselectivity tests show
The authors report that prospective tests on unseen substrates with challenging N-heteroaryl motifs identified the correct borylation positions in all reported cases. This is a promising result for those tested substrates, not a guarantee that the approach will correctly predict regioselectivity for every drug-like molecule or every reaction condition.
Best Value
The distinction matters in practice. The feasibility model helps prioritize whether a reaction is likely to produce a meaningful outcome; the regioselectivity model addresses which site reacts. A chemist still needs to conduct and validate experiments, particularly when applying a prediction to a new scaffold or a consequential synthesis decision.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Why the self-supervised tasks matter
The GNN training included online auxiliary tasks—node masking and coordinate denoising—alongside the labeled reaction tasks. Node masking asks a model to recover masked information about atoms in the molecular graph; coordinate denoising trains it to work with perturbed three-dimensional coordinates. These tasks provide additional learning signals from the reaction data being used, without requiring a separate unlabeled molecular dataset or an offline pretraining stage.
The authors report that these auxiliary tasks generally improved performance across the tested models. That finding supports their use in this setting, but does not establish that either task will help every architecture or dataset.
Free tools Windows power users keep installed
One-click scans. No signup required.
What the findings do—and do not—establish
The study’s main contribution is a data-to-model workflow: active learning directs prospective experiments, and the resulting data support geometric models for feasibility and regioselectivity. The authors also report scaffold- and cluster-split advantages for GNNs over XGBoost, while random-split results were more comparable.
- It does show: a closed-loop strategy can collect new borylation data on previously unreported substrates and use the expanded data to train reaction-prediction models.
- It does not show: that any one GNN architecture is best in every setting, that XGBoost is universally optimal as an experiment-selection oracle, or that predictions can replace laboratory validation.
- It leaves broader challenges: the authors call for larger and more diverse public reaction datasets and note class-imbalance concerns. They also report reduced accuracy for EquiformerV2 on the most structurally intricate substrates.
For researchers who want to inspect the work, the paper identifies the SURF-formatted yield and regioselectivity datasets as Zenodo record 10.5281/zenodo.20773622. It lists the reference implementation in the GitHub repository minotm/active-drug-discovery and the code and model weights in Zenodo record 10.5281/zenodo.20783136. The paper states that the reference implementation and weights are released under GPLv3.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

