Improving Synthetic Efficiency Using the Computational Prediction of Biological Activity
Bibliographic record
Abstract
A process has been developed whereby libraries of compounds for lead optimization can be synthesized and screened with greater efficiency using computational tools. In this method, analogues of a lead chemical structure are considered in the form of a virtual library. Less than 1/3 of the library is selected as a training set by clustering the compounds and choosing the centroid of each cluster. This training set is then used to generate a model using PLS regression upon the experimental values from that assay using 1D/2D descriptors. The model is applied to the remaining compounds (the test set) for which assay values are predicted and a rank ordering established. An example of this was a set of 169 PDE4 inhibitors. A predictive model was achieved using a training set of 52 compounds. When applied to the remaining 117 compounds this model allowed a rank ordering of these compounds for synthesis and testing. Selecting the top 33 compounds of the test set gives 78% of the compounds with the desired activity (hits) by synthesizing only 50% of the library, including the training set. Selecting the top 59 of the test set gives 97% of the hits from only 67% of the library. This process succeeds by avoiding two principal weaknesses of 2D descriptors: lack of interpretation and lack of extrapolation. Two principal assumptions of QSAR are shown to be unnecessary; removing descriptor redundancy does not improve fit and a predictive r2 greater than 0.5 is not necessary if rank-ordering is desired.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.006 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".