Identification of potent high-affinity secondary nucleation inhibitors of Aβ42 aggregation from an ultra-large chemical library using deep docking
Bibliographic record
Abstract
Alzheimer's disease is characterized by the aggregation of the Aβ peptide into amyloid fibrils. According to the amyloid hypothesis, pharmacologically targeting Aβ aggregation could result in disease-modifying treatments. The identification of inhibitors of Aβ aggregation, however, is complicated by complex technical challenges, which typically restrict to tens of thousands the number of compounds that can be screened in experimental aggregation assays. Here, we report a computational route to increase by 4 orders of magnitude the number of screenable compounds. We achieve this result by developing an open source pipeline version of the Deep Docking protocol, and illustrate its application to the discovery of secondary nucleation inhibitors of Aβ aggregation from an ultra-large chemical library of over 539 million compounds. The pipeline was used to prioritize 35 candidate compounds for in vitro testing in Aβ aggregation assays. We found that 19 of these compounds inhibit Aβ aggregation (54% hit rate). The two most potent compounds showed potency better than adapalene, a previously reported potent inhibitor of Aβ aggregation. Consistent with the intended mechanism of action, these two compounds also proved to be high-affinity binders of Aβ fibrils with an equilibrium dissociation constant in the low nanomolar range in surface plasmon resonance experiments. These results provide evidence that structure-based docking methods based on deep learning represent a cost-effective and rapid strategy to identify potent hits for drug development targeting protein misfolding diseases.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.002 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".