ChemVassa: A New Method for Identifying Small Molecule Hits in Drug Discovery
Notice bibliographique
Résumé
ChemVassa, a new chemical structure search technology, was developed to allow rapid in silico screening of compounds for hit and hit-to-lead identification in drug development. It functions by using a novel type of molecular descriptor that examines, in part, the structure of the small molecule undergoing analysis, yielding its "information signature." This descriptor takes into account the atoms, bonds, and their positions in 3-dimensional space. For the present study, a database of ChemVassa molecular descriptors was generated for nearly 16 million compounds (from the ZINC database and other compound sources), then an algorithm was developed that allows rapid similarity searching of the database using a query molecular descriptor (e.g., the signature of atorvastatin, below). A scoring metric then allowed ranking of the search results. We used these tools to search a subset of drug-like molecules using the signature of a commercially successful statin, atorvastatin (Lipitor™). The search identified ten novel compounds, two of which have been demonstrated to interact with HMG-CoA reductase, the macromolecular target of atorvastatin. In particular, one compound discussed in the results section tested successfully with an IC50 of less than 100uM and a completely novel structure relative to known inhibitors. Interactions were validated using computational molecular docking and an Hmg-CoA reductase activity assay. The rapidity and low cost of the methodology, and the novel structure of the interactors, suggests this is a highly favorable new method for hit generation.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,002 | 0,004 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,001 |
| Méta-épidémiologie (sens large) | 0,002 | 0,001 |
| Bibliométrie | 0,005 | 0,003 |
| Études des sciences et des technologies | 0,001 | 0,001 |
| Communication savante | 0,002 | 0,002 |
| Science ouverte | 0,002 | 0,002 |
| Intégrité de la recherche | 0,001 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,010 | 0,004 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».