MétaCan
Menu
Back to cohort
Record W4415065467

INSIGHT INTO APPLICATION OF MACHINE LEARNING IN NATURAL PRODUCTS CHEMINFORMATICS

2023· article· en· W4415065467 on OpenAlexaff
Said Moshawih, Hui Poh Goh, Nurolaini Kifli, Vijay Kotra, Long Chiau Ming

Bibliographic record

VenueDergiPark (Istanbul University) · 2023
Typearticle
Languageen
FieldComputer Science
TopicComputational Drug Discovery Methods
Canadian institutionsQuest University Canada
Fundersnot available
KeywordsCheminformaticsChemical spaceDruggabilityMolecular descriptorVirtual screeningDrug discoverySpace (punctuation)
DOInot available

Abstract

fetched live from OpenAlex

Cheminformatics utilizing machine learning (ML) techniques have opened up a new horizon in drug discovery. This is owing to vast chemical space expansion with rocketing numbers of expected hits and lead compounds that match druggable macromolecular targets, in particular from natural compounds. Due to the natural products’ (NP) structural complexity, uniqueness, and diversity, they could occupy a bigger space in pharmaceuticals, allowing the industry to pursue more selective leads in the nanomolar range of binding affinity. ML is an essential part of each step of the drug design pipeline, such as target prediction, compound library preparation, and lead optimization. Notably, molecular mechanic and dynamic simulations, induced docking, and free energy perturbations are essential in predicting best binding poses, binding free energy values, and molecular mechanics force fields. Those applications have leveraged from artificial intelligence (AI), which decreases the computational costs required for such costly simulations. This seminar aimed to describe chemical space and compound libraries related to NPs. Highthroughput virtual screening and their strategies in leveraging NPs libraries can be optimized to match the specificity of the chemical space that is occupied by such kind of complex compounds. Particular emphasis was given to AI approaches, ML tools, algorithms, and techniques, especially in drug discovery of macrocyclic compounds and approaches in computer-aided and ML-based drug discovery. The various functionalities and stereochemical complexities of macrocycles give them more selectivity and affinity to protein targets. Natural products were discussed as having the most distinct features differentiating them from synthetic compounds by the number of aromatic atoms, chiral centers, nitrogen, and oxygen atoms. Aromaticity is eminent among the synthetic compounds, while the chiral centers are more prevalent in NP compounds. Furthermore, the oxygen atoms are more prevalent in NPs, while nitrogen atoms are less. Those features make NPs as source of new lead compounds that can be developed using ML tools for diverse medicinal uses specifically in cancer, infectious diseases, and metabolic disorders.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.002
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Review · Consensus signal: none
Teacher disagreement score0.003
Threshold uncertainty score0.010

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0010.002
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0010.001
Science and technology studies0.0000.002
Scholarly communication0.0020.002
Open science0.0010.001
Research integrity0.0010.002
Insufficient payload (model declined to judge)0.0030.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.010
GPT teacher head0.232
Teacher spread0.222 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreReview

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2023
Admission routes1
Has abstractyes

Explore more

Same venueDergiPark (Istanbul University)Same topicComputational Drug Discovery MethodsFrench-language works237,207