MétaCan
Menu
Back to cohort
Record W1517465908 · doi:10.3233/ci-2008-0016

ChemModLab: A Web-Based Cheminformatics Modeling Laboratory

2012· article· en· W1517465908 on OpenAlexaff
Jacqueline M. Hughes‐Oliver, Atina D. Brooks, William J. Welch, Morteza G. Khaledi, Douglas M. Hawkins, S. Stanley Young, Kirtesh Patil, Gary Howell, Raymond T. Ng, Moody T. Chu

Bibliographic record

VenueIn Silico Biology · 2012
Typearticle
Languageen
FieldComputer Science
TopicComputational Drug Discovery Methods
Canadian institutionsUniversity of British Columbia
FundersNational Human Genome Research InstituteNational Institute of Mental HealthNational Institutes of HealthNorth Carolina State University
KeywordsComputer scienceCheminformaticsToolboxQuantitative structure–activity relationshipData miningSet (abstract data type)VisualizationSoftwareMachine learningApplicability domainDomain (mathematical analysis)Test setArtificial intelligenceBioinformatics

Abstract

fetched live from OpenAlex

ChemModLab, written by the ECCR @ NCSU consortium under NIH support, is a toolbox for fitting and assessing quantitative structure-activity relationships (QSARs). Its elements are: a cheminformatic front end used to supply molecular descriptors for use in modeling; a set of methods for fitting models; and methods for validating the resulting model. Compounds may be input as structures from which standard descriptors will be calculated using the freely available cheminformatic front end PowerMV; PowerMV also supports compound visualization. In addition, the user can directly input their own choices of descriptors, so the capability for comparing descriptors is effectively unlimited. The statistical methodologies comprise a comprehensive collection of approaches whose validity and utility have been accepted by experts in the fields. As far as possible, these tools are implemented in open-source software linked into the flexible R platform, giving the user the capability of applying many different QSAR modeling methods in a seamless way. As promising new QSAR methodologies emerge from the statistical and data-mining communities, they will be incorporated in the laboratory. The web site also incorporates links to public-domain data sets that can be used as test cases for proposed new modeling methods. The capabilities of ChemModLab are illustrated using a variety of biological responses, with different modeling methodologies being applied to each. These show clear differences in quality of the fitted QSAR model, and in computational requirements. The laboratory is web-based, and use is free. Researchers with new assay data, a new descriptor set, or a new modeling method may readily build QSAR models and benchmark their results against other findings. Users may also examine the diversity of the molecules identified by a QSAR model. Moreover, users have the choice of placing their data sets in a public area to facilitate communication with other researchers; or can keep them hidden to preserve confidentiality.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.004
metaresearch head score (Gemma)0.006
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Software · Consensus signal: none
Teacher disagreement score0.085
Threshold uncertainty score0.285

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0040.006
Meta-epidemiology (narrow)0.0030.002
Meta-epidemiology (broad)0.0030.002
Bibliometrics0.0020.003
Science and technology studies0.0010.001
Scholarly communication0.0050.004
Open science0.0090.003
Research integrity0.0020.005
Insufficient payload (model declined to judge)0.0850.052

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.030
GPT teacher head0.316
Teacher spread0.286 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreSoftware

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations15
Published2012
Admission routes1
Has abstractyes

Explore more

Same venueIn Silico BiologySame topicComputational Drug Discovery MethodsFrench-language works237,207