MétaCan
Menu
Back to cohort
Record W2170563234 · doi:10.1093/nar/gks1226

D2P2: database of disordered protein predictions

2012· article· en· W2170563234 on OpenAlexfundno aff
Matt E. Oates, Pedro Romero, Takashi Ishida, Mohamed Ghalwash, Marcin J. Mizianty, Bin Xue, Zsuzsanna Dosztányi, Vladimir N. Uversky, Zoran Obradović, Lukasz Kurgan, A. Keith Dunker, Julian Gough

Bibliographic record

VenueNucleic Acids Research · 2012
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicProtein Structure and Dynamics
Canadian institutionsnot available
FundersRussian Academy of SciencesBiotechnology and Biological Sciences Research CouncilNatural Sciences and Engineering Research Council of CanadaEngineering and Physical Sciences Research CouncilDirectorate for Biological SciencesUniversity of AlbertaNational Science Foundation
KeywordsBiologyContext (archaeology)GenomeProteomeDownloadParsingSet (abstract data type)Source codeComputational biologyTree (set theory)Computer scienceBioinformaticsGeneticsArtificial intelligenceWorld Wide WebGeneProgramming language

Abstract

fetched live from OpenAlex

We present the Database of Disordered Protein Prediction (D(2)P(2)), available at http://d2p2.pro (including website source code). A battery of disorder predictors and their variants, VL-XT, VSL2b, PrDOS, PV2, Espritz and IUPred, were run on all protein sequences from 1765 complete proteomes (to be updated as more genomes are completed). Integrated with these results are all of the predicted (mostly structured) SCOP domains using the SUPERFAMILY predictor. These disorder/structure annotations together enable comparison of the disorder predictors with each other and examination of the overlap between disordered predictions and SCOP domains on a large scale. D(2)P(2) will increase our understanding of the interplay between disorder and structure, the genomic distribution of disorder, and its evolutionary history. The parsed data are made available in a unified format for download as flat files or SQL tables either by genome, by predictor, or for the complete set. An interactive website provides a graphical view of each protein annotated with the SCOP domains and disordered regions from all predictors overlaid (or shown as a consensus). There are statistics and tools for browsing and comparing genomes and their disorder within the context of their position on the tree of life.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.004
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Dataset · Consensus signal: Dataset
Teacher disagreement score0.027
Threshold uncertainty score0.090

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0010.004
Meta-epidemiology (narrow)0.0040.001
Meta-epidemiology (broad)0.0030.001
Bibliometrics0.0060.007
Science and technology studies0.0010.000
Scholarly communication0.0030.002
Open science0.0050.003
Research integrity0.0020.002
Insufficient payload (model declined to judge)0.0270.031

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.028
GPT teacher head0.332
Teacher spread0.304 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreDataset

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations730
Published2012
Admission routes1
Has abstractyes

Explore more

Same venueNucleic Acids ResearchSame topicProtein Structure and DynamicsFrench-language works237,207