MétaCan
Menu
Back to cohort
Record W2154813243 · doi:10.1002/prot.22850

Blind predictions of protein interfaces by docking calculations in CAPRI

2010· article· en· W2154813243 on OpenAlexafffund
Marc F. Lensink, Shoshana J. Wodak

Bibliographic record

VenueProteins Structure Function and Bioinformatics · 2010
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicProtein Structure and Dynamics
Canadian institutionsHospital for Sick ChildrenUniversity of Toronto
FundersCanadian Institutes of Health Research
KeywordsDocking (animal)Computer scienceExtant taxonProtein functionMacromolecular dockingProtein–protein interactionComputational biologyData miningBiological systemProtein structureChemistryBiologyMedicineBiochemistry

Abstract

fetched live from OpenAlex

Reliable prediction of the amino acid residues involved in protein-protein interfaces can provide valuable insight into protein function, and inform mutagenesis studies, and drug design applications. A fast-growing number of methods are being proposed for predicting protein interfaces, using structural information, energetic criteria, or sequence conservation or by integrating multiple criteria and approaches. Overall however, their performance remains limited, especially when applied to nonobligate protein complexes, where the individual components are also stable on their own. Here, we evaluate interface predictions derived from protein-protein docking calculations. To this end we measure the overlap between the interfaces in models of protein complexes submitted by 76 participants in CAPRI (Critical Assessment of Predicted Interactions) and those of 46 observed interfaces in 20 CAPRI targets corresponding to nonobligate complexes. Our evaluation considers multiple models for each target interface, submitted by different participants, using a variety of docking methods. Although this results in a substantial variability in the prediction performance across participants and targets, clear trends emerge. Docking methods that perform best in our evaluation predict interfaces with average recall and precision levels of about 60%, for a small majority (60%) of the analyzed interfaces. These levels are significantly higher than those obtained for nonobligate complexes by most extant interface prediction methods. We find furthermore that a sizable fraction (24%) of the interfaces in models ranked as incorrect in the CAPRI assessment are actually correctly predicted (recall and precision ≥50%), and that these models contribute to 70% of the correct docking-based interface predictions overall. Our analysis proves that docking methods are much more successful in identifying interfaces than in predicting complexes, and suggests that these methods have an excellent potential of addressing the interface prediction challenge.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.017
metaresearch head score (Gemma)0.028
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: Simulation or modeling
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.017
Threshold uncertainty score0.088

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0170.028
Meta-epidemiology (narrow)0.0030.001
Meta-epidemiology (broad)0.0020.002
Bibliometrics0.0040.002
Science and technology studies0.0010.001
Scholarly communication0.0040.002
Open science0.0030.005
Research integrity0.0020.002
Insufficient payload (model declined to judge)0.0040.003

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.004
GPT teacher head0.215
Teacher spread0.211 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations81
Published2010
Admission routes2
Has abstractyes

Explore more

Same venueProteins Structure Function and BioinformaticsSame topicProtein Structure and DynamicsFrench-language works237,207