MétaCan
Menu
Back to cohort
Record W3195593976 · doi:10.1002/prot.26222

Prediction of protein assemblies, the next frontier: The <scp>CASP14‐CAPRI</scp> experiment

2021· article· en· W3195593976 on OpenAlexfundno aff
Marc F. Lensink, Guillaume Brysbaert, Théo Mauri, Nurul Nadzirin, Sameer Velankar, Raphaël A. G. Chaleil, Tereza Clarence, Paul A. Bates, Ren Kong, Bin Liu, Guangbo Yang, Ming Liu, Hang Shi, Xufeng Lu, Shan Chang, Raj S. Roy, Farhan Quadir, Jian Liu, Jianlin Cheng, Anna Antoniak, Cezary Czaplewski, Artur Giełdoń, Mateusz Kogut, Agnieszka G. Lipska, Adam Liwo, Emilia A. Lubecka, Martyna Maszota‐Zieleniak, Adam K. Sieradzan, Rafał Ślusarz, Patryk A. Wesołowski, Karolina Zięba, Carlos Adriel Del Carpio Munoz, Eiichiro Ichiishi, Ameya Harmalkar, Jeffrey J. Gray, Alexandre M. J. J. Bonvin, Francesco Ambrosetti, Rodrigo V. Honorato, Zuzana Jandová, Brian Jiménez‐García, Panagiotis I. Koukos, Siri van Keulen, Charlotte W. van Noort, Manon Réau, Jorge Roel‐Touris, Sergei Kotelnikov, Dzmitry Padhorny, Kathryn A. Porter, Andrey Alekseenko, Mikhail Ignatov, Israel Desta, Ryota Ashizawa, Zhuyezi Sun, Usman Ghani, Nasser Hashemi, Sándor Vajda, Dima Kozakov, Mireia Rosell, Luis Angel Rodríguez‐Lumbreras, Juan Fernández‐Recio, Agnieszka Karczyńska, Sergei Grudinin, Yumeng Yan, Hao Li, Peicong Lin, Sheng‐You Huang, Charles Christoffer, Genki Terashi, Jacob Verburgt, Daipayan Sarkar, Tunde Aderinwale, Xiao Wang, Daisuke Kihara, Tsukasa Nakamura, Yuya Hanazono, Ragul Gowthaman, Johnathan D. Guest, Rui Yin, Ghazaleh Taherzadeh, Brian G. Pierce, Didier Barradas‐Bautista, Zhen Cao, Luigi Cavallo, Romina Oliva, Yuanfei Sun, Shaowen Zhu, Yang Shen, Taeyong Park, Hyeonuk Woo, Jinsol Yang, Sohee Kwon, Jonghun Won, Chaok Seok, Yasuomi Kiyota, Shinpei Kobayashi, Yoshiki Harada, Mayuko Takeda‐Shitaka, Petras J. Kundrotas, Amar Singh, Ilya A. Vakser, Justas Dapkūnas, Kliment Olechnovič, Česlovas Venclovas, Rui Duan, Liming Qiu, Xianjin Xu, Shuang Zhang, Xiaoqin Zou, Shoshana J. Wodak

Bibliographic record

VenueProteins Structure Function and Bioinformatics · 2021
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicProtein Structure and Dynamics
Canadian institutionsnot available
FundersNational Heart, Lung, and Blood InstituteH2020 European Institute of Innovation and TechnologyLietuvos Mokslo TarybaNational Institute of General Medical SciencesMinisterio de Ciencia e InnovaciónNarodowe Centrum NaukiNational Natural Science Foundation of ChinaCancer Research UKNational Science FoundationDepartment of Energy and Climate ChangeNational Institutes of HealthInstitut national de recherche en informatique et en automatique (INRIA)Changzhou Science and Technology BureauNederlandse Organisatie voor Wetenschappelijk OnderzoekWellcome TrustMedical Research Council CanadaFrancis Crick Institute
KeywordsCASPServerComputer scienceTemplateProtein structure predictionData miningBiologyComputer networkProtein structure

Abstract

fetched live from OpenAlex

We present the results for CAPRI Round 50, the fourth joint CASP-CAPRI protein assembly prediction challenge. The Round comprised a total of twelve targets, including six dimers, three trimers, and three higher-order oligomers. Four of these were easy targets, for which good structural templates were available either for the full assembly, or for the main interfaces (of the higher-order oligomers). Eight were difficult targets for which only distantly related templates were found for the individual subunits. Twenty-five CAPRI groups including eight automatic servers submitted ~1250 models per target. Twenty groups including six servers participated in the CAPRI scoring challenge submitted ~190 models per target. The accuracy of the predicted models was evaluated using the classical CAPRI criteria. The prediction performance was measured by a weighted scoring scheme that takes into account the number of models of acceptable quality or higher submitted by each group as part of their five top-ranking models. Compared to the previous CASP-CAPRI challenge, top performing groups submitted such models for a larger fraction (70-75%) of the targets in this Round, but fewer of these models were of high accuracy. Scorer groups achieved stronger performance with more groups submitting correct models for 70-80% of the targets or achieving high accuracy predictions. Servers performed less well in general, except for the MDOCKPP and LZERD servers, who performed on par with human groups. In addition to these results, major advances in methodology are discussed, providing an informative overview of where the prediction of protein assemblies currently stands.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.013
metaresearch head score (Gemma)0.012
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: Simulation or modeling
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.013
Threshold uncertainty score0.067

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0130.012
Meta-epidemiology (narrow)0.0020.001
Meta-epidemiology (broad)0.0020.002
Bibliometrics0.0010.001
Science and technology studies0.0010.001
Scholarly communication0.0020.003
Open science0.0030.003
Research integrity0.0020.003
Insufficient payload (model declined to judge)0.0050.003

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.011
GPT teacher head0.205
Teacher spread0.194 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations127
Published2021
Admission routes1
Has abstractyes

Explore more

Same venueProteins Structure Function and BioinformaticsSame topicProtein Structure and DynamicsFrench-language works237,207