MétaCan
Menu
Retour à la cohorte
Enregistrement W2145000359 · doi:10.1371/journal.pbio.1001638

The COMBREX Project: Design, Methodology, and Initial Results

2013· article· en· W2145000359 sur OpenAlexaff
Brian P. Anton, Yi-Chien Chang, Peter Hendee Brown, Han-Pil Choi, Lina L. Faller, Jyotsna Guleria, Zhenjun Hu, Niels Klitgord, Ami Levy‐Moonshine, Almaz Maksad, Varun Mazumdar, Mark McGettrick, Lais Osmani, Revonda M. Pokrzywa, John Rachlin, Rajeswari Swaminathan, Benjamin Allen, Genevieve Housman, Caitlin Monahan, Krista Rochussen, Kevin Tao, Ashok S. Bhagwat, Steven E. Brenner, Linda Columbus, Valérie de Crécy‐Lagard, Donald J. Ferguson, Alexey Fomenkov, Giovanni Gadda, Richard Morgan, Andrei L. Osterman, Dmitry A. Rodionov, Irina A. Rodionova, Kenneth E. Rudd, Dieter Söll, James Spain, Shuang-yong Xu, Alex Bateman, Robert Blumenthal, J. Martin Bollinger, Woo‐Suk Chang, Manuel Ferrer, Iddo Friedberg, Michael Y. Galperin, Julien Gobeill, Daniel H. Haft, John Hunt, Peter D. Karp, William Klimke, Carsten Krebs, Dana Macelis, Ramana Madupu, María Martin, Jeffrey H Miller, Claire O’Donovan, Bernhard Ø. Palsson, Patrick Ruch, Aaron T. Setterdahl, Granger Sutton, John Tate, Alexander F. Yakunin, Dmitri Tchigvintsev, Germán Plata, Jie Hu, Russell Greiner, D. Horn, Kimmen Sjölander, Steven L. Salzberg, Dennis Vitkup, Stanley Letovsky, Daniel Segrè, Charles DeLisi, Richard J. Roberts, Martín Steffen, Simon Kasif

Notice bibliographique

RevuePLoS Biology · 2013
Typearticle
Langueen
DomaineBiochemistry, Genetics and Molecular Biology
ThématiqueBioinformatics and Genomic Networks
Établissements canadiensUniversity of AlbertaUniversity of Toronto
Organismes subventionnairesNational Institute of General Medical Sciences
Mots-clésBiologyEvolutionary biologyComputational biology

Résumé

récupéré en direct d'OpenAlex

Prior to the “genomic era,” when the acquisition of DNA sequence involved significant labor and expense, the sequencing of genes was strongly linked to the experimental characterization of their products. Sequencing at that time directly resulted from the need to understand an experimentally determined phenotype or biochemical activity. Now that DNA sequencing has become orders of magnitude faster and less expensive, focus has shifted to sequencing entire genomes. Since biochemistry and genetics have not, by and large, enjoyed the same improvement of scale, public sequence repositories now predominantly contain putative protein sequences for which there is no direct experimental evidence of function. Computational approaches attempt to leverage evidence associated with the ever-smaller fraction of experimentally analyzed proteins to predict function for these putative proteins. Maximizing our understanding of function over the universe of proteins in toto requires not only robust computational methods of inference but also a judicious allocation of experimental resources, focusing on proteins whose experimental characterization will maximize the number and accuracy of follow-on predictions. COMBREX (COMputational BRidges to EXperiments, http://combrex.bu.edu) is an NIH-funded enterprise that has brought computational and experimental biologists together, with the goal of greatly improving our overall understanding of microbial protein function [1],[2]. Since its inception, it has made significant progress toward the following goals: identifying the minority of proteins that have already been experimentally characterized, serving as a public repository of novel protein function predictions made by diverse methods, producing a clear chain of evidence from experiment to prediction, identifying (“recommending”) those functional predictions whose verification will contribute most to our overall understanding of protein function, and actually funding the experiments to test function. The recommendation system is a proof of concept based on active learning principles and includes, for a given protein, criteria including phylogenetic distribution of its protein family, biological and clinical phenotypes associated with it, the availability of protein structure data, and its sequence distance from experimentally determined proteins or from the other proteins in its family. COMBREX comprises several interrelated efforts. First, the project is building a community of researchers (the COMBREX Community) committed to achieving the goals above. Second, the project maintains a web-accessible database (the COMBREX Database) of known and predicted functions for microbial proteins. The database search features enable biologists to identify predictions whose experimental verification is particularly important. Finally, the project issues small monetary awards (COMBREX grants) to biologists to fund the experimental testing of such predictions. In this paper, we provide a brief review of COMBREX, focusing on its overall design, its computational resources, and the experimental results from the first phase of the project.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction distillée sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.

score de la tête « metaresearch » (Codex)0,001
score de la tête « metaresearch » (Gemma)0,000
Version: codex-gemma-dda1882f352aStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Expérimental (laboratoire) · Signal consensuel: aucune
GenreSignal candidat: Empirique · Signal consensuel: Empirique
Score de désaccord entre enseignants0,530
Score d'incertitude au seuil0,326

Scores Codex et Gemma par catégorie

CatégorieCodexGemma
Métarecherche0,0010,000
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0000,000
Bibliométrie0,0000,000
Études des sciences et des technologies0,0000,000
Communication savante0,0000,000
Science ouverte0,0000,000
Intégrité de la recherche0,0000,000
Charge utile insuffisante (le modèle a refusé de juger)0,0000,000

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,081
Tête enseignante GPT0,313
Écart entre enseignants0,232 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeExpérimental (laboratoire)
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations152
Publié2013
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revuePLoS BiologyMême sujetBioinformatics and Genomic NetworksTravaux en français237 207