MétaCan
Menu
← Retour à la cohorte
Enregistrement W4393422656 · doi:10.5281/zenodo.7589952

Cell-Cell Interaction Database

2020· dataset· en· W4393422656 sur OpenAlexaff
Ruth Isserlin, Véronique Voisin, Laurie Ailles, Gary D. Bader

Notice bibliographique

RevueZenodo (CERN European Organization for Nuclear Research) · 2020
Typedataset
Langueen
DomaineBiochemistry, Genetics and Molecular Biology
ThématiqueCell Image Analysis Techniques
Établissements canadiensLunenfeld-Tanenbaum Research InstituteSinai Health SystemPrincess Margaret Cancer CentreUniversity Health NetworkUniversity of Toronto
Organismes subventionnairesnon disponible
Mots-clésDatabaseComputer scienceCellBiologyGenetics

Résumé

récupéré en direct d'OpenAlex

Overview This page describes the automated construction of a cell-cell interaction database by filtering existing curated protein-protein interaction (PPI) data. Cell-cell interactions are important for understanding tissue organization. We and others have built cell-cell interaction databases (1-5). The resource available from this website represents an automatically built set of protein-protein interactions that can mediate cell-cell communication that is expanded compared to previous databases we have built. Receptors Receptor genes were defined based on the union of the annotations from the set of Gene Ontology (GO) terms (6,7): GO:0043235 - receptor complex, GO:0008305 - integrin complex, GO:0072657 - protein localized to membrane GO:0043113 - receptor clustering GO:0004872 - receptor activity, GO:0009897 - external side of plasma membrane) UniProt annotations search term -"Receptor [KW-0675]" go:0005886 organism:human. This created a set of 4364 receptor genes (prior to manual curation) Ligands Ligand genes were defined based on the union of the below annotations the GO terms (6,7): GO:0005102 - receptor binding the set of proteins labelled as secreted in the Secretome dataset (http://www.proteinatlas.org/humanproteome/secretome) (8). This created a set of 3209 Ligand genes (prior to manual curation) Extracellular Matrix Extracellular Matrix (ECM) genes were defined based on the union of the annotations from the GO terms (6,7): GO:0031012 - extracellular matrix GO:0005578 - proteinacious extracellular matrix GO:0005201 - extracellular matrix structural constituent GO:1990430 - extracellular matrix protein binding GO:0035426 - extracellular matrix cell signalling This created a set of 433 ECM genes (prior to manual curation) Manual Curation ECM, Receptor and ligand lists were manually curated genes that were neither receptors or ligands were removed misclassified genes were moved to the correct list (i.e. receptors found on the ligand list or vice versa) After curation, the resulting ligand, receptor and ECM sets consisted of: Receptors - 1851 genes Ligands - 1593 genes ECM - 433 genes In each of the above sets there are genes that are part of other sets (e.g. a gene can be ECM and ligand at the same time) Interaction Data The set of protein interactions were downloaded from: iRefIndex (version 14) (9). - all BioGRID interactions were excluded from the iRefIndex set as we imported the original source. Pathway Commons (version 8) (10). BioGRID (version 3.4.147) (11). The entire interaction set was filtered to only include interactions that contained receptor-ligand, receptor-receptor, ligand-ligand, receptor-ecm, ligand-ecm or ecm-ecm interactions where the receptor, ligands and ecm were defined by the above lists. The resulting Receptor-Ligand network contained 2,593 unique proteins and 38,446 unique interactions (115,900 interaction total) Data files ligands.txt - table of ligands. (contains HGNC symbol and classification (Ligand, Ligand/ECM, Ligand/Receptor, Ligand/ECM/Receptor) receptors.txt - table of receptors. (contains HGNC symbol and classification (Receptor, Receptor/ECM, Ligand/Receptor, Ligand/ECM/Receptor) ecm.txt - table of ECM. (contains HGNC symbol and classification (ECM, ECM/Receptor, ECM/Ligand, Ligand/ECM/Receptor) protein_types.txt - table of unique set of receptor, ligand and ECM genes (all of the above tables: contains HGNC symbol as well as classification (Receptor, Ligand, ECM, ECM/Receptor, ECM/Ligand, Receptor/Ligand, Ligand/ECM/Receptor) receptor_ligand_interactions_mitab_v1.0_April2017.txt(.zip/.gz) - tab delimited file in mitab 2.5 format containing the following columns: AliasA - main Alias for molecule A (often the recognized gene symbol) AliasB- main Alias for molecule B (often the recognized gene symbol) uidA - unique identifier for molecule A (depending on the source database this can be one of the following types uniprot, refseq, entrez gene id, ensembl) uidB - unique identifier for molecule A (depending on the source database this can be one of the following types uniprot, refseq, entrez gene id, ensembl) altA - list of alternate identifiers for molecule A. altB - list of alternate identifiers for molecule B. aliasA - list of alternate aliases for molecule A. aliasB - list of alternate aliases for molecule B. method - list of psi-mi terms indicating experimental methods used to discover interaction. author - text listing authors pmids - list of pmids associated with the interaction. taxa - taxon id for molecule A. taxb - taxon id for molecule B. interactionType - list of psi-mi terms indicating the type of interactions it is. sourcedb - source database. interactionIdentifier - source database interaction identifier confidence - confidence of interaction as supplied by database source References Qiao W, Wang W, Laurenti E, Turinsky AL, Wodak SJ, Bader GD, Dick JE, Zandstra PW Intercellular network structure and regulatory motifs in the human hematopoietic system Pubmed Kirouac DC, Ito C, Csaszar E, Roch A, Yu M, Sykes EA, Bader GD, Zandstra PW. Dynamic interaction networks in a hierarchically organized tissue. Mol Syst Biol. 2010 Oct 5;6:417 Pubmed Yuzwa SA, Yang G, Borrett MJ, Clarke G, Cancino GI, Zahr SK, Zandstra PW, Kaplan DR, Miller FD. Proneurogenic Ligands Defined by Modeling Developing Cortex Growth Factor Communication Networks. Neuron. 2016 Sep 7;91(5):988-1004 Pubmed Ramilowski JA, Goldberg T, Harshbarger J, Kloppmann E, Lizio M, Satagopam VP, Itoh M, Kawaji H, Carninci P, Rost B, Forrest AR. A draft network of ligand-receptor-mediated multicellular signalling in human. Nat Commun. 2015 Jul 22;6:7866. Pubmed Rieckmann JC, Geiger R, Hornburg D, Wolf T, Kveler K, Jarrossay D, Sallusto F, Shen-Orr SS, Lanzavecchia A, Mann M, Meissner F. Social network architecture of human immune cells unveiled by quantitative proteomics. Nat Immunol. 2017 May;18(5):583-593. PMID: 28263321. Ashburner M, Ball CA, Blake JA, Botstein D, Butler H, Cherry JM, Davis AP, Dolinski K, Dwight SS, Eppig JT, Harris MA, Hill DP, Issel-Tarver L, Kasarskis A, Lewis S, Matese JC, Richardson JE, Ringwald M, Rubin GM, Sherlock G. Gene ontology: tool for the unification of biology. The Gene Ontology Consortium. Nat Genet. 2000 May;25(1):25-9 Pubmed The Gene Ontology Consortium. Expansion of the Gene Ontology knowledgebase and resources. Nucleic Acids Res. 2017 Jan 4;45(D1):D331-D338 Pubmed Uhlén M, Fagerberg L, Hallström BM, Lindskog C, Oksvold P, Mardinoglu A, Sivertsson Å, Kampf C, Sjöstedt E, Asplund A, Olsson I, Edlund K, Lundberg E, Navani S, Szigyarto CA, Odeberg J, Djureinovic D, Takanen JO, Hober S, Alm T, Edqvist PH, Berling H, Tegel H, Mulder J, Rockberg J, Nilsson P, Schwenk JM, Hamsten M, von Feilitzen K, Forsberg M, Persson L, Johansson F, Zwahlen M, von Heijne G, Nielsen J, Pontén F. Proteomics. Tissue-based map of the human proteome. Science. 2015 Jan 23;347(6220) Pubmed Razick S, Magklaras G, Donaldson IM. iRefIndex: a consolidated protein interaction database with provenance. BMC Bioinformatics. 2008 Sep 30;9:405 Pubmed Cerami EG, Gross BE, Demir E, Rodchenkov I, Babur O, Anwar N, Schultz N, Bader GD, Sander C. Pathway Commons, a web resource for biological pathway data. Nucleic Acids Res. 2011 Jan;39(Database issue):D685-90.2010 Nov 10. Pubmed Stark C, Breitkreutz BJ, Reguly T, Boucher L, Breitkreutz A, Tyers M. BioGRID: a general repository for interaction datasets. Nucleic Acids Res. 2006 Jan 1;34(Database issue):D535-9. Pubmed

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,001
score de la tête « metaresearch » (Gemma)0,004
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Sans objet · Signal consensuel: Sans objet
GenreSignal candidat: Jeu de données · Signal consensuel: Jeu de données
Score de désaccord entre enseignants0,130
Score d'incertitude au seuil0,434

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0010,004
Méta-épidémiologie (sens strict)0,0020,001
Méta-épidémiologie (sens large)0,0020,001
Bibliométrie0,0070,008
Études des sciences et des technologies0,0020,000
Communication savante0,0040,002
Science ouverte0,0040,004
Intégrité de la recherche0,0020,002
Charge utile insuffisante (le modèle a refusé de juger)0,1300,132

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,021
Tête enseignante GPT0,265
Écart entre enseignants0,244 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeSans objet
Domainenon disponible
GenreJeu de données

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2020
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueZenodo (CERN European Organization for Nuclear Research)→Même sujetCell Image Analysis Techniques→Travaux en français237 207→