MétaCan
Menu
Retour à la cohorte
Enregistrement W2068807258 · doi:10.1126/science.311.5768.1709b

How Many New Genes Are There?

2006· letter· en· W2068807258 sur OpenAlexaffabout
Leo J. Lee, Timothy R. Hughes, Brendan J. Frey

Notice bibliographique

RevueScience · 2006
Typeletter
Langueen
DomaineBiochemistry, Genetics and Molecular Biology
ThématiqueGenomics and Chromatin Dynamics
Établissements canadiensUniversity of Toronto
Organismes subventionnairesnon disponible
Mots-clésGeneComputational biologyBiologyGeneticsEvolutionary biology

Résumé

récupéré en direct d'OpenAlex

In their report “The transcriptional landscape of the mammalian genome” (2 Sept. 2005, p. [1559][1]), the RIKEN Genome Exploration Research Group and Genome Science Group (Genome Network Project Core Group) and the FANTOM Consortium claim to have found 5154 new proteins in the mouse genome not encoded by previously known mRNA sequences, which could potentially correspond to a considerable number of new protein-coding genes (4311, following clustering) ([1][2]). This claim contrasts dramatically with the view of the International Human Genome Sequencing Consortium ([2][3]), which estimated that there are 20,000 to 25,000 protein-coding genes. Since there are already 22,287 genes in the Ensembl 34d catalog, this implies 0 to 2713 new genes. RIKEN/FANTOM's estimate contrasts even more strikingly with our results using exon microarrays ([3][4]), in which the number of new multi-exon protein-coding genes was estimated to be at most in the hundreds. We analyzed the putative new FANTOM proteins ([4][5]), first by comparing their sequences with RefSeq release 13 from NCBI, NIH, and including only those transcripts that are linked to a reference published no later than 1 May 2005, thus excluding all the new FANTOM proteins. Restricting our analysis to the transcripts that have strong experimental evidence (labeled Provisional, Validated, or Reviewed), we found that 2917 (56.6%) of the FANTOM proteins are in fact splice isoforms of known RefSeq transcripts, with the majority of them (2716) corresponding to exon-skipping events. By then including predicted RefSeq transcripts (labeled Genome Annotation, Inferred, Model, Predicted) in our analysis, 3568 (69.2%) were found to be splice isoforms of known transcripts. By including GenBank mRNAs linked to publications before 1 May 2005, we found an extra 303 splice isoforms, bringing the total of already-annotated genes to 3871 (75.1%). Moreover, of the 5154 FANTOM proteins, our microarray analysis detected 2293 (by two or more exons), 144 of which are among the remaining 1283 FANTOM proteins and most (131) of which are associated with known genes. We next asked whether the remaining 1193 putative proteins could be accounted for as false detections. The median open reading frame (ORF) size in this set is 119 amino acids (aa), significantly shorter than that of all the FANTOM proteins (330 aa). Although many real proteins have a length less than 119 aa, we hypothesized that such a short ORF length can arise in noncoding transcripts by chance. The FANTOM Consortium identified 23,218 nonoverlapping, noncoding transcripts, so to test this hypothesis we generated a set of 20,000 random cDNAs of 2000 bases (typical gene length) and found that 1247 of them had ORFs of 119 aa or more. Therefore, it is possible that a large portion of the remaining 1193 putative proteins arose at random from noncoding transcripts and may not encode functional polypeptides. On the basis of this analysis, the number of completely new protein-coding genes discovered by the FANTOM Consortium is at most in the hundreds, consistent with current estimates based on both sequence and microarray analysis ([2][3], [3][4]). 1. 1.[↵][6] FANTOM3 cDNA sequences are not provided, but the protein sequences can be downloaded from . 2. 2.[↵][7] Nature (2004) International Human Genome Sequencing Consortium, 431, 931. 3. 3.[↵][8] 1. B. J. Frey 2. et al. , Nat. Genet. (2005) 37, 991. 4. 4.[↵][9] See [www.psi.toronto.edu/TransLand][10] for details. # Response {#article-title-2} Lee et al. point out that the number of reported protein sequences in FANTOM3 that map to new positions on the genome appears to be too large. We are grateful to them for highlighting this discrepancy, which we investigated and thus discovered an error. For a detailed description of the correction, see the Corrections and Clarifications section in this issue. The effect of the error is somewhat less than suggested by Lee et al. In particular, our estimate of the number of new protein-coding genes found by us has been revised from 5154 to 2222, a reduction of more than half, but much less than the order of magnitude suggested by Lee et al. As correctly pointed out, the rest of the 5154 cDNAs are mainly alternatively spliced isoforms. Lee et al. present three forms of evidence: sequence similarity, exon microarray data, and ORF size. (i) The sequence homology data largely reflect the revision to the number that we mention above, except that Lee et al. used a recent RefSeq database, whereas we used Genbank (7 January 2004). There is no evidence that all RefSeq sequences correspond to real transcribed RNAs because they often include ab initio predicted exons ([1][11]). Our strategy was to construct the transcriptional frameworks entirely based on real RNA transcripts, rather than in silico reconstruction of putative gene structures. (ii) The exon microarray data concern less than 3% of the number of discussed proteins and do not have any impact on the global message of a project of the scale of FANTOM3. Despite Frey et al. 's impressive computational reconstruction of gene structure by analyzing expression patterns of ab initio predicted exons ([2][12]), we argue that this does not prove the physical structure of each mRNA and the complexity of the transcriptome with the same resolution achieved by sequencing libraries derived from mRNAs. In fact, our data show that “genes” have multiple starting and termination sites: We have conservatively identified at least 181,000 different RNA transcripts. Additionally, Frey et al. ([2][12]) used only computationally predicted exons. Rare, newly discovered transcripts are unlikely to have been in the training sets of ab initio exon identification tools, and their sensitivity to predict rare transcriptional events is not obvious. (iii) As for ORF size, 119 amino acids is a perfectly respectable size for a protein and within the bounds of statistical variation we expect. In this regard, we have further identified in the FANTOM3 dataset at least 1100 proteins shorter than 100 amino acids ([3][13]). Also, all of the novel FANTOM3 transcripts have been manually curated by individual researchers to distinguish them from novel noncoding RNAs. In any case, our final understanding of the number of protein-coding mRNAs will derive from experimental validation with full-length cDNA clones ([3][13]) rather than computational inferences. We direct interested parties to the relevant section of the FANTOM3 Web site ( ) where the updated files are available, and we thank Lee et al. for helping us to improve and update our analysis. 1. 1.[↵][14] 1. X. Pruitt 2. et al. , Nucleic Acids Res. (2005) 33, D501. 2. 2.[↵][15] 1. B. Frey 2. et al. , Nat. Genet. (2005) 37, 991. 3. 3.[↵][16] 1. M. Frith 2. et al. , Plos Genet., in press. [1]: /lookup/doi/10.1126/science.1112014 [2]: #ref-1 [3]: #ref-2 [4]: #ref-3 [5]: #ref-4 [6]: #xref-ref-1-1 View reference 1. in text [7]: #xref-ref-2-1 View reference 2. in text [8]: #xref-ref-3-1 View reference 3. in text [9]: #xref-ref-4-1 View reference 4. in text [10]: http://www.psi.toronto.edu/TransLand [11]: #ref-5 [12]: #ref-6 [13]: #ref-7 [14]: #xref-ref-5-1 View reference 1. in text [15]: #xref-ref-6-1 View reference 2. in text [16]: #xref-ref-7-1 View reference 3. in text

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,002
score de la tête « metaresearch » (Gemma)0,008
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Sans objet · Signal consensuel: aucune
GenreSignal candidat: Commentaire · Signal consensuel: aucune
Score de désaccord entre enseignants0,013
Score d'incertitude au seuil0,042

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0020,008
Méta-épidémiologie (sens strict)0,0010,000
Méta-épidémiologie (sens large)0,0010,001
Bibliométrie0,0020,004
Études des sciences et des technologies0,0010,002
Communication savante0,0020,004
Science ouverte0,0010,001
Intégrité de la recherche0,0010,001
Charge utile insuffisante (le modèle a refusé de juger)0,0130,007

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,008
Tête enseignante GPT0,207
Écart entre enseignants0,199 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeSans objet
Domainenon disponible
GenreCommentaire

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations8
Publié2006
Routes d'admission2
Résumé présentoui

Explorer davantage

Même revueScienceMême sujetGenomics and Chromatin DynamicsTravaux en français237 207