MétaCan
Menu
Back to cohort

How Many New Genes Are There?

2006· letter· en· W2068807258 on OpenAlexaffabout
Leo J. Lee, Timothy R. Hughes, Brendan J. Frey

Bibliographic record

VenueScience · 2006
Typeletter
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicGenomics and Chromatin Dynamics
Canadian institutionsUniversity of Toronto
Fundersnot available
KeywordsGeneComputational biologyBiologyGeneticsEvolutionary biology

Abstract

fetched live from OpenAlex

In their report “The transcriptional landscape of the mammalian genome” (2 Sept. 2005, p. [1559][1]), the RIKEN Genome Exploration Research Group and Genome Science Group (Genome Network Project Core Group) and the FANTOM Consortium claim to have found 5154 new proteins in the mouse genome not encoded by previously known mRNA sequences, which could potentially correspond to a considerable number of new protein-coding genes (4311, following clustering) ([1][2]). This claim contrasts dramatically with the view of the International Human Genome Sequencing Consortium ([2][3]), which estimated that there are 20,000 to 25,000 protein-coding genes. Since there are already 22,287 genes in the Ensembl 34d catalog, this implies 0 to 2713 new genes. RIKEN/FANTOM's estimate contrasts even more strikingly with our results using exon microarrays ([3][4]), in which the number of new multi-exon protein-coding genes was estimated to be at most in the hundreds. We analyzed the putative new FANTOM proteins ([4][5]), first by comparing their sequences with RefSeq release 13 from NCBI, NIH, and including only those transcripts that are linked to a reference published no later than 1 May 2005, thus excluding all the new FANTOM proteins. Restricting our analysis to the transcripts that have strong experimental evidence (labeled Provisional, Validated, or Reviewed), we found that 2917 (56.6%) of the FANTOM proteins are in fact splice isoforms of known RefSeq transcripts, with the majority of them (2716) corresponding to exon-skipping events. By then including predicted RefSeq transcripts (labeled Genome Annotation, Inferred, Model, Predicted) in our analysis, 3568 (69.2%) were found to be splice isoforms of known transcripts. By including GenBank mRNAs linked to publications before 1 May 2005, we found an extra 303 splice isoforms, bringing the total of already-annotated genes to 3871 (75.1%). Moreover, of the 5154 FANTOM proteins, our microarray analysis detected 2293 (by two or more exons), 144 of which are among the remaining 1283 FANTOM proteins and most (131) of which are associated with known genes. We next asked whether the remaining 1193 putative proteins could be accounted for as false detections. The median open reading frame (ORF) size in this set is 119 amino acids (aa), significantly shorter than that of all the FANTOM proteins (330 aa). Although many real proteins have a length less than 119 aa, we hypothesized that such a short ORF length can arise in noncoding transcripts by chance. The FANTOM Consortium identified 23,218 nonoverlapping, noncoding transcripts, so to test this hypothesis we generated a set of 20,000 random cDNAs of 2000 bases (typical gene length) and found that 1247 of them had ORFs of 119 aa or more. Therefore, it is possible that a large portion of the remaining 1193 putative proteins arose at random from noncoding transcripts and may not encode functional polypeptides. On the basis of this analysis, the number of completely new protein-coding genes discovered by the FANTOM Consortium is at most in the hundreds, consistent with current estimates based on both sequence and microarray analysis ([2][3], [3][4]). 1. 1.[↵][6] FANTOM3 cDNA sequences are not provided, but the protein sequences can be downloaded from . 2. 2.[↵][7] Nature (2004) International Human Genome Sequencing Consortium, 431, 931. 3. 3.[↵][8] 1. B. J. Frey 2. et al. , Nat. Genet. (2005) 37, 991. 4. 4.[↵][9] See [www.psi.toronto.edu/TransLand][10] for details. # Response {#article-title-2} Lee et al. point out that the number of reported protein sequences in FANTOM3 that map to new positions on the genome appears to be too large. We are grateful to them for highlighting this discrepancy, which we investigated and thus discovered an error. For a detailed description of the correction, see the Corrections and Clarifications section in this issue. The effect of the error is somewhat less than suggested by Lee et al. In particular, our estimate of the number of new protein-coding genes found by us has been revised from 5154 to 2222, a reduction of more than half, but much less than the order of magnitude suggested by Lee et al. As correctly pointed out, the rest of the 5154 cDNAs are mainly alternatively spliced isoforms. Lee et al. present three forms of evidence: sequence similarity, exon microarray data, and ORF size. (i) The sequence homology data largely reflect the revision to the number that we mention above, except that Lee et al. used a recent RefSeq database, whereas we used Genbank (7 January 2004). There is no evidence that all RefSeq sequences correspond to real transcribed RNAs because they often include ab initio predicted exons ([1][11]). Our strategy was to construct the transcriptional frameworks entirely based on real RNA transcripts, rather than in silico reconstruction of putative gene structures. (ii) The exon microarray data concern less than 3% of the number of discussed proteins and do not have any impact on the global message of a project of the scale of FANTOM3. Despite Frey et al. 's impressive computational reconstruction of gene structure by analyzing expression patterns of ab initio predicted exons ([2][12]), we argue that this does not prove the physical structure of each mRNA and the complexity of the transcriptome with the same resolution achieved by sequencing libraries derived from mRNAs. In fact, our data show that “genes” have multiple starting and termination sites: We have conservatively identified at least 181,000 different RNA transcripts. Additionally, Frey et al. ([2][12]) used only computationally predicted exons. Rare, newly discovered transcripts are unlikely to have been in the training sets of ab initio exon identification tools, and their sensitivity to predict rare transcriptional events is not obvious. (iii) As for ORF size, 119 amino acids is a perfectly respectable size for a protein and within the bounds of statistical variation we expect. In this regard, we have further identified in the FANTOM3 dataset at least 1100 proteins shorter than 100 amino acids ([3][13]). Also, all of the novel FANTOM3 transcripts have been manually curated by individual researchers to distinguish them from novel noncoding RNAs. In any case, our final understanding of the number of protein-coding mRNAs will derive from experimental validation with full-length cDNA clones ([3][13]) rather than computational inferences. We direct interested parties to the relevant section of the FANTOM3 Web site ( ) where the updated files are available, and we thank Lee et al. for helping us to improve and update our analysis. 1. 1.[↵][14] 1. X. Pruitt 2. et al. , Nucleic Acids Res. (2005) 33, D501. 2. 2.[↵][15] 1. B. Frey 2. et al. , Nat. Genet. (2005) 37, 991. 3. 3.[↵][16] 1. M. Frith 2. et al. , Plos Genet., in press. [1]: /lookup/doi/10.1126/science.1112014 [2]: #ref-1 [3]: #ref-2 [4]: #ref-3 [5]: #ref-4 [6]: #xref-ref-1-1 View reference 1. in text [7]: #xref-ref-2-1 View reference 2. in text [8]: #xref-ref-3-1 View reference 3. in text [9]: #xref-ref-4-1 View reference 4. in text [10]: http://www.psi.toronto.edu/TransLand [11]: #ref-5 [12]: #ref-6 [13]: #ref-7 [14]: #xref-ref-5-1 View reference 1. in text [15]: #xref-ref-6-1 View reference 2. in text [16]: #xref-ref-7-1 View reference 3. in text

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.002
metaresearch head score (Gemma)0.008
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Commentary · Consensus signal: none
Teacher disagreement score0.013
Threshold uncertainty score0.042

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0020.008
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0020.004
Science and technology studies0.0010.002
Scholarly communication0.0020.004
Open science0.0010.001
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0130.007

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.008
GPT teacher head0.207
Teacher spread0.199 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreCommentary

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations8
Published2006
Admission routes2
Has abstractyes

Explore more

Same venueScienceSame topicGenomics and Chromatin DynamicsFrench-language works237,207