MétaCan
Menu
← Back to cohort
Record W4414465126 · doi:10.1101/2025.09.19.677274

RNAcentral in 2026: Genes and literature integration

2025· preprint· en· W4414465126 on OpenAlexaff
Andrew Green, Carlos Eduardo Ribas, Isaac Jandalala, Philippa Muston, Colman O’Cathail, Guy Cochrane, Christina Ernst, Lingyun Zhao, Pedro Madrigal, Helen Attrill, Steven J Marygold, Doron Lancet, Niv Dobzinski, Patricia P. Chan, Todd M. Lowe, Elspeth A. Bruford, Ruth L. Seal, Henning Hermjakob, Kalpana Panneerselvam, ROBERT FINN, Tatiana A. Gurbich, Sam Griffiths‐Jones, Bastian Fromm, Kevin J. Peterson, Dominik Sordyl, Janusz M. Bujnicki, Sameer Velankar, Sri Devan Appasamy, Sudakshina Ganguly, Peng Zhang, Shunmin He, Kim Rutherford, Valerie Wood, Ruth C. Lovering, Ernesto Picardi, Nancy Ontiveros‐Palacios, Lin Huang, Zhichao Miao, Anton S. Petrov, Holly McCann, Emanuele Cavalleri, Marco Mesiti, Elena Rivas, Marcell Szikszai, Marcin Magnus, Jan Gerken, Maria Chuvochina, Danny Bergeron, Michelle S. Scott, Kelly P. Williams, Mark Quinton-Tulloch, Stavros Diamantakis, Anton I. Petrov, Blake Sweeney

Bibliographic record

VenuebioRxiv (Cold Spring Harbor Laboratory) · 2025
Typepreprint
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicRNA Research and Splicing
Canadian institutionsUniversité de Sherbrooke
FundersBiotechnology and Biological Sciences Research CouncilDirectorate for Biological SciencesNuclear Safety and Security CommissionNational Aeronautics and Space AdministrationNASA Astrobiology InstituteWellcome TrustNational Science Foundation
KeywordsSequence (biology)Interface (matter)GeneData integrationSequence analysisGene predictionBiological databaseRNA splicing

Abstract

fetched live from OpenAlex

Abstract RNAcentral was founded in 2014 to serve as a comprehensive database of non-coding RNA sequences. It began by providing a single unified interface to more specialised resources, and now contains 45 million sequences. It has grown beyond providing a single interface to many specialised resources and now provides several services and analyses. These include secondary structure prediction with R2DT, sequence search, and analysis with Rfam. Since its last publication in 2021, RNAcentral has developed two major features. First, literature integration with the development of LitScan and LitSumm. LitScan automatically identifies and links relevant publications to RNA entries, while LitSumm uses natural language processing to generate functional summaries from the literature. Together, these tools address the critical challenge of connecting sequence data with scattered functional knowledge across thousands of publications. Secondly, RNAcentral has created gene level entries. Gene level entries represent a large structural change to RNAcentral. While RNAcentral previously organized data exclusively at the sequence level, we now group related transcripts into gene-centric views. This allows researchers to explore all isoforms, splice variants, and related sequences for a gene in a unified interface, better reflecting biological organization and facilitating comparative analyses. RNAcentral is freely available at: https://rnacentral.org .

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.005
metaresearch head score (Gemma)0.015
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Methods · Consensus signal: none
Teacher disagreement score0.104
Threshold uncertainty score0.347

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0050.015
Meta-epidemiology (narrow)0.0020.001
Meta-epidemiology (broad)0.0020.002
Bibliometrics0.0270.019
Science and technology studies0.0020.001
Scholarly communication0.0080.003
Open science0.0020.006
Research integrity0.0020.002
Insufficient payload (model declined to judge)0.1040.098

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.008
GPT teacher head0.238
Teacher spread0.230 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venuebioRxiv (Cold Spring Harbor Laboratory)→Same topicRNA Research and Splicing→French-language works237,207→