Development of a highly optimized procedure for the discovery of RNA G-quadruplexes by combining several strategies
Bibliographic record
Abstract
RNA G-quadruplexes (rG4s) are non-canonical secondary structures that are formed by the self-association of guanine quartets and that are stabilized by monovalent cations (e.g. potassium). rG4s are key elements in several post-transcriptional regulation mechanisms, including both messenger RNA (mRNA) and microRNA processing, mRNA transport and translation, to name but a few examples. Over the past few years, multiple high-throughput approaches have been developed in order to identify rG4s, including bioinformatic prediction, in vitro assays and affinity capture experiments coupled to RNA sequencing. Each individual approach had its limits, and thus yielded only a fraction of the potential rG4 that are further confirmed (i.e., there is a significant level of false positive). This report aims to benefit from the strengths of several existing approaches to identify rG4s with a high potential of being folded in cells. Briefly, rG4s were pulled-down from cell lysates using the biotinylated biomimetic G4 ligand BioTASQ and the sequences thus isolated were then identified by RNA sequencing. Then, a novel bioinformatic pipeline that included DESeq2 to identify rG4 enriched transcripts, MACS2 to identify rG4 peaks, rG4-seq to increase rG4 formation probability and G4RNA Screener to detect putative rG4s was performed. This workflow uncovers new rG4 candidates whose rG4-folding was then confirmed in vitro using an array of established biophysical methods. Clearly, this workflow led to the identification of novel rG4s in a highly specific and reliable manner.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".