<i>Whippet</i> : an efficient method for the detection and quantification of alternative splicing reveals extensive transcriptomic complexity
Bibliographic record
Abstract
Abstract Alternative splicing (AS) is a widespread process underlying the generation of transcriptomic and proteomic diversity in metazoans. Major challenges in comprehensively detecting and quantifying patterns of AS are that RNA-seq datasets are expanding near exponentially, while existing analysis tools are computationally inefficient and ineffective at handling complex splicing patterns. Here, we describe Whippet , a method that rapidly, and with minimal hardware requirements, models and quantifies splicing events of any complexity without significant loss of accuracy. Using an entropic measure of splicing complexity, Whippet reveals that approximately 33% of human protein coding genes contain complex AS events that result in substantial expression of multiple splice isoforms. These events frequently affect tandem arrays of folded protein domains. Remarkably, high-entropy AS events are more prevalent in tumour relative to matched normal tissues, and these differences correlate with increased expression of proto-oncogenic splicing factors. Whippet thus affords the rapid and accurate analysis of AS events of any complexity, and as such will facilitate biomedical research.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.011 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".