Persistent hindrances to data re-use in single-cell genomics
Bibliographic record
Abstract
Abstract We report on our experience attempting to re-use published and publicly available single-cell (or single-nucleus) RNA-sequencing studies (scRNA-seq) from the Gene Expression Omnibus (GEO). We screened GEO for human, mouse and rat scRNA-seq studies as potential candidates for inclusion in the Gemma database of re-annotated and re-analyzed transcriptome studies. Using semi-automated and manual curation, we assessed whether GEO datasets included cell-level expression count matrices and cell-type annotations. We found that there are steep challenges to data reuse. Only ∼40% of studies provided readily usable processed count data that could be reliably mapped to GEO metadata, and fewer than 10% included author-provided cell-type annotations. While raw sequencing data were available for the majority of studies, only a small proportion could be re-analyzed automatically without reliance on heuristics. Our findings show that existing practices for single-cell RNA-sequencing data distribution and sharing are insufficient for effective reuse, and highlight the urgent need for repositories to strengthen and enforce submission requirements, particularly for processed data and cell-type annotations.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".