MétaCan
Menu
Back to cohort
Record W4385923611 · doi:10.1101/2023.08.15.553327

A unified model for interpretable latent embedding of multi-sample, multi-condition single-cell data

2023· preprint· en· W4385923611 on OpenAlexafffund
Ariel Madrigal, Tianyuan Lu, Larisa M. Soto, Hamed S. Najafabadi

Bibliographic record

VenuebioRxiv (Cold Spring Harbor Laboratory) · 2023
Typepreprint
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicSingle-cell and spatial transcriptomics
Canadian institutionsJewish General HospitalUniversity of TorontoMcGill University
FundersCanadian Institutes of Health ResearchAlliance de recherche numérique du Canada
KeywordsComputer scienceNormalization (sociology)Sample (material)Data miningDiscriminative modelEmbeddingGene regulatory networkCovariateLatent variableArtificial intelligenceMachine learningBiologyGene

Abstract

fetched live from OpenAlex

Abstract Analysis of single cells across multiple samples and/or conditions encompasses a series of interrelated tasks, which range from normalization and inter-sample harmonization to identification of cell state shifts associated with experimental conditions. Other downstream analyses are further needed to annotate cell states, extract pathway-level activity metrics, and/or nominate gene regulatory drivers of cell-to-cell variability or cell state shifts. Existing methods address these analytical requirements sequentially, lacking a cohesive framework to unify them. Moreover, these analyses are currently confined to specific modalities where the biological quantity of interest gives rise to a singular measurement. However, other modalities require joint consideration of dual measurements; for example, modeling the latent space of alternative splicing involves joint analysis of exon inclusion and exclusion reads. Here, we introduce a generative model, called GEDI, to identify latent space variations in multi-sample, multi-condition single cell datasets and attribute them to sample-level covariates. GEDI enables cross-sample cell state mapping on par with the state-of-the-art integration methods, cluster-free differential gene expression analysis along the continuum of cell states in the form of transcriptomic vector fields, and machine learning-based prediction of sample characteristics from single-cell data. By incorporating gene-level prior knowledge, it can further project pathway and regulatory network activities onto the cellular state space, enabling the computation of the gradient fields of transcription factor activities and their association with the transcriptomic vector fields of sample covariates. Finally, we demonstrate that GEDI surpasses the gene-centric approach by extending all these concepts to the study of alternative cassette exon splicing and mRNA stability landscapes in single cells.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.007
metaresearch head score (Gemma)0.013
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: Simulation or modeling
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.010
Threshold uncertainty score0.036

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0070.013
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0020.002
Bibliometrics0.0020.002
Science and technology studies0.0010.003
Scholarly communication0.0030.003
Open science0.0040.003
Research integrity0.0030.004
Insufficient payload (model declined to judge)0.0040.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.081
GPT teacher head0.281
Teacher spread0.201 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations2
Published2023
Admission routes2
Has abstractyes

Explore more

Same venuebioRxiv (Cold Spring Harbor Laboratory)Same topicSingle-cell and spatial transcriptomicsFrench-language works237,207