Unveiling cell-type-specific mode of evolution in comparative single-cell expression data
Bibliographic record
Abstract
While methodology for determining the mode of evolution in coding sequences has been well established, evaluation of adaptation events in emerging types of phenotype data needs further development. Here, we propose an analysis framework (expression variance decomposition, EVaDe) for comparative single-cell expression data based on phenotypic evolution theory. After decomposing the gene expression variance into separate components, we use two strategies to identify genes exhibiting large between-taxon expression divergence and small within-cell-type expression noise in certain cell types, attributing this pattern to putative adaptive evolution. In a dataset of primate prefrontal cortex, we find that such human-specific key genes enrich with neurodevelopment-related functions, while most other genes exhibit neutral evolution patterns. Specific neuron types are found to harbor more of these key genes than other cell types, thus likely to have experienced more extensive adaptation. Reassuringly, at the molecular sequence level, the key genes are significantly associated with the rapidly evolving conserved non-coding elements. An additional case analysis comparing the naked mole-rat (NMR) with the mouse suggests that innate-immunity-related genes and cell types have undergone putative expression adaptation in NMR. Overall, the EVaDe framework may effectively probe adaptive evolution mode in single-cell expression data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".