OmniPert: A Deep Learning Foundation Model for Predicting Responses to Genetic and Chemical Perturbations in Single Cancer Cells
Bibliographic record
Abstract
In cancer, intra- and inter-patient heterogeneity presents a significant challenge for therapeutic management, as patients with apparently similar profiles often exhibit divergent responses to the same therapies. This heterogeneity is primarily attributed to genetic and molecular variations among individuals and their tumors. Understanding the impact of these differences on treatment outcomes is widely believed to be a key step for developing effective precision medicine strategies. However, the complexity of most biological pathways makes it difficult to predict the effect of genetic variation on cells and tissues, let alone predict a patient's response to therapy. As a result, high-throughput genetic and chemical perturbation screens have emerged as valuable tools for precision medicine-related tasks, such as disease modeling, target discovery, cellular programming, and pathway reconstruction. This approach is fundamentally limited, however, because the number of possible combinations of cell types, cell states, perturbation targets, and perturbation types is huge and cannot be exhaustively tested experimentally. This calls for computational approaches that can simulate such experiments in silico, guiding in vitro experiments towards perturbations that are more likely to produce the desired effect. Here we describe OmniPert, a novel generative AI tool, which utilizes a deep learning, transformer-based architecture to model the effects of genetic and chemical perturbations on single-cell transcriptomes. Trained on millions of diverse cellular profiles, this approach allows for more granular analysis of cellular responses, thereby facilitating downstream applications in cell-specific gene-gene and gene-drug interaction networks, biomarker and drug target discovery, drug repurposing, and in silico perturbation reverse-engineering. In the context of oncology, OmniPert promises to facilitate the discovery of novel cell type- and state-specific targets, ultimately contributing to more effective and personalized cancer treatments.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".