Abstract IA03: Accelerating oncology drug discovery with the power of microscopy & AI
Bibliographic record
Abstract
Abstract Cell images contain a vast amount of quantifiable information about the status of the cell: for example, whether it is diseased, whether it is responding to a drug treatment, or whether a pathway has been disrupted by a genetic mutation. We aim to go beyond measuring individual cell phenotypes that biologists already know are relevant to a particular disease. Instead, in a strategy called image-based profiling, often using the Cell Painting assay, we extract hundreds of features of cells from microscopy images. Just like transcriptional or proteomic profiling, the similarities and differences in the patterns of extracted features reveal connections among diseases, drugs, and genes, with many applications in cancer research. In fact, these strategies underpin drug discovery platform companies such as Recursion and SyzOnc. Because images are inexpensive and high-throughput, we can carry out experiments at very large scale, yielding single-cell profiles for hundreds of thousands of samples through public-private consortia (JUMP, OASIS, VISTA, NIH IGVF) and pooled barcode-based optical screens. Cell morphology is therefore a powerful data source for cancer systems biology, alongside molecular omics. Citation Format: Anne E. Carpenter. Accelerating oncology drug discovery with the power of microscopy & AI [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Artificial Intelligence and Machine Learning; 2025 Jul 10-12; Montreal, QC, Canada. Philadelphia (PA): AACR; Clin Cancer Res 2025;31(13_Suppl):Abstract nr IA03.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.007 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.005 | 0.002 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.075 | 0.037 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".