Abstract PO-049: Pseudo-alignment resolution in discriminating mouse reads
Bibliographic record
Abstract
Abstract Patient-Derived Xenografts (PDX) are the best preclinical models to study cancer biology preclinical drug screening. PDXs are generated by implanting a piece of fresh cancer tissue directly into an immunodeficient mouse to grow and imitate a human tumor. Previous studies have shown that PDXs are able to accurately recapitulate the (epi-)genomic characteristics of the original tumor. Yet the inclusion of the mouse cells in tumor biopsies can introduce unwanted noise and bias in high throughput (epi-)genomics analyses and interpretation. Thus, filtering the mouse reads is one of the first and important steps in analyzing high throughput PDX sequencing data. There are two main approaches to account for mouse reads in analyzing bulk RNAseq data generated from PDXs. In the first approach, an external tool is used to classify the raw or aligned reads into a human, mouse, or ambiguous reads. Examples include XenoFilter, BBsplit, disambiguate, and Xenome. In the second approach, a mixed human-mouse reference genome is used when aligning the raw reads. Multiple studies have shown that the mixed reference approach produces better or similar results compared to when an external tool is used for mouse read discrimination. However, these comparisons have employed exact alignment methods when using a mixed reference approach. Recently, the significant deluge of sequencing data has brought attention to pseudo-alignment based methods, specifically Kallisto and Salmon. By taking advantage of pseudo-alignment to a reference transcriptome, these tools are capable of analyzing millions of reads in minutes while demonstrating near-optimal quantification performance. Yet their performance in analyzing PDX RNAseq data while using a mixed reference is unclear. Moreover, among the external methods, only Xenome has a pseudo-alignment strategy for read classification and the rest of the methods use exact alignment making them labor-intensive and intractable for large datasets. In this study, we compare two main mouse read discrimination pipelines when using Kallisto and Salmon. By sampling from two human and mouse lung RNAseq datasets, we have simulated five pseudo-PDX RNAseq datasets with different contamination ratios. By comparing outputs of different discriminating strategies to the ground truth, we show that Kallisto and Salmon when used with mixed human-mouse transcriptome produce results that are more correlated with the ground truth. Moreover, by differential gene expression analyses we show that fewer genes are significantly affected when using the mixed reference approach. Citation Format: Soheil Jahangiri-Tazehkand, Farnoosh Agha-Babazadeh, Benjamin Haibe-Kains. Pseudo-alignment resolution in discriminating mouse reads [abstract]. In: Proceedings of the AACR Virtual Special Conference on Artificial Intelligence, Diagnosis, and Imaging; 2021 Jan 13-14. Philadelphia (PA): AACR; Clin Cancer Res 2021;27(5_Suppl):Abstract nr PO-049.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.010 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".