Personalized genome assembly for accurate cancer somatic mutation discovery using cancer-normal paired reference samples
Bibliographic record
Abstract
Abstract The use of personalized genome assembly as a reference for detecting the full spectrum of somatic events from cancers has long been advocated but never been systematically investigated. Here we address the critical need of assessing the accuracy of somatic mutation detection using personalized genome assembly versus the standard human reference assembly (i.e. GRCh38). We first obtained massive whole genome sequencing data using multiple sequencing technologies, and then performed de novo assembly of the first tumor-normal paired genomes, both nuclear and mitochondrial, derived from the same donor with triple negative breast cancer. Compared to standard human reference assembly, the haplotype phased chromosomal-scale personalized genome was best demonstrated with individual specific haplotypes for some complex regions and medical relevant genes. We then used this well-assembled personalized genome as a reference for read mapping and somatic variant discovery. We showed that the personalized genome assembly results in better alignments of sequencing reads and more accurate somatic mutation calls. Direct comparison of mitochondrial genomes led to discovery of unreported nonsynonymous somatic mutations. Our findings provided a unique resource and proved the necessity of personalized genome assembly as a reference in improving somatic mutation detection at personal genome level not only for breast cancer reference samples, but also potentially for other cancers.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".