Clinical-Grade Genomic Analysis of a 99-patient Cohort for the Identification of Therapeutic Targets
Bibliographic record
Abstract
Several large-scale studies have reported the application of holistic bioinformatics pipeline approaches for the comprehensive detection of genomic drivers, routinely utilizing whole-genome sequencing (WGS) and whole-exome sequencing (WES). However, a general pipeline methodology for efficient analysis of small-scale genomics studies, utilizing targeted genomic sequencing (TGS) remains to be accomplished. Here we report the application of Illumina’s TruSight 170 panel, 5 database aggregations (ClinVar, COSMIC, TCGA, GTEX and DGidb) and basic computational resources upon a 99-patient cohort, with the aim of describing the underlying causative mechanisms of cancer. The pipeline initially reported a gene pool of 133 genes across all 13 cancer variations. After extensive DNA mutation profiling TP53, CHEK2, BARD1, ATM and NOTCH1 were determined to be the most frequently mutated within the dataset. Additional, clinical significance analysis revealed the pathogenesis of all such genes resulting in further stratification of highly pathogenic genes, which included TP53. Individual variants were analyzed to determine the frequency and distributions of such variants across the human genome, via COSMIC. MAFtools analysis was conducted to investigate the somatic mutational landscape of the dataset consisting of patient cohort patterns, mutation classes and genomic driver co-occurrence. More than 95% of genes co-occur within the dataset, indicating the application of double-target agents. Multiple oncogenic mechanisms and single-target drug propositions include loss of heterozygosity - Olaparib/Veliparib/Rucaparib (BRCA1/2) and RTK-RAS pathway - Afatinib/Lazertinib (EGFR). Multiple double-target driver combination propositions included BRCA1 - BRCA2, MSH3 - KMT2A and TP53 - BARD1. The development of such efficient genomic pipeline frameworks allows for rapid genomic driver detection, ultimately leading to drug discovery, prognostic modeling and a higher quality of life for cancer patients.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".