Abstract A036: Genomic alterations associate with distinct tumor dynamics in NSCLC: A genomics-informed Bayesian hierarchical modeling approach
Bibliographic record
Abstract
Abstract Understanding and predicting individual tumor dynamics is critical for personalizing cancer treatment, yet clinical application of mathematical models remains limited due to sparse longitudinal data and failure to integrate biological features beyond tumor size. We hypothesized that genomic alterations could inform patient-specific tumor dynamics and developed an integrated computational framework to test this relationship. We analyzed 844 non-small cell lung cancer (NSCLC) patients treated with Osimertinib who underwent MSK-IMPACT gene profiling (505 cancer-related genes tested for mutations, copy number alterations, and rearrangements). Longitudinal tumor measurements were extracted from radiology reports using our validated natural language processing (NLP) algorithm. We modeled tumor dynamics using an ordinary differential equation (ODE) system capturing sensitive and resistant cell populations, with parameters estimated through a novel genomics-informed Bayesian hierarchical framework. This approach clusters patients based on genomic alteration profiles and prioritizes information from patients with dense measurements (≥4 timepoints) to inform parameter estimation in patients with sparse data. The latter mitigates the risk of identifiability issues in patients with sparse measurements. The NLP method achieved 95.2% accuracy in tumor size detection and 94.9% accuracy in location matching against 62 manually annotated cases. Principal component and K-means analysis identified four patient groups with distinct alteration frequencies in TP53, RB1, CTNNB1, SMAD4, and PIK3CA (Fisher’s exact test with Benjamini-Hochberg correction; all p < 0.0001). Hierarchical parameter estimation of the ODE model revealed significant inter-cluster differences in growth rates (Kruskal-Wallis, H=9.01 and 10.83, both p<0.05), treatment response (Kruskal-Wallis, H=21.04, p<0.001), and the estimated initial resistant cell fraction (H=26.92, p<0.001). In particular, clusters enriched in TP53 alterations had higher treatment response rates; however, they tended towards a higher resistance growth rate compared to other genomic subgroups. Our integrated framework combining automated tumor size extraction with genomics-informed hierarchical modeling uncovered distinct tumor dynamics across genomically defined NSCLC subgroups. This approach lays the foundation for the prediction of treatment response trajectories from baseline molecular data—potentially facilitating personalized lung cancer treatments and follow-up guidance based on patient-specific tumor trajectories. Citation Format: Nikolaos M. Dimitriou, Jonas Willmann, Edward Christopher. Dee, Karl Pichotta, David Ma, Sohrab Salehi, Kevin Boehm, Puneeth Iyengar, Francisco Sanchez-Vega, Sohrab P. Shah, Nikolaus Schultz, Jian Carrot-Zhang. Genomic alterations associate with distinct tumor dynamics in NSCLC: A genomics-informed Bayesian hierarchical modeling approach [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Artificial Intelligence and Machine Learning; 2025 Jul 10-12; Montreal, QC, Canada. Philadelphia (PA): AACR; Clin Cancer Res 2025;31(13_Suppl):Abstract nr A036.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".