Abstract A043: Mutational profiling and machine learning for risk stratification and biomarker identification in intraductal papillary mucinous neoplasms progressing to pancreatic cancer
Bibliographic record
Abstract
Abstract Intraductal papillary mucinous neoplasms (IPMNs) are common precursors to invasive pancreatic ductal adenocarcinoma (PDAC), with the risk of progression varying by IPMN type and anatomic location. Despite some IPMNs having a high risk of progression, there are limited non-surgical options for precision prevention in patients with IPMNs. One of the main challenges is that existing radiologic and molecular markers are insufficient for reliably assessing the risk of progression, mostly due to a lack of validated intervention targets. Currently, cancer prevention for patients with IPMN is centered on surgery or surveillance using a risk-based strategy; the clinical ability to stratify risk of cancer progression of individual IPMN tumors is poor and essentially no effective non-surgical interventions exist. Thus, the objective of this work is to apply statistical features extraction and use latent features derived from sequencing data as input to machine learning (ML) models to identify unexplored markers driving the progression of IPMNs to invasive PDAC. Data consisting of formalin-fixed, paraffin-embedded tissue cores sampled from 34 unique patients with the following pathological diagnoses: 8 non-IPMN-derived PDAC, 7 IPMN-derived PDAC, 7 high-grade IPMNs, and 12 low-grade IPMNs, were analyzed using the Moffitt STAR 2.0 Cancer Mutation and Molecular Biomarker Profiling panel. This STAR 2.0 next generation sequencing method performed using the TruSight Oncology 500 panel from Illumina, Inc., is designed to interpret sequence information for over 500 somatically altered genes. The analyses of sequencing data incorporated statistical and ML analyses of the mutational profiles of patients’ genomes followed by integration of trinucleotide sequence-derived mutational features. By applying non-negative matrix factorization to DNA trinucleotide motif mutational data, we identified 4 distinct mutational signatures. These signatures exhibited varying degrees of similarity to the Single Base Substitution Signatures from COSMIC (Sanger Institute). However, the contribution of these signatures to the mutational profile in each sample is more complex, as samples may carry a combination of signatures rather than a single, defining signature. Preliminary results from the discriminative ML models showed high performance in predicting type of malignancy (multiclass area under the curve > 0.8), using the mutational counts and each sample-to-signature contribution. We demonstrate that insights extracted from mutational profiles have the potential to enhance the interpretation of mutational patterns and improve the stratification of IPMNs. Further analyses are needed to fully understand the complex interplay of mutational processes across the samples, beyond the initial identification of signatures. We aim to uncover key gene patterns, focusing on codon mutations and their mapping to protein alterations, to better understand the molecular mechanisms driving the progression of IPMNs and identify potential therapeutic targets for early intervention. Citation Format: Aleksandra Karolak, Evan W. Davis, Mouktik Isukapalli, Rohit Veligeti, Margaret A. Park, Jamie K. Teer, Daniel K. Jeong, Kun Jiang, Dung-Tsa Chen, Jennifer B. Permuth, Ghulam Rasool. Mutational profiling and machine learning for risk stratification and biomarker identification in intraductal papillary mucinous neoplasms progressing to pancreatic cancer [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Artificial Intelligence and Machine Learning; 2025 Jul 10-12; Montreal, QC, Canada. Philadelphia (PA): AACR; Clin Cancer Res 2025;31(13_Suppl):Abstract nr A043.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".