Formalin fixation increases deamination mutation signature but should not lead to false positive mutations in clinical practice
Bibliographic record
Abstract
Genomic analysis of cancer tissues is an essential aspect of personalized oncology treatment. Though it has been suggested that formalin fixation of patient tissues may be suboptimal for molecular studies, this tissue processing approach remains the industry standard. Therefore clinical molecular laboratories must be able to work with formalin fixed, paraffin embedded (FFPE) material. This study examines the effects of pre-analytic variables introduced by routine pathology processing on specimens used for clinical reports produced by next-generation sequencing technology. Tissue resected from three colorectal cancer patients was subjected to 2, 15, 24, and 48 hour fixation times in neutral buffered formalin. DNA was extracted from all tissues twice, once with uracil-N-glycosylase (UNG) treatment to counter deamination effects, and once without. Of note, deamination events at methylated cytosine, as found at CpG sites, remains unaffected by UNG. After extraction a two-step PCR targeted sequencing method was performed using the Illumina MiSeq and the data was analyzed via a custom-built bioinformatics pipeline, including filtration of reads with mapping quality <30. A larger baseline group of samples (n = 20) was examined to establish if there was a sample performance difference between the two DNA extraction methods, with/without UNG treatment. There was no statistical difference between sequencing performance of the two extraction methods when comparing read counts (raw, mapped, and filtered) and read quality (% mapped, % filtered). Analyzing mutation type, there was no significant difference between mutation calls until the 48 hour fixation treatment. At 48 hours there is a significant increase in C/G->T/A mutations that is not represented in DNA treated with UNG. This suggests these errors may be due to deamination events triggered by a longer fixation time. However the allelic frequency of these events remained below the limit of detection for reportable mutations in this assay (<2%). We do however recommend that suspected intratumoral heterogeneity events be verified by re-sequencing the same FFPE block.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.009 | 0.028 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".