MétaCan
Menu
Back to cohort
Record W4294875893 · doi:10.1186/s12920-022-01348-z

TMBur: a distributable tumor mutation burden approach for whole genome sequencing

2022· article· en· W4294875893 on OpenAlexafffund
Emma Titmuss, Richard Corbett, Scott Davidson, Sanna Abbasi, Laura Williamson, Erin Pleasance, Adam Shlien, Daniel J. Renouf, Steven J.M. Jones, Janessa Laskin, Marco A. Marra

Bibliographic record

VenueBMC Medical Genomics · 2022
Typearticle
Languageen
FieldMedicine
TopicCancer Immunotherapy and Biomarkers
Canadian institutionsUniversity of British ColumbiaSickKids FoundationHospital for Sick ChildrenCanada's Michael Smith Genome Sciences Centre
FundersGenome British ColumbiaCanada Foundation for InnovationGenome Canada
KeywordsComputer scienceComputational biologyContext (archaeology)Whole genome sequencingWorkflowGenomeBiologyReplicateData miningGeneticsStatisticsMathematicsDatabase

Abstract

fetched live from OpenAlex

BACKGROUND: Tumor mutation burden (TMB) is a key characteristic used in a tumor-type agnostic context to inform the use of immune checkpoint inhibitors (ICI). Accurate and consistent measurement of TMB is crucial as it can significantly impact patient selection for therapy and clinical trials, with a threshold of 10 mutations/Mb commonly used as an inclusion criterion. Studies have shown that the most significant contributor to variability in mutation counts in whole genome sequence (WGS) data is differences in analysis methods, even more than differences in extraction or library construction methods. Therefore, tools for improving consistency in whole genome TMB estimation are of clinical importance. METHODS: We developed a distributable TMB analysis suite, TMBur, to address the need for genomic TMB estimate consistency in projects that span jurisdictions. TMBur is implemented in Nextflow and performs all analysis steps to generate TMB estimates directly from fastq files, incorporating somatic variant calling with Manta, Strelka2, and Mutect2, and microsatellite instability profiling with MSISensor. These tools are provided in a Singularity container downloaded by the workflow at runtime, allowing the entire workflow to be run identically on most computing platforms. To test the reproducibility of TMBur TMB estimates, we performed replicate runs on WGS data derived from the COLO829 and COLO829BL cell lines at multiple research centres. The clinical value of derived TMB estimates was then evaluated using a cohort of 90 patients with advanced, metastatic cancer that received ICIs following WGS analysis. Patients were split into groups based on a threshold of 10/Mb, and time to progression from initiation of ICIs was examined using Kaplan-Meier and cox-proportional hazards analyses. RESULTS: TMBur produced identical TMB estimates across replicates and at multiple analysis centres. The clinical utility of TMBur-derived TMB estimates were validated, with a genomic TMB ≥ 10/Mb demonstrating improved time to progression, even after correcting for differences in tumor type (HR = 0.39, p = 0.012). CONCLUSIONS: TMBur, a shareable workflow, generates consistent whole genome derived TMB estimates predictive of response to ICIs across multiple analysis centres. Reproducible TMB estimates from this approach can improve collaboration and ensure equitable treatment and clinical trial access spanning jurisdictions.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.769
Threshold uncertainty score0.771

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0010.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0010.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.032
GPT teacher head0.272
Teacher spread0.240 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations7
Published2022
Admission routes2
Has abstractyes

Explore more

Same venueBMC Medical GenomicsSame topicCancer Immunotherapy and BiomarkersFrench-language works237,207