Fnih MRD in AML Biomarkers Consortium: Evaluation of Methods for Detection of FLT3 internal Tandem Duplication Measurable Residual Disease in Acute Myeloid Leukemia
Bibliographic record
Abstract
Introduction FLT3 internal tandem duplication (ITD) is one of the most frequently observed mutations in acute myeloid leukemia (AML) and persistent detection in complete remission (CR) prior to allogeneic hematopoietic transplant (alloHCT) is highly predictive of increased incidence of relapse and death. A variant allele fraction (VAF) of ≥0.01% is associated with the greatest risk, but there is increasing evidence that some of this risk can be mitigated by intervention. Objective There is an immediate need for commercially available assays for FLT3 -ITD measurable residual disease (MRD) detection. The FNIH MRD in AML Biomarkers Consortium here benchmarked technologies for FLT3 -ITD MRD detection. Methods Assay characteristics from five next-generation sequencing (NGS)-based assays for sensitive FLT3 -ITD detection, available commercially as a service or kit, were collected from the providers. In phase 1, 48 contrived standards were generated by serially diluting DNA from FLT3 -ITD cell lines into control DNA in isolation or multiplexed (VAF 0.01-1%). The top four performing assays were carried forward to phase 2, where testing of DNA from peripheral blood or bone marrow of 16 AML patients with FLT3 -ITD(s) (12-192bp) detected by capillary gel electrophoresis at a low allelic ratio (0.01-0.18) and four controls was performed. All testing was performed blinded. Results The five NGS-based technologies evaluated utilized a variety of enrichment strategies, error-correction methods, and sequencing platforms ( Figure 1 ). In phase 1 testing of standards, four technologies reliably detected all ITDs down to 0.05-0.01% VAF, while 2 ITDs were missed by one technology due to primer design and bioinformatic filtering. Significant correlation between expected and observed VAFs was seen for three technologies ( Figure 2) , while one technology exhibited conflicting results related to ITD length. In all cases, the presence of multiple ITDs did not impact performance. In phase 2 testing of primary samples, all technologies correctly assigned patients as FLT3 -ITD positive or negative. While 18 of 20 ITDs were identified by all four assays, one technology missed 2 complex ITDs due to bioinformatic filtering. A significant correlation between expected and observed VAF was observed for all but one method. Additional analyses including further assay optimization and comparison to reported assay performance will be presented. Conclusion The evaluation of NGS-based technologies for FLT3 -ITD MRD detection showed the ability of multiple strategies to detect and quantify ITDs of different lengths down to a clinically relevant level. However, multiple factors, both technical and bioinformatic, can impact assay performance. Given test heterogeneity, careful consideration of the performance characteristics of any assay proposed for clinical decision making prior to alloHCT is recommended.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.012 | 0.011 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".