A sensitive and scalable microsatellite instability assay to diagnose constitutional mismatch repair deficiency by sequencing of peripheral blood leukocytes
Bibliographic record
Abstract
Constitutional mismatch repair deficiency (CMMRD) is caused by germline pathogenic variants in both alleles of a mismatch repair gene. Patients have an exceptionally high risk of numerous pediatric malignancies and benefit from surveillance and adjusted treatment. The diversity of its manifestation, and ambiguous genotyping results, particularly from PMS2, can complicate diagnosis and preclude timely patient management. Assessment of low-level microsatellite instability in nonneoplastic tissues can detect CMMRD, but current techniques are laborious or of limited sensitivity. Here, we present a simple, scalable CMMRD diagnostic assay. It uses sequencing and molecular barcodes to detect low-frequency microsatellite variants in peripheral blood leukocytes and classifies samples using variant frequencies. We tested 30 samples from 26 genetically-confirmed CMMRD patients, and samples from 94 controls and 40 Lynch syndrome patients. All samples were correctly classified, except one from a CMMRD patient recovering from aplasia. However, additional samples from this same patient tested positive for CMMRD. The assay also confirmed CMMRD in six suspected patients. The assay is suitable for both rapid CMMRD diagnosis within clinical decision windows and scalable screening of at-risk populations. Its deployment will improve patient care, and better define the prevalence and phenotype of this likely underreported cancer syndrome.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".