An international reproducibility study validating quantitative determination of ERBB2, ESR1, PGR, and MKI67 mRNA in breast cancer using MammaTyper®
Bibliographic record
Abstract
BACKGROUND: Accurate determination of the predictive markers human epidermal growth factor receptor 2 (HER2/ERBB2), estrogen receptor (ER/ESR1), progesterone receptor (PgR/PGR), and marker of proliferation Ki67 (MKI67) is indispensable for therapeutic decision making in early breast cancer. In this multicenter prospective study, we addressed the issue of inter- and intrasite reproducibility using the recently developed reverse transcription-quantitative real-time polymerase chain reaction-based MammaTyper® test. METHODS: Ten international pathology institutions participated in this study and determined messenger RNA expression levels of ERBB2, ESR1, PGR, and MKI67 in both centrally and locally extracted RNA from formalin-fixed, paraffin-embedded breast cancer specimens with the MammaTyper® test. Samples were measured repeatedly on different days within the local laboratories, and reproducibility was assessed by means of variance component analysis, Fleiss' kappa statistics, and interclass correlation coefficients (ICCs). RESULTS: Total variations in measurements of centrally and locally prepared RNA extracts were comparable; therefore, statistical analyses were performed on the complete dataset. Intersite reproducibility showed total SDs between 0.21 and 0.44 for the quantitative single-marker assessments, resulting in ICC values of 0.980-0.998, demonstrating excellent agreement of quantitative measurements. Also, the reproducibility of binary single-marker results (positive/negative), as well as the molecular subtype agreement, was almost perfect with kappa values ranging from 0.90 to 1.00. CONCLUSIONS: On the basis of these data, the MammaTyper® has the potential to substantially improve the current standards of breast cancer diagnostics by providing a highly precise and reproducible quantitative assessment of the established breast cancer biomarkers and molecular subtypes in a decentralized workup.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".