Disparity between Ki67 measurements and tumor gene expression tests in patients with hormone-sensitive early breast cancer from the OPTIMA preliminary trial.
Bibliographic record
Abstract
567 Background: Tumor gene expression tests are increasingly used in breast cancer management. The Ki67 biomarker has been proposed as an inexpensive alternative for making chemotherapy decisions, has demonstrated utility for determination of endocrine therapy responsiveness and is included in the FDA license for adjuvant abemaciclib. We have compared Ki67 measurements with tumor gene expression test results for patients included in the OPTIMA prelim trial. Methods: We compared Ki67 %staining with the results of Oncotype DX, Prosigna and MammaPrint performed by the test vendor. Ki67 was determined in a single laboratory on triplicate tissue micro-arrays using quantitative image analysis including a 10% manual quality control check. We used kappa statistics to measure agreement between tests, divided into groups using the pre-defined test score boundaries for high vs. not high risk. Results: Data were available for 259 patients. Using ≥20% staining to define a high Ki67 score, kappa values (95% CI) for agreement with Prosigna were: 0.39 (0.28-0.49); Oncotype DX: 0.27 (0.18-0.36); and MammaPrint: 0.38 (0.27-0.49). Kappa values <0.2 are conventionally interpreted as showing slight agreement and 0.21-0.4 as fair agreement. A detailed breakdown of the comparisons of Ki67 with Prosigna and Oncotype DX is tabulated. Conclusions: Agreement between Ki67 and tumor gene expression tests is limited. Therefore, Ki67 values cannot accurately be used to reflect any of the molecular scores assessed here, all of which are well validated prognostic biomarkers. The use of Ki67 to determine suitability for adjuvant chemotherapy requires validation before it can replace the existing tests. Tumor gene expression tests may prove superior to Ki67 for the identification of patients likely to benefit from adjuvant abemaciclib. OPTIMA prelim is registered as ISRCTN42400492 and funded by the UK NIHR Health Technology Assessment Programme, award number 10/34/01. Clinical trial information: ISRCTN42400492. [Table: see text]
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".