Response to Zhang and Yang
Bibliographic record
Abstract
We thank Drs Zhang and Yang for their comments regarding Ki67 immunohistochemical assessment. We would like to emphasize that we are not claiming there is stronger concordance in absolute Ki67 index scores when they are below 5% or above 30%, but rather that those values are sufficiently different from prognostic decision cutpoints that, despite the residual analytical variability that remains after standardization, results are sufficient for the purpose of deciding on the need for adjuvant chemotherapy in anatomically favorable ER-positiveand HER2-negative patients. To illustrate this, consider the scores reported in Figure 1, republished from Leung et al. (1). For all cases with at least 1 laboratory reporting a score of 5% or less, the median values of all laboratories in these cases range from 4% to 8%, which is below typical clinical thresholds (10%-15%). Similarly, in this same data set, the cases with at least 1 laboratory reporting a score of 30% or greater, the median values of all laboratories in these cases ranges from 19% to 88%. These conclusions not only are based on the data from our own group but also are supported by several other published studies (2,3). Immunohistochemical testing for Ki67 remains a very practical and widely available test applicable in many jurisdictions, some of which do not have ready access to more complex testing modalities. Despite its limitations in quantitative precision, when technical quality assurance is in place and scoring is standardized, this test retains real value for patient care in specific clinical contexts. Heat map of Ki67 scores. Rows represent cases and columns represent laboratories. The color green indicates that the score is less than 10%, yellow 10% to 20%, and red greater than 20%. Cases are ordered by the median scores (across laboratories), which are shown in parentheses beside the specimen number. Laboratories are ordered (within each group) by the median scores (across cases). The 3 colon-separated numbers to the right of the table represent the number of laboratories giving scores falling into different ranges: less than 10% (left-most), 10% to 20% (middle), and greater than 20% (right-most). For example, 15:6:1 indicates that 15 laboratories gave a score of less than 10%, 6 laboratories from 10% to 20%, and 1 laboratory greater than 20%. Republished from Leung et al. (1). Please refer to the original article (1) for the color version of this figure. Heat map of Ki67 scores. Rows represent cases and columns represent laboratories. The color green indicates that the score is less than 10%, yellow 10% to 20%, and red greater than 20%. Cases are ordered by the median scores (across laboratories), which are shown in parentheses beside the specimen number. Laboratories are ordered (within each group) by the median scores (across cases). The 3 colon-separated numbers to the right of the table represent the number of laboratories giving scores falling into different ranges: less than 10% (left-most), 10% to 20% (middle), and greater than 20% (right-most). For example, 15:6:1 indicates that 15 laboratories gave a score of less than 10%, 6 laboratories from 10% to 20%, and 1 laboratory greater than 20%. Republished from Leung et al. (1). Please refer to the original article (1) for the color version of this figure. None. Role of the funder: Not applicable. Disclosures: Torsten O. Nielsen (T.O.N.) received royalty from NanoString Technologies. T.O.N. has intellectual property rights and hold patent with Bioclassifier LLC. Mitch Dowsett received lecture fees from Nanostring and Myriad; participated in advisory/consultancy role with Radius, Lilly, AbbVie, H3 Biomedicine and Zentalis. His institution received grants from Pfizer and Lilly on studies that includes Ki67 analysis. Daniel F. Hayes (D.F.H.) reports research support from Menarini Silicon BioSystems (MSB). The University of Michigan (UM) holds patent US 8,790,878 B2 for which D.F.H. is designated as inventor, and that is licensed to MSB with annual royalties paid to UM and D.F.H. Outside the submitted work D.F.H. holds stock options from OncImmune LLC, InBiomotion, and serves on advisory boards for Cepheid, Freenome, CellWorks, Agendia, Salutogenic, EPIC Sciences and L-Nutra and UM receives research funding on his behalf from Merrimack, Eli Lilly, Puma Biotechnology, Pfizer, AstraZeneca. The remaining authors have no conflicts of interest to disclose. Author contributions: Torsten O. Nielsen contributed to conceptualization and writing—original draft, review and editing. Samuel C.Y. Leung contributed to conceptualization and writing—original draft, review and editing. Lisa M. McShane contributed to conceptualization and writing—original draft, review and editing. Mitch Dowsett contributed to conceptualization and writing—original draft review and editing. Daniel F. Hayes contributed to conceptualization and writing—original draft, review and editing. Not applicable.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.017 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.003 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".