Quantitative Justification of the Change From 10% to 30% for Human Epidermal Growth Factor Receptor 2 Scoring in the American Society of Clinical Oncology/College of American Pathologists Guidelines: Tumor Heterogeneity in Breast Cancer and Its Implications for Tissue Microarray–Based Assessment of Outcome
Bibliographic record
Abstract
PURPOSE: The variability in scoring of immunohistochemistry, whether a result of true heterogeneity or artifacts in preparation, has led to decreased reliability in companion diagnostics and the recommendation for new standards (eg, the American Society of Clinical Oncology/College of American Pathologists [ASCO-CAP] guidelines). The basis of this problem is the amount of tissue required to be representative of an entire tumor. Because protein expression on tissue microarrays (TMAs) can be rigorously measured and one 0.6-mm spot is equivalent to two to three high-power fields, we used TMAs to assess levels of heterogeneity and to determine optimal representation as a function of outcome. PATIENTS AND METHODS: We analyzed estrogen receptor (ER), progesterone receptor, and human epidermal growth factor receptor 2 (HER-2) expression in two cohorts (n = 676 and n = 152) on a series of four to five separate TMA cores and assessed heterogeneity by linear regression analysis. Minimum, average, and maximum scores were generated for each set, which were then assessed for prognostic and predictive value. RESULTS: Each marker shows some heterogeneity, but average r values between 0.7 and 0.8 are seen between TMA spots. Analysis for prognostic value shows that the highest maximum score (of five spots) is the most prognostic for ER, whereas a high HER-2 minimum score is most prognostic for poor outcome and most predictive of response to trastuzumab. CONCLUSION: These results suggest that the representivity required for each biomarker may be a function of its role in tumorigenesis. Furthermore, these results provide scientific basis for the ASCO-CAP guidelines for assessment of HER-2 expression but perhaps suggest that the 30% figure is still too conservative.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.013 | 0.008 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".