Is “Polysomy” in Breast Carcinoma the “New Equivocal” in<i>HER2</i>Testing?
Bibliographic record
Abstract
It is reasonable to assume that the ultimate purpose of evidence-based consensus guidelines in pathology practice is to improve patient care. When the American Society of Clinical Oncology and the College of American Pathologists (ASCO/CAP) convened an expert panel to consider recommendations for HER2 testing in breast cancer, the HER2 testing landscape was more a mosaic of random color and form than a coherent composition. Despite using common immunohistochemistry (IHC) and in situ hybridization (ISH) tools, each with defined interpretative criteria, diagnostic laboratories were unable to achieve levels of interlaboratory concordance that could provide a meaningful level of confidence in the reliability of HER2 test results. This lack of clear concordance was complicated by overall positive rates of up to 30% (despite an expected rate closer to 12%–16%), and this separation of actual performance from expected values reflected both unacceptably high levels of false-positive and false-negative results.1,2 The first published ASCO/CAP guideline recommendations, although perhaps a flawed redefinition of the original HER2 IHC interpretative criteria, largely reiterated interpretative guidelines created and validated for initial companion testing for trastuzumab clinical validation trials while incorporating separately approved interpretative schemes for HER2 and dual HER2 /Cep17 ISH modalities.2 The improvement in laboratory performance following release of these guidelines in 2007 was thus less likely related to interpretive guidelines than to those elements of testing that related to preanalytic factors, tissue selection, assay validation, and emphasis on reflex testing for equivocal results. Whatever the basis, substantial improvement in testing performance did occur: overall positive rates aligned more closely with expected values, false-positive and false-negative rates decreased, and concordance both between laboratories and (when comparing different testing modalities) within laboratories also improved.1 With the release of updates to these original recommendations, the ASCO/CAP panel took pains to create a testing …
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.013 | 0.029 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.004 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.002 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.010 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".