New standards in HER2-low testing: the CASI-01 comparative methods study
Bibliographic record
Abstract
BACKGROUND: The introduction of Trastuzumab deruxtecan (T-Dxd) has exposed clinically significant limitations in accurately detecting HER2-low expression testing when using immunohistochemistry (IHC) assays originally developed to detect HER2 over-expression. While HER2 testing is widely used to determine T-Dxd eligibility, no HER2-low assay was ever validated against HER2 protein expression. METHODS: To address this pressing need, the Consortium for Analytic Standardization in Immunohistochemistry (CASI) conducted the CASI-01 study, involving 54 IHC laboratories across Europe and the U.S. The study aimed to identify optimal assay conditions for accurate HER2 testing, differentiating between HER2 overexpression (3+) for Trastuzumab eligibility and HER2-low expression (1+ or ultra-low) for T-Dxd eligibility. The conventional FDA-cleared HER2 assay ("predicate") was compared with higher-sensitivity assays using pathologist versus image analysis readouts. HER2 overexpression was validated against HER2 gene amplification via in situ hybridisation (ISH), while HER2-low accuracy was evaluated using newly introduced HER2 reference standards and a novel IHC parameter-dynamic range. FINDINGS: CASI-01 revealed variability in predicate HER2 assays, with detection thresholds ranging from 30,000 to 60,000 among laboratories. Despite this variability, these assays demonstrated high accuracy for identifying HER2 overexpression (3+), with 85.7% (18/21) sensitivity (95% confidence limits 63.66-96.95%) and 100% (49/49) specificity (95% confidence limits 92.75-100%), though sensitivity may have been limited by the use of older tissue specimens, with loss or reduced expression levels of the HER2 protein. However, these same assays exhibited poor dynamic range for detecting HER2-low scores. Enhanced analytic sensitivity of IHC assays combined with image analysis overcame this limitation with HER2-low scores, achieving a six-fold improvement (p = 0.0017). INTERPRETATION: IHC assays with detection thresholds in the range of 30,000-60,000 HER2 molecules per cell yield accurate results for determination of Trastuzumab eligibility (HER2 3+) but fail to demonstrate the dynamic range for accurate HER2-low scores. Enhanced analytic sensitivity of HER2 assays combined with image analysis addresses this critical gap in HER2-low testing. More generally, CASI-01 introduces pivotal advancements in precision medicine: (a) the importance of reporting IHC analytic sensitivity and ability to demonstrate an assay dynamic range, and (b) image analysis can surpass pathologist readout accuracy in specific clinical contexts. FUNDING: This work was supported by the National Cancer Institute of the National Institutes of Health under Award Number R44CA268484.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.003 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".