Abstract P1-03-28: Agreement between the DESTINY-Breast04/06 VENTANA 4B5 HER2 IHC clinical trial assay and other comparator assays for HER2-low breast cancer: Overall results of a large-scale, multicenter global ring study
Bibliographic record
Abstract
Abstract Background: Treatment with trastuzumab deruxtecan improved outcomes compared with standard of care for human epidermal growth factor receptor 2 (HER2)-low (immunohistochemistry [IHC] 1+ or IHC 2+ with in situ hybridization negative) metastatic breast cancer (BC) in DESTINY-Breast04 and DESTINY-Breast06. The VENTANA Pathway 4B5 IHC assay, approved in the US as a companion diagnostic, was used in both trials. This global ring study assessed concordance between VENTANA Pathway 4B5 and comparator assays (CAs) in identifying HER2-low BC, with the first phase of the study including laboratories in the US, Canada, and Europe. Concordance in the first phase varied, with positive percent agreement (PPA) tending to be high, especially with 4B5 laboratory-developed tests (LDT). Here, we report the overall results combining the first and second phases of the ring study. Methods: 50 clinical BC samples were chosen from a cohort of 300 by a steering committee of expert pathologists. Samples were stained using VENTANA Pathway 4B5 and scored as HER2 IHC 0, 1+, 2+, and 3+ by a central laboratory and a panel of experts before being sent to laboratories for HER2 IHC testing per American Society of Clinical Oncology-College of American Pathologists (ASCO-CAP) 2018 guidelines. Laboratories were actively scoring HER2 IHC for BC in a clinical setting, had two independent pathologists, and did not routinely use VENTANA Pathway 4B5. The second phase analyzed data from laboratories in Australia, New Zealand, Brazil, Chile, China, Hong Kong, Taiwan, Malaysia, and the Philippines. Pathologists first scored samples with their routine protocols and assays (HercepTest [Omnis or Link48], Leica Oracle, non-4B5 LDT, or 4B5 LDT); then, following virtual alignment on interpretation of HER2 IHC scoring guidelines, they rescored the samples 2 weeks later. Postalignment scores were compared with the reference VENTANA Pathway 4B5 scores. The primary endpoint was PPA and negative percent agreement (NPA) for HER2-low versus HER2 IHC 0 (includes both ≤10% faint, incomplete membrane staining and no membrane staining) based on postalignment scoring. Results: A combined 6580 scores from 135 pathologists at 70 laboratories were recorded before virtual alignment. Of these, 129 pathologists from 68 laboratories received alignment guidance, and 6270 postalignment scores were available for analysis. Following alignment, PPA (agreement in identifying HER2-low), NPA (agreement in identifying HER2 IHC 0), and overall agreement were 84.8% (95% CI, 83.6-86.0), 69.2% (95% CI, 67.0-71.2), and 79.4% (95% CI, 78.3-80.5), respectively, with Cohen’s κ of 0.54 (corresponding to moderate agreement). Virtual alignment did not substantially affect PPA, NPA, or overall agreement (prealignment scores were 85.1%, 69.5%, and 79.7%, respectively). Postalignment, PPA by assay type ranged from 61.6% (Leica Oracle, N = 196) to 95.5% (HercepTest Omnis, N = 467) and NPA ranged from 36.9% (HercepTest Omnis) to 81.7% (Leica Oracle). PPA by regional subgroup ranged from 62.0% (Latin America, N = 335) to 95.6% (France, N = 391), while NPA ranged from 52.8% (Europe [Other], N = 776) to 89.0% (Australia/New Zealand, N = 395). Conclusions: Interassay and interlaboratory variability was observed in concordance between VENTANA Pathway 4B5 and CAs in identifying HER2 low versus HER2 IHC 0. PPA tended to be higher than NPA, suggesting more consistent detection of HER2-low compared to HER2 IHC 0. Awareness of available novel treatment options, deliberate pathologist training, and optimization of analytical assay methods and their choice are indicated for accurate identification of patients with clinically actionable HER2 expression levels.= Citation Format: Sunil Badve, Corrado D’Arrigo, Gelareh Farshid, Annette Lebeau, Vicente Peg, Fréderique Penault-Llorca, Josef Rueschoff, Wentao Yang, Neil Atkey, Jessica Baumann, Anika Altenfeld, Elisabeth Beyerlein, Amy Hanlon Newell, Alexander Penner, Akira Moh, Giuseppe Viale. Agreement between the DESTINY-Breast04/06 VENTANA 4B5 HER2 IHC clinical trial assay and other comparator assays for HER2-low breast cancer: Overall results of a large-scale, multicenter global ring study [abstract]. In: Proceedings of the San Antonio Breast Cancer Symposium 2024; 2024 Dec 10-13; San Antonio, TX. Philadelphia (PA): AACR; Clin Cancer Res 2025;31(12 Suppl):Abstract nr P1-03-28.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.014 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".