52 Digital pathology training effectiveness for the evaluation of PD-L1 expression in multiple tumor indications
Bibliographic record
Abstract
<h3>Background</h3> In-person pathologist trainings during the COVID-19 pandemic became impossible, necessitating a shift to remote-digital whole slide image (WSI) training. High concordance between WSI and glass slide scores from the same specimens stained with PD-L1 IHC 22C3 pharmDx (SK006) across multiple tumor indications supported the validity of digital training.<sup>1</sup> However, in-person microscope (glass-slide) training versus remote-digital (WSI) training effectiveness must be assessed. Collated testing data on specimens (SK006 stained) spanning multiple indications scored by external pathologists during Agilent led training and testing (T&T) sessions via glass slides were compared to sessions utilizing WSIs. <h3>Methods</h3> Stained slides (30 unique specimens per tumor indication) were scanned on an Aperio AT2 scanner to generate WSIs for digital T&T. Remote T&T sessions used WebEx and PathcoreScholar’s online platform to discuss scoring guidelines and WSI training cases. Subsequently, external pathologists evaluated WSIs in PathcoreScholar for PD-L1 expression using either Tumor Proportion Score (TPS) or Combined Positive Score (CPS) scoring algorithms and interpreted these scores at predefined cutoffs (figure 1). In both glass and WSI scoring test modalities, passing is defined as inter and intra-observer overall agreement (OA) ≥85%. Training effectiveness pass rates from glass slide data (2018–2020) and WSI data (2021–2022) spanning multiple indications and scoring algorithms were calculated and then compared using the Fisher-Freeman-Halton test, with a significance threshold of 0.05. Only data from initial pathologist tests were included in the pass rate calculation; data from re-tests executed after initial test failure were excluded. <h3>Results</h3> The differences between pass rates for microscope (glass slide) and digital (WSI) testing were not statistically significant (p-value > 0.05) (tables 1 and 2). Testing pass rates for indications scored with TPS or CPS using microscope glass slide vs digital WSI T&T was not statistically significant (p-value > 0.05) (table 3). <h3>Conclusions</h3> No statistically significant differences in pathologist training effectiveness for PD-L1 were observed between remote and in-person trainings across multiple tumor indications, scoring algorithms, and cutoffs. These results demonstrate the effectiveness and equivalency of remote-digital pathologist trainings for evaluation of PD-L1 expression as detected by PD-L1 IHC 22C3 pharmDx in multiple tumor indications when compared to in-person-microscope glass slide T&T. Use of digital training and scoring proficiency testing can provide pathologists around the world with access to high-quality, interactive training from leading experts in PD-L1 expression evaluation. <h3>Acknowledgements</h3> We would like to thank our colleagues at Agilent Technologies, Inc. and all the pathologists who completed Agilent scoring certification training and testing for their valuable contributions to this study. Tissue samples were provided by the Cooperative Human Tissue Network which is funded by the National Cancer Institute. Other investigators may have received specimens from the same subjects. Tissue samples supplied by BioIVT (Hicksville, NY, USA). The data and biospecimens used in this project were provided by Centre Hospitalier Universitaire (CHU) de Nice (Nice, France), US Biolab (Gaithersburg, MD, USA), Contract Research Ltd (Charlestown, Nevis), Centre Hospitalier Universitaire (CHU) de Nice (Nice, France), IOM Ricera (Viagrande, Italy), National BioService LLC (Saint Petersburg, Russia), SageBio LLC (Sharon, MA, USA, Tumorothèque Régionale de Franche-Comté (Besançon, France), Centre Antoine Lacassagne (CAL; Nice, France, GLAS (Winston-Salem, NC, USA), Maine Medical, Hospices Civils de Lyon (Lyon, France), Sofia Bio LLC (New York, NY, USA), SELARL DIAG (Nice, France), and Clin-Path Diagnostics (Tempe, AZ, USA) with appropriate ethics approval and through Azenta Life Sciences. Biological materials were provided by the Ontario Tumour Bank, which is supported by the Ontario Institute for Cancer Research (Toronto, Ontario, Canada) through funding provided by the Government of Ontario. <h3>Reference</h3> Adams M, Moquin D, Littrell J, et al 36 Digital Whole Slide Image (WSI) scoring is equivalent to microscope glass slide scoring for evaluation of programmed death-ligand 1 (PD-L1) expression across multiple tumor indications. <i>Journal for ImmunoTherapy of Cancer</i> 2021;<b>9</b>: doi: 10.1136/jitc-2021-SITC2021.036.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".