Pan-tumor harmonization of pathologic response assessment for standardized data collection in neoadjuvant IO trials (PATHdata): Final analysis of a multi-institutional reproducibility study.
Bibliographic record
Abstract
2515 Background: Immunotherapeutic agents are now being investigated for treating earlier-stage cancers. Radiographic assessment by RECIST, widely used to assess treatment response in clinical trials for advanced cancers, has limitations in the neoadjuvant setting; and pathologic response assessment is increasingly being used as a primary and/or secondary endpoint. To that end, a pan-tumor scoring system for assessing pathologic response was developed (1,2). This scoring system allows for the quantitative assessment of residual viable tumor (RVT) in multiple locations: i.e. primary and lymph node (LN) or distant metastases, akin to RECIST. %RVT scored using this system been associated with patient outcomes after treatment with anti-PD-1-based therapies. Additionally, %RVT in LN has been shown to have additive value to %RVT in the primary tumor when predicting patient survival (3). As a result, pathologists are now being asked to score pathologic response in the primary tumor and LN as a part of ongoing clinical trials and routine clinical care. Here, we evaluated the reproducibility of %RVT scoring using pan-tumor immune-related pathologic response criteria (irPRC). Methods: A multi-institutional, international study led by the Society for Immunotherapy of Cancer was initiated to assess the concordance of pathologic response assessment in resection specimens from patients treated with anti-PD-1-based therapies. Online lecture-based modules for irPRC scoring were developed, and 14 pathologists from multiple institutions, including academic and industry partners, were trained to score H&E-stained slides. The pathologists have scored n=37 pathology cases from resection specimens and on-treatment biopsies from >10 different tumor types, in part derived from phase II/III clinical trials. %RVT in the primary tumor and LN from patient specimens were scored separately (total of n=374 slides scored by each pathologist). Results: At the first interim analysis, scoring of pathologic response using irPRC was shown to be highly reproducible, irrespective of disease location (i.e. primary tumor vs lymph node metastasis). The second half of the study is nearing completion, and these reproducibility numbers will be finalized and presented in the final abstract. Extended analyses will also be presented that include subset analyses by tumor type. Conclusions: The results will be interpreted and presented in the context of the larger field for pathologic response assessment. A post-study survey completed by the participating pathologists will be used to refine irPRC training materials prior to dissemination to the wider immuno-oncology community. 1. Cottrell et al. Ann Oncol2018. 2. Stein et al. Clin Can Res 2020. 3. Deutsch, et al. Nat Med 2023.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.336 | 0.355 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.005 |
| Bibliometrics | 0.004 | 0.007 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.004 | 0.002 |
| Open science | 0.003 | 0.007 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.006 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".