Abstract P5-02-04: Upfront dichotomous histopathological assessment of ductal carcinoma in situ of the breast to reduce inter-observer variability: The DCISion study
Bibliographic record
Abstract
Abstract Background. Ductal carcinoma in situ (DCIS) of the breast is considered to be a non-obligatory precursor of invasive breast cancer. Its histopathological assessment is characterized by considerable inter-observer discordance. In a previous study, we performed post hoc dichotomization of multi-categorical variables to determine the ‘ideal’ cut-offs for dichotomous histopathological assessment. In the present multicenter study, inter-observer variability is evaluated among 39 pathologists who performed upfront dichotomous evaluation of a consecutive series of 149 DCIS. Methods. All participants were board-certified pathologists with a special interest in breast disease. The participants assessed at least 50 primary oncologic breast cancer resection specimens per year, in accordance with the EUSOMA criteria for dedicated breast pathologists. No training set was used. Instead, a written guideline with associated DCISion poster was provided, which contained all definitions of the histopathological features of interest. Representative digital slides of 149 DCIS were accessible via an online platform. All pathologists independently assessed the following histopathological features: nuclear atypia, necrosis, solid DCIS architecture, calcifications, stromal architecture and lobular cancerization. Stromal inflammation was assessed semi-quantitatively. Stromal tumor-infiltrating lymphocytes (TILs) were quantified as percentages and were also dichotomously assessed with a cut-off at 50%. Krippendorff’s alpha (KA), Cohen’s kappa (K) and intraclass correlation coefficient (ICC) were calculated for the appropriate variables. Results. Intraductal calcifications (KA 0,676) and solid DCIS architecture (KA 0,602) were characterized by the highest inter-observer concordance. Stromal inflammation (KA 0,564), dichotomously assessed TILs (KA 0,520) and comedonecrosis (KA 0,539) showed slightly higher inter-observer agreement. Lobular cancerization (KA 0,396), nuclear atypia (KA 0,422) and stromal architecture (KA 0,450) showed the lowest inter-observer concordance. Assessment of TILs as a percentage showed good overall agreement with a mean ICC of 0,821 (range 0,566 - 0,933). Semi-quantitative assessment of stromal inflammation (KA 0,564) resulted in somewhat lower inter-observer variability than upfront dichotomous TILs assessment (KA 0,520). High stromal inflammation corresponded best with dichotomously assessed TILs when the TILs cut-off was set at 10% (K 0,881). Nevertheless, a post hoc TILs cut-off set at 20% resulted in the highest inter-observer agreement (KA 0,669). Experience and time dedicated to breast pathology did not influence the degree of concordance. Conclusion. The DCSion study shows that, despite upfront dichotomous evaluation, the inter-observer variability remains considerable and is at most acceptable. Nevertheless, the discordance rate varies among the different histopathological features. Future studies should investigate its impact on DCIS risk stratification. Differences in prognostic value among the different methods to quantify TILs are of particular interest, since inter-observer variability may partly explain different outcomes among different studies. Artificial intelligence might be able to tackle this diagnostic challenge. Development of deep learning algorithms could result in more objective histopathological assessment. Although machine learning might represent the next “pathologist’s best friend”, we should be careful not to introduce inter-observer variability into these deep learning algorithms. The DCISion study therefore provides an excellent setting to investigate the value of such algorithms in rendering the final diagnosis more robust. Citation Format: Mieke Rosalie Van Bockstal, Hélène Dano, Serdar Altinay, Laurent Arnould, Noella Bletard, Cecile Colpaert, Franceska Dedeurwaerdere, Benjamin Dessauvagie, Valérie Duwel, Giuseppe Floris, Stephen Fox, Clara Gerosa, Shabnam Jaffer, Eline Kurpershoek, Magali Lacroix-Triki, Andoni Laka, Kathleen Lambein, Gaëtan Marie MacGrogan, Caterina Marchió, Dolores Martin Martinez, Sharon Nofech-Mozes, Dieter Peeters, Alberto Ravarino, Emily Reisenbichler, Erika Resetkova, Souzan Sanati, Anne-Marie Schelfhout, Vera Schelfhout, Abeer M Shaaban, Renata Sinke, Claudia Maria Stanciu-Pop, Claudia Stobbe, Carolien HM van Deurzen, Koen Van de Vijver, Anne-Sophie Van Rompuy, Stephanie Verschuere, Anne Vincent-Salomon, Hannah Wen, Caroline Bouzin, Christine Galant. Upfront dichotomous histopathological assessment of ductal carcinoma in situ of the breast to reduce inter-observer variability: The DCISion study [abstract]. In: Proceedings of the 2019 San Antonio Breast Cancer Symposium; 2019 Dec 10-14; San Antonio, TX. Philadelphia (PA): AACR; Cancer Res 2020;80(4 Suppl):Abstract nr P5-02-04.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.031 | 0.049 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".