MétaCan
Menu
Retour à la cohorte
Enregistrement W4415900298 · doi:10.1136/jitc-2025-sitc2025.0097

97 Statistical and systematic uncertainties affect biomarker reliability: a case study in multiplex immunofluorescence

2025· article· W4415900298 sur OpenAlexaff
Jeffrey S Roskes, Benjamin Green, Emily B. Cohen, Margaret Eminizer, Sam J Tabrisky, Sigfredo Soto-Diaz, Boyang Zhang, Daphne Wang, Daniel Jiménez‐Sánchez, Justina X. Caushi, Jiajia Zhang, Nina M. D’Amiano, Joel Sunshine, J.S. Deutsch, Sonali Uttam, Alexa Fiorante, Nicole Espinosa, Teodora Popa, Aleksandra Ogurtsova, Andrew Jorquera, Jamie E. Chaft, Julie R. Brahmer, Michael Conroy, Joshua E. Reuss, Hongkai Ji, Patrick M. Forde, Drew M. Pardoll, Kellie N. Smith, Alexander S. Szalay, Janis M. Taube, Tricia R. Cottrell

Notice bibliographique

RevueRegular and Young Investigator Award Abstracts · 2025
Typearticle
Langue
DomaineDecision Sciences
ThématiqueReliability and Agreement in Measurement
Établissements canadiensQueen's University
Organismes subventionnairesnon disponible
Mots-clésMultiplexBiomarkerImmunofluorescenceAffect (linguistics)

Résumé

récupéré en direct d'OpenAlex

Background Every measurement comes with statistical uncertainties, arising from the finite data sample, and systematic uncertainties, from errors in the measurement or model. Biomedical analyses typically account only for the statistical uncertainty from the number of patients. These are reported as p-values, power calculations, and other similar metrics.As datasets grow and this statistical uncertainty shrinks, it is not necessarily the dominant uncertainty anymore. A closer look at other uncertainties that affect analyses is critically needed.1 Methods We processed mIF slides from four cohorts containing a total of n=75 pre-treatment specimens from patients with NSCLC using the AstroPath platform, 2 which applies numerous image corrections and produces comprehensive cell maps.We identified CD8+FoxP3+ cells as a positive predictor of response to anti-PD1 immunotherapy in non-small-cell lung cancer (NSCLC). We further developed the Diversity of Niches Unlocking Treatment Sensitivity (DONUTS) biomarker. The DONUTS are niches that mimic the neighborhoods around CD8+FoxP3+ cells but are 150x more common than the cells themselves.3 4 We evaluated both biomarkers’ performance and calculated the statistical uncertainties resulting both from the finite number of patients and from the finite number of cells or DONUTS within each patient‘s biopsy.We processed our data through the AstroPath pipeline again, removing some of the corrections. We estimated the systematic uncertainties that would have been present had we not applied the corrections and the remaining systematic uncertainties resulting from imperfect corrections.We developed the ROC Picker package5 to propagate these uncertainties to ROC and Kaplan-Meier curves.Results With the relatively small cohort sizes (largest: n=25) used in this analysis, the statistical uncertainty from the number of patients still dominates, but other uncertainties contribute appreciably.The statistical uncertainty from the number of CD8+FoxP3+ cells is comparable to the statistical uncertainty from the number of patients (figure 1A), but the statistical uncertainty from the more common DONUTS is much smaller (figure 1B).The systematic uncertainties were also substantial before the corrections in the pipeline, but dropped by more than half after applying those corrections.Conclusions We demonstrate a proof of concept for evaluating statistical and systematic uncertainties in mIF. While standard practice is to consider only one statistical uncertainty, other uncertainties are becoming more important as datasets grow and, if not accounted for, will impact biomarkers and clinical decision making. 6 Our methodology can be applied to biomarkers from all data modalities, ensuring that they remain reliable in the era of big data.References Berry S, Giraldo N, Green B, et al. Analysis of multispectral imaging with the AstroPath platform informs efficacy of PD-1 blockade. Science. 2021;372.Green B, et al. AstroPath Pipeline v0.1.0. Zenodo. 2021.Cottrell T, Roskes J, Cohen E, et al. Early, effector CD8+FoxP3+ cells and their topology associate with outcomes in patients with non-small cell lung carcinoma (NSCLC) receiving neoadjuvant anti-PD-1-based therapy. In preparation. 2025.Cohen E, et al. CD8+FoxP3+ cells represent early, effector T-cells and predict outcomes in patients with resectable non-small cell lung carcinoma (NSCLC) receiving neoadjuvant anti-PD-1-based therapy. J Immunother Cancer. 2022;10:A63.Roskes J. ROC Picker v1.1.0. GitHub. 2024.Taube J, Sunshine J, et al. Society for immunotherapy of cancer: updates and best practices for multiplex immunohistochemistry (IHC) and immunofluorescence (IF) image analysis and data sharing. J Immunother Cancer. 2025;13.Abstract 97 Figure 1Kaplan-Meier curves for regression-free survival for patients with NSCLC, stratified by density of (A) CD8+FoxP3+ cells or (B) DONUTs. The error bands show the statistical error resulting from the finite number of cells or DONUTs

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,070
score de la tête « metaresearch » (Gemma)0,255
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesMétarecherche
Catégories consensuellesaucune
DomaineSignal candidat: Méthodes · Signal consensuel: aucune
Devis d'étudeSignal candidat: Observationnel · Signal consensuel: Observationnel
GenreSignal candidat: Empirique · Signal consensuel: Empirique
Score de désaccord entre enseignants0,930
Score d'incertitude au seuil0,368

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0700,255
Méta-épidémiologie (sens strict)0,0010,001
Méta-épidémiologie (sens large)0,0010,002
Bibliométrie0,0020,002
Études des sciences et des technologies0,0020,003
Communication savante0,0030,002
Science ouverte0,0020,001
Intégrité de la recherche0,0030,002
Charge utile insuffisante (le modèle a refusé de juger)0,0010,000

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,069
Tête enseignante GPT0,344
Écart entre enseignants0,275 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Devis d'étudeObservationnel
DomaineMéthodes
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2025
Routes d'admission1
Résumé présentnon

Explorer davantage

Même revueRegular and Young Investigator Award AbstractsMême sujetReliability and Agreement in MeasurementTravaux en français237 207