97 Statistical and systematic uncertainties affect biomarker reliability: a case study in multiplex immunofluorescence
Notice bibliographique
Résumé
Background Every measurement comes with statistical uncertainties, arising from the finite data sample, and systematic uncertainties, from errors in the measurement or model. Biomedical analyses typically account only for the statistical uncertainty from the number of patients. These are reported as p-values, power calculations, and other similar metrics.As datasets grow and this statistical uncertainty shrinks, it is not necessarily the dominant uncertainty anymore. A closer look at other uncertainties that affect analyses is critically needed.1 Methods We processed mIF slides from four cohorts containing a total of n=75 pre-treatment specimens from patients with NSCLC using the AstroPath platform, 2 which applies numerous image corrections and produces comprehensive cell maps.We identified CD8+FoxP3+ cells as a positive predictor of response to anti-PD1 immunotherapy in non-small-cell lung cancer (NSCLC). We further developed the Diversity of Niches Unlocking Treatment Sensitivity (DONUTS) biomarker. The DONUTS are niches that mimic the neighborhoods around CD8+FoxP3+ cells but are 150x more common than the cells themselves.3 4 We evaluated both biomarkers’ performance and calculated the statistical uncertainties resulting both from the finite number of patients and from the finite number of cells or DONUTS within each patient‘s biopsy.We processed our data through the AstroPath pipeline again, removing some of the corrections. We estimated the systematic uncertainties that would have been present had we not applied the corrections and the remaining systematic uncertainties resulting from imperfect corrections.We developed the ROC Picker package5 to propagate these uncertainties to ROC and Kaplan-Meier curves.Results With the relatively small cohort sizes (largest: n=25) used in this analysis, the statistical uncertainty from the number of patients still dominates, but other uncertainties contribute appreciably.The statistical uncertainty from the number of CD8+FoxP3+ cells is comparable to the statistical uncertainty from the number of patients (figure 1A), but the statistical uncertainty from the more common DONUTS is much smaller (figure 1B).The systematic uncertainties were also substantial before the corrections in the pipeline, but dropped by more than half after applying those corrections.Conclusions We demonstrate a proof of concept for evaluating statistical and systematic uncertainties in mIF. While standard practice is to consider only one statistical uncertainty, other uncertainties are becoming more important as datasets grow and, if not accounted for, will impact biomarkers and clinical decision making. 6 Our methodology can be applied to biomarkers from all data modalities, ensuring that they remain reliable in the era of big data.References Berry S, Giraldo N, Green B, et al. Analysis of multispectral imaging with the AstroPath platform informs efficacy of PD-1 blockade. Science. 2021;372.Green B, et al. AstroPath Pipeline v0.1.0. Zenodo. 2021.Cottrell T, Roskes J, Cohen E, et al. Early, effector CD8+FoxP3+ cells and their topology associate with outcomes in patients with non-small cell lung carcinoma (NSCLC) receiving neoadjuvant anti-PD-1-based therapy. In preparation. 2025.Cohen E, et al. CD8+FoxP3+ cells represent early, effector T-cells and predict outcomes in patients with resectable non-small cell lung carcinoma (NSCLC) receiving neoadjuvant anti-PD-1-based therapy. J Immunother Cancer. 2022;10:A63.Roskes J. ROC Picker v1.1.0. GitHub. 2024.Taube J, Sunshine J, et al. Society for immunotherapy of cancer: updates and best practices for multiplex immunohistochemistry (IHC) and immunofluorescence (IF) image analysis and data sharing. J Immunother Cancer. 2025;13.Abstract 97 Figure 1Kaplan-Meier curves for regression-free survival for patients with NSCLC, stratified by density of (A) CD8+FoxP3+ cells or (B) DONUTs. The error bands show the statistical error resulting from the finite number of cells or DONUTs
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,070 | 0,255 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,001 |
| Méta-épidémiologie (sens large) | 0,001 | 0,002 |
| Bibliométrie | 0,002 | 0,002 |
| Études des sciences et des technologies | 0,002 | 0,003 |
| Communication savante | 0,003 | 0,002 |
| Science ouverte | 0,002 | 0,001 |
| Intégrité de la recherche | 0,003 | 0,002 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,001 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».