Abstract B003: Towards machine learning fairness in glioblastoma: An evaluation of protected attributes in publicly available clinical datasets
Notice bibliographique
Résumé
Abstract Introduction: As machine learning (ML) algorithms are increasingly developed for clinical applications, there are growing concerns over the real-world benefits of ML-assisted decision-making applications. These issues include reduced generalizability across medical institutions, lack of clinician uptake during algorithmic deployment, and observed disparate performances across various demographics, including race, gender, and socioeconomic status. Glioblastoma (GBM) is a rare brain cancer with poor outcomes and few publicly available datasets, resulting in limited opportunities for algorithmic validation and fairness evaluation. As ML research calls for improved fairness assessments and improved cross-institutional validation, we set out to characterize publicly available GBM datasets and explore fairness evaluations enabled by the protected attribute clinical data available and/or missing in these datasets. Methods: We identified and assessed 16 publicly available glioma and/or GBM clinical datasets (with non-overlapping patient cohorts) for protected attribute availability and potential issues inhibiting ML fairness methods. We also investigated patient treatment timelines and longitudinal sample pairing to assess dataset shift, cross-institutional variability, and ability to assess disease recurrence. Results and Discussion: Datasets contained an average of 343 patients (range 26-1480), with imaging (n=12) and genomic (n=11) data most commonly included, followed by histopathology (n=6) and transcriptomic (n=4) data. Treatment timelines were reported for 8 datasets, with patient treatment data spanning an average of of 9.9 years (range 2-24). Age and sex/gender variables were available across 88% of datasets (n=14), with self-reported race (56%, n=9) and ethnicity (38%, n=6) attributes less commonly available. Three datasets reported patients from multiple countries and only one dataset reported patient insurance type, indicating challenges in evaluating cross-regional generalizability and performance across socioeconomic status. Datasets with limited racial/ethnic information may result in "fairness through unawareness" evaluation approaches, which have demonstrated disparate impacts on protected groups in other ML domains. Given documented sex differences in GBM and "negative legacy" issues amongst sampling minoritized racial/ethnic groups, our results demonstrate barriers towards applying current ML bias mitigation methods. Conclusion: With increased calls for ML projects to publish and validate their algorithms on publicly available datasets, we indicate gaps between current fairness methodologies and the clinical data attribute landscape of publicly available GBM clinical data. Our results discuss how current fairness methodologies can be applied to existing datasets or may be limited by attribute availability. As clinical algorithms are increasingly developed for GBM applications, we advocate for further enrichment of publicly available datasets with socioeconomic, regional, and genetic ancestry data for improved fairness assessment. Citation Format: Shreya Chappidi, Andra V. Krauze. Towards machine learning fairness in glioblastoma: An evaluation of protected attributes in publicly available clinical datasets [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Artificial Intelligence and Machine Learning; 2025 Jul 10-12; Montreal, QC, Canada. Philadelphia (PA): AACR; Clin Cancer Res 2025;31(13_Suppl):Abstract nr B003.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,070 | 0,160 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,002 |
| Bibliométrie | 0,002 | 0,002 |
| Études des sciences et des technologies | 0,002 | 0,002 |
| Communication savante | 0,004 | 0,003 |
| Science ouverte | 0,002 | 0,004 |
| Intégrité de la recherche | 0,002 | 0,002 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,003 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».