A novel federated learning framework for medical imaging: Resource‐efficient approach combining PCA with early stopping
Notice bibliographique
Résumé
BACKGROUND: Federated learning (FL) facilitates collaborative model training across multiple institutions while preserving privacy by avoiding the sharing of raw data, a critical consideration in medical imaging applications. Despite its potential, FL faces challenges such as high-dimensional data, heterogeneity among datasets from different centers, and resource constraints, which limit its efficiency and effectiveness in healthcare settings. PURPOSE: This study aims to present a novel adaptive FL framework to address the challenges of data heterogeneity and resource constraints in medical imaging. The proposed framework is designed to optimize computational efficiency, enhance training processes, improve model performance, and ensure robustness against non-independent and identically distributed (non-IID) data across decentralized data sources. METHODS: The proposed adaptive FL framework addresses the challenges of high-dimensional data and heterogeneity in nonuniform and decentralized data sources through a key innovation. First, Federated incremental principal component analysis (FIPCA) achieves privacy-preserving dimensionality reduction by aggregating local scatter matrices and means from participating centers, enabling the computation of a global PCA model. This process ensures data alignment across centers, mitigates heterogeneity, and significantly reduces computational complexity. We evaluated the framework's ability to generalize across institutions in a cross-site classification task distinguishing clinically significant prostate cancer (csPCa) from non-csPCa. This assessment used 1500 T2-weighted (T2W) prostate MRI images from three institutions, where two centers (800 + 350 cases) were used for training and validation, and one center (350 cases) served as an independent test site. RESULTS: The proposed method significantly reduced the number of global training rounds from 200 to 38, achieving a 98% reduction in energy consumption compared to the standard FedAvg algorithm. The effective use of FIPCA for dimensionality reduction enhanced generalizability, while adaptive early stopping prevented overfitting, leading to an improvement in model performance, with the area under the curve (AUC) on the unseen test center increasing from 0.68 to 0.73 (95 % CI 0.70 - 0.77) on the test center's data. Additionally, the method demonstrated improved sensitivity and specificity, indicating superior classification performance. The integration of FIPCA accelerated convergence by reducing data dimensionality, while the adaptive early-stopping mechanism further optimized resource utilization and prevented overfitting. CONCLUSIONS: Our adaptive FL approach efficiently handles large, heterogeneous medical imaging data, reducing training time and computational overhead, while improving model accuracy. The substantial reduction in energy consumption and accelerated convergence make it suitable for real-world healthcare settings.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,006 | 0,009 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,001 |
| Méta-épidémiologie (sens large) | 0,002 | 0,001 |
| Bibliométrie | 0,001 | 0,001 |
| Études des sciences et des technologies | 0,001 | 0,002 |
| Communication savante | 0,002 | 0,003 |
| Science ouverte | 0,004 | 0,003 |
| Intégrité de la recherche | 0,002 | 0,002 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,002 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».