MétaCan
Menu
Back to cohort
Record W4412078570 · doi:10.1038/s41467-025-60466-1

Towards fair decentralized benchmarking of healthcare AI algorithms with the Federated Tumor Segmentation (FeTS) challenge

2025· article· en· W4412078570 on OpenAlexafffund
Maximilian Zenk, Ujjwal Baid, Sarthak Pati, Akis Linardos, Brandon Edwards, Micah Sheller, Patrick Foley, Alejandro Aristizábal, David Zimmerer, А. Д. Груздев, Jason Martin, Russell T. Shinohara, Annika Reinke, Fabian Isensee, Santhosh Parampottupadam, Ralf Floca, Hasan Kassem, Bhakti Baheti, Siddhesh Thakur, Verena Chung, Kaisar Kushibar, Karim Lekadir, Meirui Jiang, Youtan Yin, Hongzheng Yang, Quande Liu, Cheng Chen, Qi Dou, Pheng‐Ann Heng, Xiaofan Zhang, Shaoting Zhang, Muhammad Irfan Khan, Mohammad Ayyaz Azeem, Mojtaba Jafaritadi, Esa Alhoniemi, Elina Kontio, Suleiman A. Khan, Leon Mächler, Ivan Ezhov, Florian Kofler, Suprosanna Shit, Johannes C. Paetzold, Timo Loehr, Benedikt Wiestler, Himashi Peiris, Kamlesh Pawar, Shenjun Zhong, Zhaolin Chen, Munawar Hayat, Gary F. Egan, Mehrtash Harandi, Ece Isik-Polat, Görkem Polat, Altan Koçyiğit, Alptekin Temizel, Anup Tuladhar, Lakshay Tyagi, Raissa Souza, Nils D. Forkert, Pauline Mouchès, Matthias Wilms, Vishruth Shambhat, Akansh Maurya, Shubham Subhas Danannavar, Rohit Kalla, Vikas Kumar Anand, Ganapathy Krishnamurthi, Sahil Nalawade, Chandan Ganesh, Benjamin Wagner, Divya Reddy, Yudhajit Das, Fang Yu, Baowei Fei, Ananth J. Madhuranthakam, Joseph A. Maldjian, Gaurav Singh, Jianxun Ren, Wei Zhang, Ning An, Qingyu Hu, Youjia Zhang, Ying Zhou, Vasilis Siomos, Giacomo Tarroni, Jonathan Passerat‐Palmbach, Ambrish Rawat, Giulio Zizzo, Swanand Kadhe, Jonathan P. Epperlein, Stefano Braghin, Yuan Wang, Renuga Kanagavelu, Qingsong Wei, Yechao Yang, Yong Liu, Krzysztof Kotowski, Szymon Adamski, Bartosz Machura, Wojciech Malara, Łukasz Zarudzki, Jakub Nalepa, Yaying Shi, Hongjian Gao, Salman Avestimehr, Yonghong Yan, Agus Subhan Akbar, Ekaterina Kondrateva, Hua Yang, Zhaopei Li, Hung-Yu Wu, Johannes Roth, Camillo Saueressig, Alexandre Milesi, Quôc Dinh Nguyên, Nathan J Gruenhagen, Tsung-Ming Huang, Jun Ma, Har Shwinder H Singh, Nai-Yu Pan, Dingwen Zhang, Ramy A. Zeineldin, Michal Futrega, Yading Yuan, Gian Marco Conte, Xue Feng, Quan D. Pham, Yong Xia, Zhifan Jiang, Huan Minh Luu, Mariia Dobko, Alexandre Carré, Bair Tuchinov, Hassan Mohy‐ud‐Din, Saruar Alam, Anup K. Singh, Nameeta Shah, Weichung Wang, Chiharu Sako, Michel Bilello, Satyam Ghodasara, Suyash Mohan, Christos Davatzikos, Evan Calabrese, Jeffrey D. Rudie, Javier Villanueva‐Meyer, Soonmee Cha, Christopher P. Hess, John Mongan, Madhura Ingalhalikar, Manali Jadhav, Umang Pandey, Jitender Saini, Raymond Y. Huang, Ken Chang, Minh‐Son To, Sargam Bhardwaj, Chee Chong, Marc Agzarian, Michal Kozubek, Filip Lux, Jan Michálek, Petr Matula, Miloš Keřkovský, Tereza Kopřivová, Marek Dostál, Václav Vybíhal, Marco C. Pinho, James Holcomb, Marie‐Christin Metz, Rajan Jain, Matthew Lee, Yvonne W. Lui, Pallavi Tiwari, Ruchika Verma, Rohan Bareja, Ipsa Yadav, Jonathan Chen, Neeraj Kumar, Yuriy Gusev, Krithika Bhuvaneshwar, Anousheh Sayah, Camelia Bencheqroun, Anas Belouali, Subha Madhavan, Rivka R. Colen, Aikaterini Kotrotsou, Gianluca Brugnara, Chandrakanth Jayachandran Preetha, Felix Sahm, Martin Bendszus, Wolfgang Wick, Abhishek Mahajan, Carmen Balañá, Jaume Capellades, Josep Puig, Yoon Seong Choi, Seung‐Koo Lee, Jong Hee Chang, Sung Soo Ahn, Hassan M. Fathallah‐Shaykh, Alejandro Herrera-Trujillo, María Trujillo, William Escobar, Ana Lorena Abello, José Bernal, Jhon Gómez, Pamela LaMontagne, Daniel S. Marcus, Mikhail Milchenko, Arash Nazeri, Bennett A. Landman, Karthik Ramadass, Kaiwen Xu, Silky Chotai, Lola B. Chambless, Akshitkumar M. Mistry, Reid C. Thompson, Ashok Srinivasan, Jayapalli Rajiv Bapuraj, Arvind Rao, Nicholas Wang, Yoshiaki Ota, Toshio Moritani, Sevcan Türk, Joonsang Lee, Snehal Prabhudesai, John W. Garrett, Matthew Larson, Robert Jeraj, Hongwei Li, Tobias Weiß, Michael Weller, Andrea Bink, Bertrand Pouymayou, Sonam Sharma, Tzu-Chi Tseng, Saba Adabi, Alexandre X. Falcão, Samuel Botter Martins, Bernardo Corrêa de Almeida Teixeira, F Sprenger, David Menotti, Diego Rafael Lucio, Simone P. Niclou, Olivier Keunen, Ann‐Christin Hau, Enrique Peláez, Heydy Sailé Franco Maldonado, Francis R. Loayza, Sebastián Quevedo, Richard McKinley, Johannes Slotboom, Piotr Radojewski, Raphaël Meier, Roland Wiest, Johannes Trenkler, Josef Pichler, Georg Necker, Andreas Haunschmidt, Stephan Meckel, Pamela Guevara, Esteban Torche, Cristóbal Mendoza, Franco Vera, Elvis Ríos, Eduardo López, Sergio A. Velastín, Joseph Choi, Stephen Baek, Yusung Kim, Heba Ismael, Bryan G. Allen, John M. Buatti, P Zampakis, Vasileios Panagiotopoulos, Panagiotis Tsiganos, Sotiris Alexiou, Ilias Haliassos, Evangelia I. Zacharaki, Κωνσταντίνος Μουστάκας, Christina Kalogeropoulou, Dimitrios Kardamakis, Bing Luo, Laila Poisson, Ning Wen, Martin Vallières, Mahdi Ait Lhaj Loutfi, David Fortin, Martin Lepage, Fanny Morón, Jacob Mandel, Gaurav Shukla, Spencer Liem, Joseph S. Lombardo, Joshua D. Palmer, Adam E. Flanders, Adam P. Dicker, Godwin Ogbole, Dotun Oyekunle, Olubunmi Odafe-Oyibotha, Babatunde Osobu, Mustapha Shu'aibu Hikima, Mayowa Soneye, Farouk Dako, Adeleye Dorcas, Derrick Murcia, Eric Fu, Rourke Haas, John A. Thompson, D. Ryan Ormond, Stuart Currie, Kavi Fatania, Russell Frood, Amber L. Simpson, Jacob Peoples, Ricky Hu, Danielle Cutler, Fábio Ynoe de Moraes, Anh Tran, Mohammad Hamghalam, Michael A. Boss, James F. Gimpel, Deepak Kattil Veettil, Kendall Schmidt, Lisa Cimino, Cynthia Price, Brian Bialecki, Sailaja Marella, Charles Apgar, András Jakab, Marc‐André Weber, Errol Colak, Jens Kleesiek, John Freymann, Justin Kirby, Jake Albrecht, Peter Mattson, Alexandros Karargyris, Prashant Shah, Bjoern Menze, Klaus Maier‐Hein, Spyridon Bakas

Bibliographic record

VenueNature Communications · 2025
Typearticle
Languageen
FieldMedicine
TopicRadiomics and Machine Learning in Medical Imaging
Canadian institutionsUniversity of TorontoMcGill University Health CentreCentre Hospitalier Universitaire de SherbrookeQueen's UniversityUniversité de SherbrookeAlberta Children's HospitalHotchkiss Brain InstituteUniversity of AlbertaUniversity of CalgaryArtificial Intelligence in Medicine (Canada)
FundersNational Center for Advancing Translational SciencesNational Institute of Biomedical Imaging and BioengineeringNational Institute of General Medical SciencesNational Cancer InstituteCanadian Institutes of Health ResearchAnschutz Medical Campus, University of ColoradoSidney Kimmel Comprehensive Cancer CenterFaculty of Medicine and Health, University of SydneyUniversité de SherbrookeSchool of Electronic Engineering and Computer Science, Queen Mary University of LondonNational Institutes of HealthAgencia Nacional de Investigación y DesarrolloUniversitätsklinikum HeidelbergUniversitätsspital ZürichCanadian Institute for Advanced ResearchInselspital, Universitätsspital BernUniversität ZürichChina Scholarship CouncilUniversidad de ConcepciónMcGill UniversityConselho Nacional de Desenvolvimento Científico e TecnológicoQueen's UniversityIslamic Azad UniversityDeutsche ForschungsgemeinschaftSchool of Medicine, Indiana UniversityUniversity of BernUniversidad Carlos III de MadridMinisterstvo Zdravotnictví Ceské RepublikyBusiness FinlandFrederick National Laboratory for Cancer ResearchCancer Research UKVarian Medical SystemsWellcome TrustUniversidad Católica de CuencaDeutsches KrebsforschungszentrumThomas Jefferson UniversityLeidosU.S. Department of Veterans AffairsUniversity of PennsylvaniaEscuela Superior Politécnica del LitoralKepler UniversitätsklinikumQueen Mary University of LondonUniversity of PatrasOhio State UniversityHigher Education Commision, PakistanUniversity of TorontoFaculty of Arts and SciencesSilesian University of TechnologyUniversité du LuxembourgNational Science Foundation
KeywordsBenchmarkingBenchmark (surveying)Computer scienceSegmentationArtificial intelligenceGeneralizationMachine learningAlgorithmFederated learningData mining

Abstract

fetched live from OpenAlex

Computational competitions are the standard for benchmarking medical image analysis algorithms, but they typically use small curated test datasets acquired at a few centers, leaving a gap to the reality of diverse multicentric patient data. To this end, the Federated Tumor Segmentation (FeTS) Challenge represents the paradigm for real-world algorithmic performance evaluation. The FeTS challenge is a competition to benchmark (i) federated learning aggregation algorithms and (ii) state-of-the-art segmentation algorithms, across multiple international sites. Weight aggregation and client selection techniques were compared using a multicentric brain tumor dataset in realistic federated learning simulations, yielding benefits for adaptive weight aggregation, and efficiency gains through client sampling. Quantitative performance evaluation of state-of-the-art segmentation algorithms on data distributed internationally across 32 institutions yielded good generalization on average, albeit the worst-case performance revealed data-specific modes of failure. Similar multi-site setups can help validate the real-world utility of healthcare AI algorithms in the future.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Other design · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.942
Threshold uncertainty score0.631

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.014
GPT teacher head0.355
Teacher spread0.340 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designOther design
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations10
Published2025
Admission routes2
Has abstractyes

Explore more

Same venueNature CommunicationsSame topicRadiomics and Machine Learning in Medical ImagingFrench-language works237,207