Fast, Low-resource, and Accurate oRgan and Pan-cancer sEgmentation in Abdomen CT
Notice bibliographique
Résumé
Abdomen organs are quite common cancer sites, such as colorectal cancer and pancreatic cancer, which are the 2nd and 3rd most common cause of cancer death [1,2]. Computed Tomography (CT) scanning yields important prognostic information for cancer patients and is a widely used technology for treatment monitoring. In both clinical trials and daily clinical practice, radiologists and clinicians measure the tumor and organ on CT scans based on manual two-dimensional measurements (e.g., Response Evaluation Criteria In Solid Tumors (RECIST) criteria) [3]. However, this manual assessment is inherently subjective with considerable inter- and intra-expert variability. Moreover, existing challenges mainly focus on one type of tumor (e.g., liver cancer, kidney cancer). There are still no general and publicly available models for universal abdominal organ and cancer segmentation at present. In this challenge, we aim to promote the development of universal organ and tumor segmentation in abdominal CT scans. This is an extension of FLARE2021 and FLARE2022 challenge. In FLARE2021, the challenge task is to segment four abdominal organs in a fully supervised setting. In FLARE2022, the challenge task is to segment 13 organs in a semi-supervised setting. In FLARE2023, we will add the lesion segmentation task. Different from existing tumor segmentation challenges [4,5], we focus on pan-cancer segmentation, which covers various abdominal cancer types. Specifically, the segmentation algorithm should segment 13 organs ( liver, spleen, pancreas, right kidney, left kidney, stomach, gallbladder, esophagus, aorta, inferior vena cava, right adrenal gland, left adrenal gland, and duodenum) and a single tumor class that includes all kinds of cancer types (such as liver cancer, kidney cancer, stomach cancer, pancreas cancer, colon cancer) in abdominal CT scans. Compared to the dataset in FLARE2021 (361 CT scans) and FLARE2022 (2300 CT scans), we further increase the dataset in FLARE2023 to 4500 CT scans (training/validation/testing: 4000, 100, 400), which is a multi-racial, multicenter, multi-disease, multi-phase, and multi-vendor dataset. To the best knowledge, this will be the largest and most diverse publicly available dataset for abdominal cancer analysis. The challenge will employ a partial-label learning setting where limited targets are annotated in each CT scan. This setting is in line with real-world settings because each medical department mainly focuses on one specific cancer. We aim to benchmark general abdominal cancer segmentation models that can handle multiple cancer types. Based on the results in FLARE2021 and FLARE2022, we found that segmentation models can achieve a good tradeoff between segmentation accuracy and efficiency. Thus, we will continue to evaluate both segmentation accuracy and efficiency. In particular, we will use Dice Similarity Coefficient (DSC), Normalized Surface Dice (NSD), and lesion-wise F1 score to evaluate segmentation accuracy, which is motivated by the metrics reloaded [6]. The segmentation efficiency is evaluated by running time and GPU memory consumption. In summary, the FLARE 2023 challenge has three main features: (1) Task: this is the first challenge for pan-cancer segmentation in abdominal CT scans (2) Dataset: we provide the largest abdomen CT dataset, including 4500 3D CT scans from 30+ medical centers. (3) Evaluation measures: we focus on both segmentation accuracy and segmentation efficiency [1] Nation Cancer Institute. "Cancer Stat Facts: Common Cancer Sites", 18 November 2022, https://seer.cancer.gov/statfacts/html/common.html. [2] Siegel, Rebecca, et al. "Cancer statistics, 2022. " CA: A Cancer Journal for Clinicians. 72 (2022): 7-33. [3] Eisenhauer, Elizabeth., et al. "New response evaluation criteria in solid tumours: revised RECIST guideline (version 1.1)." European journal of cancer 45.2 (2009): 228-247. [4] Bilic, Patrick, et al. The Liver Tumor Segmentation Benchmark (LiTS), Medical Image Analysis, (2022): 102680. [5] Heller, Nicholas, et al. "The state of the art in kidney and kidney tumor segmentation in contrast-enhanced CT imaging: Results of the KiTS19 challenge." Medical Image Analysis 67 (2021): 101821. [6] Maier-Hein, Lena, et al. "Metrics reloaded: Pitfalls and recommendations for image analysis validation." arXiv preprint arXiv:2206.01653 (2022).
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,001 | 0,004 |
| Méta-épidémiologie (sens strict) | 0,002 | 0,001 |
| Méta-épidémiologie (sens large) | 0,002 | 0,002 |
| Bibliométrie | 0,002 | 0,002 |
| Études des sciences et des technologies | 0,001 | 0,001 |
| Communication savante | 0,003 | 0,002 |
| Science ouverte | 0,003 | 0,002 |
| Intégrité de la recherche | 0,004 | 0,002 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,003 | 0,004 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».