ChatGPT in a gynaecologic oncology multidisciplinary team tumour board: A feasibility study
Notice bibliographique
Résumé
The practical medical use of artificial intelligence is rapidly progressing. Specifically, the application of ChatGPT was explored in medical education and even medical clinical data evaluation.1, 2 Tumour board is an integral and pivotal part of patient treatment and management in gynaecologic oncology.3 It entails the processing of various pathological and clinical parameters, coupled with the familiarity with treatment guidelines in accordance with the various parameters. The participation of ChatGPT in breast cancer tumour board was previously studied, with contrasting results.4, 5 We aim to study the feasibility of ChatGPT (Versions 3.5 and 4) as a support tool for endometrial cancer (EC) and ovarian cancer (OC) according to the NCCN and ESGO guidelines. Ten EC cases and ten OC cases were fabricated based on experience of authors pertaining to the most complex scenarios discussed in real practice. For EC the following data was formulated: age, histology, stage, grade, lymphovascular space invasion, tumour size and molecular classification—MMR, p53 and POLE mutation status. For OC, the following data was formulated: age, histology and stage. We created a new account for ChatGPT 3.5 and purchased and created an account for ChatGPT 4. We used generic prompts for all the cases. The ChatGPT 3.5 and ChatGPT 4 prompt are described (Appendix S1). For each tumour board case, we accessed the NCCN and ESGO guidelines separately and recorded their recommendation. All ChatGPT recommendations were judged as correct or incorrect by two independent reviewers (G.L. and Y.B.). Data analysis is described in detail in the Appendix S1. We used SPSS 29 for the statistical analysis. As no patient information was used—no ethical board review was needed for this study. There were ten cases of EC cancer, stages IA-IIIC with four different histology, and ten cases of OC stages IA-IC3 with five different histology. ChatGPT 3.5 was unable to give a concrete recommendation, and ChatGPT 4 gave a recommendation to all cases. No disagreements between reviewers were noted for all 40 evaluations. The rate of correct recommendations was 70% (14/20) for NCCN guidelines and 60% (12/20) for ESGO guidelines (p = 0.512). (Table 1). There were 55% (11/20) of cases with correct recommendations for both guidelines, 20% (4/20) of cases in which a correct recommendation was given only according to one guideline (Figure S1), and 25% (5/20) of cases in which an incorrect recommendation was given. Of those with an incorrect recommendation, 80% (4/5) were EC, stages IA-II, of all histology, and one case of OC, stage IA. Of the four single guidelines correct recommendations, all were EC, with three incorrect recommendations according to ESGO guidelines, including the only two cases with a positive POLE mutation. OC had higher complete correct recommendation as compared to EC (90% vs. 20%, p = 0.005). ChatGPT 4 suggestions for adjuvant treatment are presented in Tables S1 and S2. In this feasibility study, we showed that ChatGPT 4 provided correct recommendations in two-thirds of the cases evaluated, however in 25% of cases, mostly endometrial cancer, there was an incorrect recommendation. Endometrial cancer had a lower complete rate of correct recommendations, likely due to the complexity of stage, histology and grade in early stages and in the integration of molecular characterisation of endometrial cancer. More research is required to assess the credibility and configure protocols for the potential use of this tool. However, in a setting of high-volume clinics, or in regions where resources are limiting in terms of expertise, such tools may aid physicians maintain evidenced-based care. Further studies should focus on ChatGPT familiarity with ongoing clinical trials to assess for possible patient eligibility. Our limitations include the small number of cases studied and limiting our study to endometrial and ovarian cancer. Additionally, we have used the generic ChatGPT tool without any specific training for our data. Moreover, we have used only two AI platforms in this study, this may limit the generalisability of our results. Importantly, we did not compare the AI-generated recommendation to a multidisciplinary Tumor Board recommendation, which is the 'gold standard' in real practice. Finally, all data is correct to the time this manuscript was written. As ChatGPT is a large language model, he is constantly trains on prompts and his output may change and evolve over time. Future prospective real-life evaluation of gynaecologic oncology tumour board is encouraged to better delineate advantages and pitfalls of artificial intelligence tools and their impact on practice. Gabriel Levin: conception, design, acquisition of data, analysis and interpretation of data, drafting the article, approval of the final version. Walter Gotlieb: acquisition of data, critical revision of the article, approval of the final version. Pedro Ramirez: acquisition of data, critical revision of the article, approval of the final version. Raanan Meyer: acquisition of data, critical revision of the article, approval of the final version. Yoav Brezinov: conception and design, analysis and interpretation of data, critical revision of the article, approval of the final version. This research received no external funding. None. The authors report no conflict of interest. As no patient information was used—no ethical board review was needed for this study. The data that support the findings of this study are available from the corresponding author upon reasonable request. Appendix S1. Table S1-ChatGPT 4 suggestions for adjuvant treatment for ovarian cancer Table S2-GhatGPT 4 suggestions for adjuvant treatment for endomerial cancer Table S2. Please note: The publisher is not responsible for the content or functionality of any supporting information supplied by the authors. Any queries (other than missing content) should be directed to the corresponding author for the article.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,014 | 0,025 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,001 |
| Méta-épidémiologie (sens large) | 0,001 | 0,001 |
| Bibliométrie | 0,001 | 0,001 |
| Études des sciences et des technologies | 0,002 | 0,001 |
| Communication savante | 0,002 | 0,002 |
| Science ouverte | 0,002 | 0,003 |
| Intégrité de la recherche | 0,002 | 0,002 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,009 | 0,002 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».