AI-assisted clinical summary and treatment planning for cancer care: A comparative study of human vs. AI-based approaches.
Notice bibliographique
Résumé
1523 Background: Understanding a patient's clinical narrative, timeline, and history is critical for accurate treatment decision-making. However, reviewing and summarizing complex records is time-consuming and error-prone. Recent advancements in artificial intelligence (AI), specifically large language models (LLM), offer paths to improve quality and efficiency. Methods: A study was conducted on 50 breast cancer cases from an academic medical institution, utilizing all medical records—clinic, pathology, and radiology reports—up until the point of the initial treatment decision. All cases were processed using three different approaches: AI-assisted; full-AI; and human-only. In the AI-assisted method, two oncology physician assistants (PAs) revised AI-generated summaries to create clinical summaries. The full-AI method had AI independently produce clinical summaries, while the human-only method had the PAs compile summaries without AI. Eight board-certified international oncology specialists blindly evaluated summaries for faithfulness, completeness, and succinctness using a 3-point scale, ranked their preferences, and tried to predict which summaries were full-AI. Rankings were assessed using a Friedman test followed by a Wilcoxon signed-rank test, and full-AI prediction was assessed using a two-sided one-sample binomial test. After summarization, a distinct AI system with access to clinical guidelines provided treatment plans. These plans were then evaluated by a board-certified oncologist with access to the original treatment decision. Results: The study found specialists favored AI-assisted, followed by full-AI, and then human-only summaries, with average ranks of 1.73, 1.93, 2.34 respectively (lower is better, p<0.001). The difference between full-AI and AI-assisted was not significant (p=0.11). Evaluation scores (mean±95%CI, higher is better) showed AI-assisted, full-AI, and human-only scored 2.35±0.13, 2.14±0.14, 2.17±0.14 for faithfulness; 2.28±0.12, 2.01±0.12, 1.93±0.14 for completeness; and 2.33±0.12, 2.21±0.12, 1.99±0.13 for succinctness. The average summarization time was 19.71, 1.17, 26.03 minutes. Full-AI identification accuracy was 0.28 (not different from chance 0.33, p=0.46). With AI-assisted summaries, the treatment plans were accurate in 45 cases (90%) and partially accurate in 5 cases (10%). In the 5 partially accurate cases, the system was accurate with the provided input data, but there were inaccuracies with the input data, including incorrect formats or missing data. Conclusions: Incorporating LLMs into the creation of medical summaries has shown improvements in both quality and efficiency, achieving up to 22.2x speed up with full-AI, indicating that AI-assisted summarization tools can potentially enhance care quality. AI-assisted summaries yield accurate treatment plans when the input data is accurate.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,017 | 0,079 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,001 |
| Bibliométrie | 0,002 | 0,002 |
| Études des sciences et des technologies | 0,001 | 0,001 |
| Communication savante | 0,002 | 0,002 |
| Science ouverte | 0,001 | 0,002 |
| Intégrité de la recherche | 0,001 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,004 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».