RE: Use of artificial intelligence for cancer clinical trial enrollment
Notice bibliographique
Résumé
Clinical trials represent a pivotal step in advancing novel cancer therapies from development to clinical application. Despite a strong willingness among patients to participate, however, less than 5% of adult patients with cancer are enrolled in oncology trials. Enhancing enrollment workflows can expedite treatment advances and lead to faster patient outcome improvements (1). In their review in this issue of the Journal, Chow et al. (2) found that artificial intelligence (AI) workflows for trial enrollment outperformed manual methods, with industry-developed systems having higher positive predictive values than in-house systems. Although we commend Chow et al. for their comprehensive review, we believe that certain concerns warrant further discussion. First, the AI workflows examined had substantial heterogeneities, which complicates the meta-analysis and interpretation of AI performance. AI enrollment workflows can generally be broken down into 3 steps: 1) extracting eligibility criteria from protocols, 2) extracting data from electronic health records, and 3) matching the extracted data with the eligibility criteria (3). The studies included in the review differed greatly in how they approached these steps. For instance, where Meystre et al. (3) applied AI only to the latter 2 steps, Beck et al. (4) automated all 3. Thus, the accuracy of Meystre et al.’s workflow would be less affected by AI than that of Beck et al. Meanwhile, other studies, such as Calaprice-Whitty et al. (5), incorporated auxiliary systems such as optical character recognition, which introduced additional failure points that could reduce the study’s accuracy. The workflow assessed may differ even when the same AI platform was assessed in different studies. For example, both Alexander et al. (6) and Beck et al. (4) evaluated IBM’s Watson for Clinical Trial Matching (WCTM) system (IBM Corp, Armonk, NY). Yet, where Alexander et al. manually entered input parameters into WCTM, Beck et al. used WCTM’s natural language processing system to extract the input parameters from unstructured electronic health records. Hence, Alexander et al.’s study accuracy would be less affected by the performance of the WCTM natural language processing system than the study by Beck et al. Given such variations, summary statistics from meta-analyses become less insightful. What do the pooled metrics actually mean? Which enrollment step has the greatest implications for AI performance? Which AI algorithm is the best for each step? How do auxiliary systems such as optical character recognition affect the overall AI system’s performance? These are important questions that were overlooked in Chow et al.’s (2) review and require closer assessment in future reviews. Second, although the authors found that industry-developed systems exhibited higher positive predictive values than did in-house systems, they did not discuss the potential impact of conflict of interest on these findings. Industry sponsorship has been a historical sticking point in oncology research, with industry-funded studies more likely to yield positive outcomes and publish in high-impact journals (7). In Chow et al.’s review, studies assessing industry-developed systems all received industry funding, often from the very developers of the platforms under review (see Table 1). Many authors of these studies were also employed by the developers or hold stock options in the developers. Given that these preliminary studies are not registered as clinical trials, they are prone to biases. Hence, caution should be exercised when interpreting these results. Funding sources and conflict of interest of studies assessing the industry-developed systems included in the Chow et al. review Funding sources and conflict of interest of studies assessing the industry-developed systems included in the Chow et al. review In conclusion, although we share the authors’ belief in AI’s potential role in clinical trial enrollment, further investigations from nonindustry sources—with particular focus on each specific step of the AI enrollment process—would be needed to advance the field of AI applications in oncology trials. No data are reported in this correspondence. Jiawen Deng, MS-2 (Conceptualization; Investigation; Writing—original draft; Writing—review & editing), Kiyan Heybati, MS-3, MSc(c) (Investigation; Writing—review & editing). This correspondence and its authors received no funding support from any funding agency in the public, commercial, or not-for-profit sectors. The authors declare no conflicts of interest.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Étiquettes directes de modèles (non validées)
Étiquettes de catégorie et de devis d'étude par modèle, issues des rondes d'étiquetage. C'est une sortie machine, non validée, et le désaccord entre modèles est livré comme donnée. Aucun devis ici n'est encore validé contre MEDLINE.
| Bras | Catégories | Devis d'étude | Confiance |
|---|---|---|---|
| gemma | aucune catégorie Domaine: non disponible · Genre: Commentaire Porte sur le système de recherche canadien: non · Porte sur un sujet canadien: non | Sans objet | low |
| gpt | aucune catégorie Domaine: non disponible · Genre: Commentaire Porte sur le système de recherche canadien: non · Porte sur un sujet canadien: non | Sans objet | high |
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,132 | 0,376 |
| Méta-épidémiologie (sens strict) | 0,002 | 0,001 |
| Méta-épidémiologie (sens large) | 0,003 | 0,006 |
| Bibliométrie | 0,010 | 0,008 |
| Études des sciences et des technologies | 0,001 | 0,002 |
| Communication savante | 0,006 | 0,006 |
| Science ouverte | 0,005 | 0,004 |
| Intégrité de la recherche | 0,002 | 0,004 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,012 | 0,003 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéeÉtiqueté directement par 2 modèles lisant le dossier complet.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».