RE: Use of artificial intelligence for cancer clinical trial enrollment
Bibliographic record
Abstract
Clinical trials represent a pivotal step in advancing novel cancer therapies from development to clinical application. Despite a strong willingness among patients to participate, however, less than 5% of adult patients with cancer are enrolled in oncology trials. Enhancing enrollment workflows can expedite treatment advances and lead to faster patient outcome improvements (1). In their review in this issue of the Journal, Chow et al. (2) found that artificial intelligence (AI) workflows for trial enrollment outperformed manual methods, with industry-developed systems having higher positive predictive values than in-house systems. Although we commend Chow et al. for their comprehensive review, we believe that certain concerns warrant further discussion. First, the AI workflows examined had substantial heterogeneities, which complicates the meta-analysis and interpretation of AI performance. AI enrollment workflows can generally be broken down into 3 steps: 1) extracting eligibility criteria from protocols, 2) extracting data from electronic health records, and 3) matching the extracted data with the eligibility criteria (3). The studies included in the review differed greatly in how they approached these steps. For instance, where Meystre et al. (3) applied AI only to the latter 2 steps, Beck et al. (4) automated all 3. Thus, the accuracy of Meystre et al.’s workflow would be less affected by AI than that of Beck et al. Meanwhile, other studies, such as Calaprice-Whitty et al. (5), incorporated auxiliary systems such as optical character recognition, which introduced additional failure points that could reduce the study’s accuracy. The workflow assessed may differ even when the same AI platform was assessed in different studies. For example, both Alexander et al. (6) and Beck et al. (4) evaluated IBM’s Watson for Clinical Trial Matching (WCTM) system (IBM Corp, Armonk, NY). Yet, where Alexander et al. manually entered input parameters into WCTM, Beck et al. used WCTM’s natural language processing system to extract the input parameters from unstructured electronic health records. Hence, Alexander et al.’s study accuracy would be less affected by the performance of the WCTM natural language processing system than the study by Beck et al. Given such variations, summary statistics from meta-analyses become less insightful. What do the pooled metrics actually mean? Which enrollment step has the greatest implications for AI performance? Which AI algorithm is the best for each step? How do auxiliary systems such as optical character recognition affect the overall AI system’s performance? These are important questions that were overlooked in Chow et al.’s (2) review and require closer assessment in future reviews. Second, although the authors found that industry-developed systems exhibited higher positive predictive values than did in-house systems, they did not discuss the potential impact of conflict of interest on these findings. Industry sponsorship has been a historical sticking point in oncology research, with industry-funded studies more likely to yield positive outcomes and publish in high-impact journals (7). In Chow et al.’s review, studies assessing industry-developed systems all received industry funding, often from the very developers of the platforms under review (see Table 1). Many authors of these studies were also employed by the developers or hold stock options in the developers. Given that these preliminary studies are not registered as clinical trials, they are prone to biases. Hence, caution should be exercised when interpreting these results. Funding sources and conflict of interest of studies assessing the industry-developed systems included in the Chow et al. review Funding sources and conflict of interest of studies assessing the industry-developed systems included in the Chow et al. review In conclusion, although we share the authors’ belief in AI’s potential role in clinical trial enrollment, further investigations from nonindustry sources—with particular focus on each specific step of the AI enrollment process—would be needed to advance the field of AI applications in oncology trials. No data are reported in this correspondence. Jiawen Deng, MS-2 (Conceptualization; Investigation; Writing—original draft; Writing—review & editing), Kiyan Heybati, MS-3, MSc(c) (Investigation; Writing—review & editing). This correspondence and its authors received no funding support from any funding agency in the public, commercial, or not-for-profit sectors. The authors declare no conflicts of interest.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Direct model labels (unvalidated)
Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.
| Model arm | Categories | Study design | Confidence |
|---|---|---|---|
| gemma | no category Domain: not available · Genre: Commentary About the Canadian research system: no · About a Canadian topic: no | Not applicable | low |
| gpt | no category Domain: not available · Genre: Commentary About the Canadian research system: no · About a Canadian topic: no | Not applicable | high |
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.132 | 0.376 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.006 |
| Bibliometrics | 0.010 | 0.008 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.006 | 0.006 |
| Open science | 0.005 | 0.004 |
| Research integrity | 0.002 | 0.004 |
| Insufficient payload (model declined to judge) | 0.012 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedLabeled directly by 2 models reading the full record.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".