Re: Deng and Heybati
Bibliographic record
Abstract
We appreciate the comments submitted by Deng and Heybati, who agree that artificial intelligence (AI) has a potential role in clinical trial enrollment but noted two limitations of our review, namely, the heterogeneity of AI workflows and funding sources across the published studies (1). We thank them for their interest and would like to respond as follows. As mentioned in our Limitations section, we agree that there is heterogeneity in the types of AI algorithms used across published studies and where they are integrated into the workflow. Given the potential of AI to improve multiple steps of the workflow and ongoing investigations in these areas, we would likely be able to better answer questions such as “Which enrollment step has the greatest implications for AI performance?” in future reviews. These questions are unfortunately just not possible to answer with the current state of the literature. Additionally, detailed information regarding the workings of several algorithms was not provided, and we hope subsequent studies would share this information publicly to facilitate answers to these important questions. As mentioned in our article, and reiterated by Deng and Heybati, we encourage further investigations into the parameters of AI algorithms/workflows that are critical for test characteristics. Our noting of the higher performance of industry algorithms observed post hoc should be considered exploratory. We agree that conflicts of interest should be considered in the interpretation of any results, either in-house or industry-developed studies. However, we caution the outright dismissal of study results and validity simply because they are industry supported. Although potential biases may contribute to the improved reported performance of the industry algorithms, it is also possible that these algorithms were more sophisticated given the increased resources devoted to software development and deployment that might have contributed to their success. Increased transparency on the workings of these algorithms and validation of these results would be required to confirm these findings, which we encourage as part of future investigations into the parameters of AI algorithms. We would be remiss if we did not mention that there may be other biases at play, which are common to review articles but not necessarily mentioned, including positive results bias, selective outcome reporting bias, journal/location bias, and time lag bias, particularly for clinical trials. We always encourage critical thinking and dialogue, about not only conflicts of interest and potential biases but also individual study validity, results, and generalizability. An omission of mentioning a bias does not necessarily mean that a bias might not exist. We again thank Deng and Heybati for their comments and interest in our article. No new data were generated or analyzed in support of this letter. Ronald Chow, MS, MEng (Conceptualization; Writing—original draft), Fei-Fei Liu, MD (Writing—review & editing), Benjamin Haibe-Kains, PhD (Writing—review & editing), Michael Lock, MD (Writing—review & editing), Srinivas Raman, MD, MASc (Conceptualization; Writing—review & editing). This work was partially funded by the CARO-CROF Pamela Catton Summer Studentship Award and the Robert L. Tundermann and Christine E. Couturier philanthropic funds. None. The funders had no influence on the design and conduct of the study; collection, management, analysis, and interpretation of the data; preparation, review, or approval of the manuscript; and decision to submit the manuscript for publication.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.007 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.003 | 0.003 |
| Insufficient payload (model declined to judge) | 0.123 | 0.057 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".