Bibliographic record
Abstract
EditorialsJuly 2022Friend or Foe? The Role of Robots in Systematic ReviewsLisa Hartling, PhD and Allison Gates, PhDLisa Hartling, PhDAlberta Research Centre for Health Evidence, Department of Pediatrics, Faculty of Medicine & Dentistry, University of Alberta, Edmonton, Alberta, Canada and Allison Gates, PhDAlberta Research Centre for Health Evidence, Department of Pediatrics, Faculty of Medicine & Dentistry, University of Alberta, Edmonton, Alberta, CanadaAuthor, Article, and Disclosure Informationhttps://doi.org/10.7326/M22-1439 SectionsAboutFull TextPDF ToolsAdd to favoritesDownload CitationsTrack CitationsPermissions ShareFacebookTwitterLinkedInRedditEmail The dramatic increase in publication of health literature has generated a growing need for evidence syntheses to support decision making. Most recently, the COVID-19 pandemic has created a remarkable and unprecedented demand for the rapid production of reliable evidence syntheses (1). Systematic reviewers recognize the need to create efficiencies in review production while maintaining methodological rigor to ensure valid conclusions. To this end, technologies (often supported by machine learning and artificial intelligence) are being developed and used to fully or partially automate various stages of the systematic review process (2).One such technology, RobotReviewer, was evaluated in a trial by ...References1. Global Commission on Evidence to Address Societal Challenges. The Evidence Commission report: a wake-up call and path forward for decision-makers, evidence intermediaries, and impact-oriented evidence producers. McMaster Health Forum; 2022. Google Scholar2. Khalil H, Ameen D, Zarnegar A. Tools to support the automation of systematic reviews: a scoping review. J Clin Epidemiol. 2022;144:22-42. [PMID: 34896236] doi:10.1016/j.jclinepi.2021.12.005 CrossrefMedlineGoogle Scholar3. Arno A, Thomas J, Wallace B, et al. Accuracy and efficiency of machine learning–assisted risk-of-bias assessments in “real-world” systematic reviews. A noninferiority randomized controlled trial. Ann Intern Med. 2022;175:1001-9. doi:10.7326/M22-0092 LinkGoogle Scholar4. Higgins JPT, Savović J, Page MJ, et al. Chapter 8: Assessing risk of bias in a randomized trial. In: Higgins JPT, Thomas J, Chandler J, et al, eds. Cochrane Handbook for Systematic Reviews of Interventions. Version 6.3 (updated February 2022). The Cochrane Collaboration; 2022. Google Scholar5. O’Connor AM, Tsafnat G, Thomas J, et al. A question of trust: can we build an evidence base to gain trust in systematic review automation technologies. Syst Rev. 2019;8:143. [PMID: 31215463] doi:10.1186/s13643-019-1062-0 CrossrefMedlineGoogle Scholar6. Schumi J, Wittes JT. Through the looking glass: understanding non-inferiority. Trials. 2011;12:106. [PMID: 21539749] doi:10.1186/1745-6215-12-106 CrossrefMedlineGoogle Scholar7. Soboczenski F, Trikalinos TA, Kuiper J, et al. Machine learning to help researchers evaluate biases in clinical trials: a prospective, randomized user study. BMC Med Inform Decis Mak. 2019;19:96. [PMID: 31068178] doi:10.1186/s12911-019-0814-z CrossrefMedlineGoogle Scholar8. Arno A, Elliott J, Wallace B, et al. The views of health guideline developers on the use of automation in health evidence synthesis. Syst Rev. 2021;10:16. [PMID: 33419479] doi:10.1186/s13643-020-01569-2 CrossrefMedlineGoogle Scholar9. Higgins JP, Altman DG, Gøtzsche PC, et al; Cochrane Bias Methods Group. The Cochrane Collaboration's tool for assessing risk of bias in randomised trials. BMJ. 2011;343:d5928. [PMID: 22008217] doi:10.1136/bmj.d5928 CrossrefMedlineGoogle Scholar10. Marshall IJ, Wallace BC. Toward systematic review automation: a practical guide to using machine learning tools in research synthesis [Editorial]. Syst Rev. 2019;8:163. [PMID: 31296265] doi:10.1186/s13643-019-1074-9 CrossrefMedlineGoogle Scholar Author, Article, and Disclosure InformationAffiliations: Alberta Research Centre for Health Evidence, Department of Pediatrics, Faculty of Medicine & Dentistry, University of Alberta, Edmonton, Alberta, CanadaNote: Dr. Gates is employed by the Canadian Agency for Drugs and Technologies in Health (CADTH). This work was unrelated to her employment, and CADTH had no role in the work reported. Drs. Hartling and Gates have collaborated on papers with Joanne McKenzie, an author of the RobotReviewer trial referenced in this editorial.Financial Support: Dr. Hartling is supported by a Canada Research Chair in Knowledge Synthesis and Translation.Disclosures: Disclosures can be viewed at www.acponline.org/authors/icmje/ConflictOfInterestForms.do?msNum=M22-1439.Corresponding Author: Lisa Hartling, PhD, Alberta Research Centre for Health Evidence, Department of Pediatrics, Faculty of Medicine & Dentistry, University of Alberta, 4-472 ECHA, 11405 87 Avenue, Edmonton, AB T6G 2J3, Canada; e-mail, lisa.[email protected]ca.This article was published at Annals.org on 31 May 2022. PreviousarticleNextarticle Advertisement FiguresReferencesRelatedDetailsSee AlsoAccuracy and Efficiency of Machine Learning–Assisted Risk-of-Bias Assessments in “Real-World” Systematic Reviews Anneliese Arno , James Thomas , Byron Wallace , Iain J. Marshall , Joanne E. McKenzie , and Julian H. Elliott Metrics July 2022Volume 175, Issue 7Page: 1045-1046KeywordsClinical epidemiologyMachine learningRandomized trialsResearch quality assessmentSystematic reviews ePublished: 31 May 2022 Issue Published: July 2022 Copyright & PermissionsCopyright © 2022 by American College of Physicians. All Rights Reserved.PDF downloadLoading ...
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.556 | 0.884 |
| Meta-epidemiology (narrow) | 0.004 | 0.006 |
| Meta-epidemiology (broad) | 0.016 | 0.008 |
| Bibliometrics | 0.023 | 0.021 |
| Science and technology studies | 0.006 | 0.019 |
| Scholarly communication | 0.031 | 0.027 |
| Open science | 0.014 | 0.011 |
| Research integrity | 0.023 | 0.023 |
| Insufficient payload (model declined to judge) | 0.028 | 0.014 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".