Computer-Aided Systematic Review Screening Comes of Age
Bibliographic record
Abstract
Editorials1 August 2017Computer-Aided Systematic Review Screening Comes of AgeBrian J. Hemens, BScPhm, MSc, RPh and Alfonso Iorio, MD, PhDBrian J. Hemens, BScPhm, MSc, RPhFrom McMaster University, Hamilton, Ontario, Canada.Search for more papers by this author and Alfonso Iorio, MD, PhDFrom McMaster University, Hamilton, Ontario, Canada.Search for more papers by this authorAuthor, Article, and Disclosure Informationhttps://doi.org/10.7326/M17-1295 SectionsAboutFull TextPDF ToolsAdd to favoritesDownload CitationsTrack CitationsPermissions ShareFacebookTwitterLinkedInRedditEmail Publication of the 2017 update of the clinical practice guideline on the treatment of low bone density or osteoporosis to prevent fractures from the American College of Physicians (1) marks an important advancement in systematic review methodology. In addition to altering advice that will improve patient care, this update is based on new evidence found by training a computer to automatically identify relevant references.Shekelle and colleagues (2) describe important refinements to machine-learning software designed to reduce the need for human activity to identify relevant studies for a systematic review update. Those who conduct systematic reviews know that as the ...References1. Qaseem A, Forciea MA, McLean RM, Denberg TD; Clinical Guidelines Committee of the American College of Physicians. Treatment of low bone density or osteoporosis to prevent fractures in men and women: a clinical practice guideline update from the American College of Physicians. Ann Intern Med. 2017;166:818-39. [PMID: 28492856]. doi:10.7326/M15-1361 LinkGoogle Scholar2. Shekelle PG, Shetty K, Newberry S, Maglione M, Motala A. Machine learning versus standard techniques for updating searches for systematic reviews: a diagnostic accuracy study [Letter]. Ann Intern Med. 2017;167:213-5. doi:10.7326/L17-0124 LinkGoogle Scholar3. Shojania KG, Sampson M, Ansari MT, Ji J, Doucette S, Moher D. How quickly do systematic reviews go out of date? A survival analysis. Ann Intern Med. 2007;147:224-33. [PMID: 17638714] LinkGoogle Scholar4. Garner P, Hopewell S, Chandler J, MacLehose H, Schünemann HJ, Akl EA, et al; Panel for Updating Guidance for Systematic Reviews (PUGs). When and how to update systematic reviews: consensus and checklist. BMJ. 2016;354:i3507. [PMID: 27443385] doi:10.1136/bmj.i3507 CrossrefMedlineGoogle Scholar5. Wilczynski NL, McKibbon KA, Haynes RB. Enhancing retrieval of best evidence for health care from bibliographic databases: calibration of the hand search of the literature. Stud Health Technol Inform. 2001;84:390-3. [PMID: 11604770] MedlineGoogle Scholar6. Wilczynski NL, McKibbon KA, Haynes RB. Search filter precision can be improved by NOTing out irrelevant content. AMIA Annu Symp Proc. 2011;2011:1506-13. [PMID: 22195215] MedlineGoogle Scholar7. Sampson M, Tetzlaff J, Urquhart C. Precision of healthcare systematic review searches in a cross-sectional sample. Res Synth Methods. 2011;2:119-25. [PMID: 26061680] doi:10.1002/jrsm.42 CrossrefMedlineGoogle Scholar8. Hemens BJ, Haynes RB. McMaster Premium LiteratUre Service (PLUS) performed well for identifying new studies for updated Cochrane reviews. J Clin Epidemiol. 2012;65:62-72. [PMID: 21856121] doi:10.1016/j.jclinepi.2011.02.010 CrossrefMedlineGoogle Scholar9. Eden J, Levit L, Berg A, Morton S, eds. Institute of Medicine; Committee on Standards for Systematic Reviews of Comparative Effectiveness Research. Finding What Works in Health Care: Standards for Systematic Reviews. Washington, DC: National Academies Pr; 2011. [PMID: 24983062] doi:10.17226/13059 CrossrefMedlineGoogle Scholar10. Paynter R, Banez LL, Berliner E, Erinoff E, Lege-Matsuura J, Potter S, et al. EPC Methods: An Exploration of the Use of Text-Mining Software in Systematic Reviews. Report no. 16-EHC023-EF. Rockville: Agency for Healthcare Research and Quality; 2016. [PMID: 27195359] Google Scholar Author, Article, and Disclosure InformationAffiliations: From McMaster University, Hamilton, Ontario, Canada.Disclosures: Authors have disclosed no conflicts of interest. Forms can be viewed at www.acponline.org/authors/icmje/ConflictOfInterestForms.do?msNum=M17-1295.Corresponding Author: Alfonso Iorio, MD, PhD, Health Information Research Unit, McMaster University, 1280 Main Street W, CRL-140, Hamilton, Ontario L8S 4K1, Canada; e-mail, [email protected]ca.Current Author Addresses: Mr. Hemens: McMaster University, 1280 Main Street W, HSC 2C, Hamilton, Ontario L3N 8Z5, Canada.Dr. Iorio: Health Information Research Unit, McMaster University, 1280 Main Street W, CRL-140, Hamilton, Ontario M4Y 2X6, Canada.This article was published at Annals.org on 13 June 2017. PreviousarticleNextarticle Advertisement FiguresReferencesRelatedDetailsSee AlsoMachine Learning Versus Standard Techniques for Updating Searches for Systematic Reviews: A Diagnostic Accuracy Study Paul G. Shekelle , Kanaka Shetty , Sydne Newberry , Margaret Maglione , and Aneesa Motala Metrics Cited byrevtools: An R package to support article screening for evidence synthesisA question of trust: can we build an evidence base to gain trust in systematic review automation technologies? 1 August 2017Volume 167, Issue 3Page: 210-211KeywordsComputersHealth information technologyInformation technologyMachine learningOsteoporosisPrecision medicineSoftware designSpecificitySystematic reviewsTreatment guidelines ePublished: 13 June 2017 Issue Published: 1 August 2017 Copyright & PermissionsCopyright © 2017 by American College of Physicians. All Rights Reserved.PDF downloadLoading ...
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.285 | 0.745 |
| Meta-epidemiology (narrow) | 0.005 | 0.005 |
| Meta-epidemiology (broad) | 0.020 | 0.013 |
| Bibliometrics | 0.068 | 0.042 |
| Science and technology studies | 0.004 | 0.004 |
| Scholarly communication | 0.025 | 0.018 |
| Open science | 0.010 | 0.010 |
| Research integrity | 0.010 | 0.007 |
| Insufficient payload (model declined to judge) | 0.132 | 0.024 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".