Implementation of Machine Learning Applications in Health Care Organizations: Protocol for a Systematic Review of Empirical Studies
Bibliographic record
Abstract
BACKGROUND: An increasing interest in machine learning (ML) has been observed among scholars and health care professionals. However, while ML-based applications have been shown to be effective and have the potential to change the delivery of patient care, their implementation in health care organizations is complex. There are several challenges that currently hamper the uptake of ML in daily practice, and there is currently limited knowledge on how these challenges have been addressed in empirical studies on implemented ML-based applications. OBJECTIVE: The aim of this systematic literature review is twofold: (1) to map the ML-based applications implemented in health care organizations, with a focus on investigating the organizational dimensions that are relevant in the implementation process; and (2) to analyze the processes and strategies adopted to foster a successful uptake of ML. METHODS: We developed this protocol following the PRISMA-P (Preferred Reporting Items for Systematic Review and Meta-Analysis Protocols) guidelines. The search was conducted on 3 databases (PubMed, Scopus, and Web of Science), considering a 10-year time frame (2013-2023). The search strategy was built around 4 blocks of keywords (artificial intelligence, implementation, health care, and study type). Based on the detailed inclusion criteria defined, only empirical studies documenting the implementation of ML-based applications used by health care professionals in clinical settings will be considered. The study protocol was registered in PROSPERO (International Prospective Register of Systematic Reviews). RESULTS: The review is ongoing and is expected to be completed by September 2023. Data analysis is currently underway, and the first results are expected to be submitted for publication in November 2023. The study was funded by the European Union within the Multilayered Urban Sustainability Action (MUSA) project. CONCLUSIONS: ML-based applications involving clinical decision support and automation of clinical tasks present unique traits that add several layers of complexity compared with earlier health technologies. Our review aims at contributing to the existing literature by investigating the implementation of ML from an organizational perspective and by systematizing a conspicuous amount of information on factors influencing implementation. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): DERR1-10.2196/47971.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.121 | 0.152 |
| Meta-epidemiology (narrow) | 0.007 | 0.006 |
| Meta-epidemiology (broad) | 0.022 | 0.019 |
| Bibliometrics | 0.018 | 0.018 |
| Science and technology studies | 0.004 | 0.007 |
| Scholarly communication | 0.008 | 0.010 |
| Open science | 0.005 | 0.005 |
| Research integrity | 0.009 | 0.008 |
| Insufficient payload (model declined to judge) | 0.056 | 0.009 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".