Opportunities, challenges, and requirements for Artificial Intelligence (AI) implementation in Primary Health Care (PHC): a systematic review
Bibliographic record
Abstract
BACKGROUND: Artificial Intelligence (AI) has significantly reshaped Primary Health Care (PHC), offering various possibilities and complexities across all functional dimensions. The objective is to review and synthesize available evidence on the opportunities, challenges, and requirements of AI implementation in PHC based on the Primary Care Evaluation Tool (PCET). METHODS: We conducted a systematic review, following the Cochrane Collaboration method, to identify the latest evidence regarding AI implementation in PHC. A comprehensive search across eight databases- PubMed, Web of Science, Scopus, Science Direct, Embase, CINAHL, IEEE, and Cochrane was conducted using MeSH terms alongside the SPIDER framework to pinpoint quantitative and qualitative literature published from 2000 to 2024. Two reviewers independently applied inclusion and exclusion criteria, guided by the SPIDER framework, to review full texts and extract data. We synthesized extracted data from the study characteristics, opportunities, challenges, and requirements, employing thematic-framework analysis, according to the PCET model. The quality of the studies was evaluated using the JBI critical appraisal tools. RESULTS: In this review, we included a total of 109 articles, most of which were conducted in North America (n = 49, 44%), followed by Europe (n = 36, 33%). The included studies employed a diverse range of study designs. Using the PCET model, we categorized AI-related opportunities, challenges, and requirements across four key dimensions. The greatest opportunities for AI integration in PHC were centered on enhancing comprehensive service delivery, particularly by improving diagnostic accuracy, optimizing screening programs, and advancing early disease prediction. However, the most challenges emerged within the stewardship and resource generation functions, with key concerns related to data security and privacy, technical performance issues, and limitations in data accessibility. Ensuring successful AI integration requires a robust stewardship function, strategic investments in resource generation, and a collaborative approach that fosters co-development, scientific advancements, and continuous evaluation. CONCLUSIONS: Successful AI integration in PHC requires a coordinated, multidimensional approach, with stewardship, resource generation, and financing playing key roles in enabling service delivery. Addressing existing knowledge gaps, examining interactions among these dimensions, and fostering a collaborative approach in developing AI solutions among stakeholders are essential steps toward achieving an equitable and efficient AI-driven PHC system. PROTOCOL: Registered in Open Science Framework (OSF) ( https://doi.org/10.17605/OSF.IO/HG2DV ).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.057 | 0.178 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.007 | 0.008 |
| Bibliometrics | 0.017 | 0.018 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.007 | 0.007 |
| Open science | 0.003 | 0.004 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".