MétaCan
Menu
← Back to cohort
Record W4387931634 · doi:10.2196/47105

Guidelines and Standard Frameworks for AI in Medicine: Protocol for a Systematic Literature Review

2023· article· en· W4387931634 on OpenAlexvenueno aff
Kirubel Biruk Shiferaw, Moritz Roloff, Dagmar Waltemath, Atinkut Alamirrew Zeleke

Bibliographic record

VenueJMIR Research Protocols · 2023
Typearticle
Languageen
FieldMedicine
TopicArtificial Intelligence in Healthcare and Education
Canadian institutionsnot available
FundersDeutscher Akademischer Austauschdienst
KeywordsSystematic reviewProtocol (science)Computer scienceTransparency (behavior)Health careMEDLINEBest practiceArtificial intelligenceData scienceMedicineAlternative medicine

Abstract

fetched live from OpenAlex

BACKGROUND: Applications of artificial intelligence (AI) are pervasive in modern biomedical science. In fact, research results suggesting algorithms and AI models for different target diseases and conditions are continuously increasing. While this situation undoubtedly improves the outcome of AI models, health care providers are increasingly unsure which AI model to use due to multiple alternatives for a specific target and the "black box" nature of AI. Moreover, the fact that studies rarely use guidelines in developing and reporting AI models poses additional challenges in trusting and adapting models for practical implementation. OBJECTIVE: This review protocol describes the planned steps and methods for a review of the synthesized evidence regarding the quality of available guidelines and frameworks to facilitate AI applications in medicine. METHODS: We will commence a systematic literature search using medical subject headings terms for medicine, guidelines, and machine learning (ML). All available guidelines, standard frameworks, best practices, checklists, and recommendations will be included, irrespective of the study design. The search will be conducted on web-based repositories such as PubMed, Web of Science, and the EQUATOR (Enhancing the Quality and Transparency of Health Research) network. After removing duplicate results, a preliminary scan for titles will be done by 2 reviewers. After the first scan, the reviewers will rescan the selected literature for abstract review, and any incongruities about whether to include the article for full-text review or not will be resolved by the third and fourth reviewer based on the predefined criteria. A Google Scholar (Google LLC) search will also be performed to identify gray literature. The quality of identified guidelines will be evaluated using the Appraisal of Guidelines, Research, and Evaluation (AGREE II) tool. A descriptive summary and narrative synthesis will be carried out, and the details of critical appraisal and subgroup synthesis findings will be presented. RESULTS: The results will be reported using the PRISMA (Preferred Reporting Items for Systematic Review and Meta-Analyses) reporting guidelines. Data analysis is currently underway, and we anticipate finalizing the review by November 2023. CONCLUSIONS: Guidelines and recommended frameworks for developing, reporting, and implementing AI studies have been developed by different experts to facilitate the reliable assessment of validity and consistent interpretation of ML models for medical applications. We postulate that a guideline supports the assessment of an ML model only if the quality and reliability of the guideline are high. Assessing the quality and aspects of available guidelines, recommendations, checklists, and frameworks-as will be done in the proposed review-will provide comprehensive insights into current gaps and help to formulate future research directions. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): DERR1-10.2196/47105.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.136
metaresearch head score (Gemma)0.254
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Systematic review · Consensus signal: Systematic review
GenreCandidate signal: Protocol · Consensus signal: Protocol
Teacher disagreement score0.136
Threshold uncertainty score0.718

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.1360.254
Meta-epidemiology (narrow)0.0060.007
Meta-epidemiology (broad)0.0170.018
Bibliometrics0.0210.018
Science and technology studies0.0060.007
Scholarly communication0.0110.012
Open science0.0060.008
Research integrity0.0110.009
Insufficient payload (model declined to judge)0.0790.017

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.714
GPT teacher head0.740
Teacher spread0.026 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSystematic review
Domainnot available
GenreProtocol

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations11
Published2023
Admission routes1
Has abstractyes

Explore more

Same venueJMIR Research Protocols→Same topicArtificial Intelligence in Healthcare and Education→French-language works237,207→