MétaCan
Menu
← Back to cohort
Record W3154024252 · doi:10.2196/25377

Measuring the Quality of Clinical Skills Mobile Apps for Student Learning: Systematic Search, Analysis, and Comparison of Two Measurement Scales

2021· review· en· W3154024252 on OpenAlexvenueno aff
Tehmina Gladman, Grace Tylee, Steve Gallagher, Jonathan Mair, Rebecca Grainger

Bibliographic record

VenueJMIR mhealth and uhealth · 2021
Typereview
Languageen
FieldHealth Professions
TopicMobile Health and mHealth Applications
Canadian institutionsnot available
FundersUniversity of Otago
KeywordsMobile appsQuality (philosophy)mHealthMobile deviceComputer scienceMobile technologyPsychologyMedical educationMultimediaWorld Wide WebMedicinePsychological intervention

Abstract

fetched live from OpenAlex

BACKGROUND: Mobile apps are widely used in health professions, which increases the need for simple methods to determine the quality of apps. In particular, teachers need the ability to curate high-quality mobile apps for student learning. OBJECTIVE: This study aims to systematically search for and evaluate the quality of clinical skills mobile apps as learning tools. The quality of apps meeting the specified criteria was evaluated using two measures-the widely used Mobile App Rating Scale (MARS), which measures general app quality, and the Mobile App Rubric for Learning (MARuL), a recently developed instrument that measures the value of apps for student learning-to assess whether MARuL is more effective than MARS in identifying high-quality apps for learning. METHODS: Two mobile app stores were systematically searched using clinical skills terms commonly found in medical education and apps meeting the criteria identified using an approach based on PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines. A total of 9 apps were identified during the screening process. The apps were rated independently by 2 reviewers using MARS and MARuL. RESULTS: The intraclass correlation coefficients (ICCs) for the 2 raters using MARS and MARuL were the same (MARS ICC [two-way]=0.68; P<.001 and MARuL ICC [two-way]=0.68; P<.001). Of the 9 apps, Geeky Medics-OSCE revision (MARS Android=3.74; MARS iOS=3.68; MARuL Android=75; and MARuL iOS=73) and OSCE PASS: Medical Revision (MARS Android=3.79; MARS iOS=3.71; MARuL Android=69; and MARuL iOS=73) scored highly on both measures of app quality and for both Android and iOS. Both measures also showed agreement for the lowest rated app, Patient Education Institute (MARS Android=2.21; MARS iOS=2.11; MARuL Android=18; and MARuL iOS=21.5), which had the lowest scores in all categories except information (MARS) and professional (MARuL) in both operating systems. MARS and MARuL were both able to differentiate between the highest and lowest quality apps; however, MARuL was better able to differentiate apps based on teaching and learning quality. CONCLUSIONS: This systematic search and rating of clinical skills apps for learning found that the quality of apps was highly variable. However, 2 apps-Geeky Medics-OSCE revision and OSCE PASS: Medical Revision-rated highly for both versions and with both quality measures. MARS and MARuL showed similar abilities to differentiate the quality of the 9 apps. However, MARuL's incorporation of teaching and learning elements as part of a multidimensional measure of quality may make it more appropriate for use with apps focused on teaching and learning, whereas MARS's more general rating of quality may be more appropriate for health apps targeting a general health audience. Ratings of the 9 apps by both measures also highlighted the variable quality of clinical skills mobile apps for learning.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.106
metaresearch head score (Gemma)0.314
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Systematic review · Consensus signal: Systematic review
GenreCandidate signal: Review · Consensus signal: Review
Teacher disagreement score0.894
Threshold uncertainty score0.562

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.1060.314
Meta-epidemiology (narrow)0.0020.001
Meta-epidemiology (broad)0.0120.015
Bibliometrics0.0270.021
Science and technology studies0.0010.003
Scholarly communication0.0040.005
Open science0.0030.004
Research integrity0.0020.001
Insufficient payload (model declined to judge)0.0020.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.488
GPT teacher head0.648
Teacher spread0.160 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designSystematic review
DomainMethods
GenreReview

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations7
Published2021
Admission routes1
Has abstractyes

Explore more

Same venueJMIR mhealth and uhealth→Same topicMobile Health and mHealth Applications→French-language works237,207→