Improving Skin Cancer Diagnostics Through a Mobile App With a Large Interactive Image Repository: Randomized Controlled Trial
Bibliographic record
Abstract
BACKGROUND: Skin cancer diagnostics is challenging, and mastery requires extended periods of dedicated practice. OBJECTIVE: The aim of the study was to determine if self-paced pattern recognition training in skin cancer diagnostics with clinical and dermoscopic images of skin lesions using a large-scale interactive image repository (LIIR) with patient cases improves primary care physicians' (PCPs') diagnostic skills and confidence. METHODS: A total of 115 PCPs were randomized (allocation ratio 3:1) to receive or not receive self-paced pattern recognition training in skin cancer diagnostics using an LIIR with patient cases through a quiz-based smartphone app during an 8-day period. The participants' ability to diagnose skin cancer was evaluated using a 12-item multiple-choice questionnaire prior to and 8 days after the educational intervention period. Their thoughts on the use of dermoscopy were assessed using a study-specific questionnaire. A learning curve was calculated through the analysis of data from the mobile app. RESULTS: On average, participants in the intervention group spent 2 hours 26 minutes quizzing digital patient cases and 41 minutes reading the educational material. They had an average preintervention multiple choice questionnaire score of 52.0% of correct answers, which increased to 66.4% on the postintervention test; a statistically significant improvement of 14.3 percentage points (P<.001; 95% CI 9.8-18.9) with intention-to-treat analysis. Analysis of participants who received the intervention as per protocol (500 patient cases in 8 days) showed an average increase of 16.7 percentage points (P<.001; 95% CI 11.3-22.0) from 53.9% to 70.5%. Their overall ability to correctly recognize malignant lesions in the LIIR patient cases improved over the intervention period by 6.6 percentage points from 67.1% (95% CI 65.2-69.3) to 73.7% (95% CI 72.5-75.0) and their ability to set the correct diagnosis improved by 10.5 percentage points from 42.5% (95% CI 40.2%-44.8%) to 53.0% (95% CI 51.3-54.9). The diagnostic confidence of participants in the intervention group increased on a scale from 1 to 4 by 32.9% from 1.6 to 2.1 (P<.001). Participants in the control group did not increase their postintervention score or their diagnostic confidence during the same period. CONCLUSIONS: Self-paced pattern recognition training in skin cancer diagnostics through the use of a digital LIIR with patient cases delivered by a quiz-based mobile app improves the diagnostic accuracy of PCPs. TRIAL REGISTRATION: ClinicalTrials.gov NCT05661370; https://classic.clinicaltrials.gov/ct2/show/NCT05661370.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.004 | 0.003 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.003 | 0.003 |
| Insufficient payload (model declined to judge) | 0.012 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".