Challenges in Teledermoscopy Diagnostic Outcome Studies: Scoping Review of Heterogeneous Study Characteristics
Bibliographic record
Abstract
BACKGROUND: Teledermoscopy has demonstrated benefits such as decreased costs and enhanced access to dermatology care for skin cancer detection. However, the heterogeneity among teledermoscopy studies hinders the systematic reviews' synopsis of diagnostic outcomes, impeding trust and adoption in general practice and limiting overall health care benefits. OBJECTIVE: This study aims to improve understanding and standardization of teledermoscopy diagnostic studies, by identifying and categorizing study characteristics contributing to heterogeneity. Subsequently, the variability and consistency of these characteristics were assessed. METHODS: A review of systematic reviews regarding the diagnostic outcomes of teledermoscopy was performed to discern reported study characteristics contributing to heterogeneity. These characteristics were thematically grouped into 3 domains (population, index test, and reference standard), forming a data extraction framework. A scoping review on teledermoscopy diagnostic outcomes studies was performed, guided by the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) checklist. Data pertaining to study characteristics from included studies were extracted and analyzed through descriptive content analysis. Systematic reviews' reference lists validated the scoping review query. RESULTS: The literature search yielded 4 systematic reviews, revealing 15 heterogeneous studies across the population, index test, and reference standard domains. The scoping review identified 49 studies, with 27 overlapping with the systematic reviews. Population characteristics varied, with one-third (16/49, 33%) of studies reporting fewer than 100 samples; most studies (41/49, 84%) reported on the type of lesion, and most (20/49, 41%) teledermoscopy consultations took place in secondary care. One-fifth (11/49, 22%) did not describe inclusion or exclusion criteria, or the criteria varied highly. Index test characteristics showed differences in clinical expertise, profession, and training in dermatoscopic photography, and 59% (29/49) did not report on 1 or more index test characteristics. Image quality and clinical information reporting likewise varied. Reference standard characteristics involved teledermatologists' assessment, but 16 studies did not report teledermatologists' experience levels. Most studies (26/49, 53%) used histopathology as a gold standard. CONCLUSIONS: The heterogeneity in the population, index tests, and reference standard domains across teledermoscopy diagnostic outcome studies underscores the need for standardized reporting. This hinders the synopsis of teledermoscopy diagnostic outcomes in systematic reviews and limits the integration of research results into practice. Adopting a (tailored) STARD (Standards for Reporting Diagnostic Accuracy Studies) checklist for teledermoscopy diagnostic outcome studies is recommended to enhance the consistency and comparability of outcomes. We suggest performing a Delphi study to gather consensus on the tailored STARD guideline.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".