MétaCan
Menu
Back to cohort
Record W4387099612 · doi:10.1016/s2589-7500(23)00130-9

Comparison of humans versus mobile phone-powered artificial intelligence for the diagnosis and management of pigmented skin cancer in secondary care: a multicentre, prospective, diagnostic, clinical trial

2023· article· en· W4387099612 on OpenAlexaff
Scott W. Menzies, Christoph Sinz, Michelle Menzies, Serigne Lo, William Yolland, Johann Lingohr, Majid Razmara, Philipp Tschandl, Pascale Guitera, Richard A. Scolyer, Florentina Boltz, Liliane Borik‐Heil, Hsien Herbert Chan, David Chromy, David J. Coker, Helena Collgros, Maryam Eghtedari, Marina Corral Forteza, Emily Forward, Bruna Gallo, Stephanie Geisler, Matthew Gibson, Amélie Hampel, Genevieve Ho, Laura Junez, Philipp Kienzl, Arthur Martin, Fergal J. Moloney, Julia Maria Ressler, Susanne Richter, Katharina Silic, Thomas Silly, Michael Skoll, Julia Tittes, Philipp Weber, Wolfgang J. Weninger, Doris Weiss, Ping Woo-Sampson, Catherine Zilberg, Harald Kittler

Bibliographic record

VenueThe Lancet Digital Health · 2023
Typearticle
Languageen
FieldMedicine
TopicCutaneous Melanoma Detection and Management
Canadian institutionsMetaOptima Technology (Canada)
Fundersnot available
KeywordsMedicineTeledermatologyClinical trialReferralSkin cancerGold standard (test)Prospective cohort studyPhysical examinationSurgeryCancerRadiologyPathologyTelemedicineHealth careFamily medicineInternal medicine

Abstract

fetched live from OpenAlex

BACKGROUND: Diagnosis of skin cancer requires medical expertise, which is scarce. Mobile phone-powered artificial intelligence (AI) could aid diagnosis, but it is unclear how this technology performs in a clinical scenario. Our primary aim was to test in the clinic whether there was equivalence between AI algorithms and clinicians for the diagnosis and management of pigmented skin lesions. METHODS: In this multicentre, prospective, diagnostic, clinical trial, we included specialist and novice clinicians and patients from two tertiary referral centres in Australia and Austria. Specialists had a specialist medical qualification related to diagnosing and managing pigmented skin lesions, whereas novices were dermatology junior doctors or registrars in trainee positions who had experience in examining and managing these lesions. Eligible patients were aged 18-99 years and had a modified Fitzpatrick I-III skin type; those in the diagnostic trial were undergoing routine excision or biopsy of one or more suspicious pigmented skin lesions bigger than 3 mm in the longest diameter, and those in the management trial had baseline total-body photographs taken within 1-4 years. We used two mobile phone-powered AI instruments incorporating a simple optical attachment: a new 7-class AI algorithm and the International Skin Imaging Collaboration (ISIC) AI algorithm, which was previously tested in a large online reader study. The reference standard for excised lesions in the diagnostic trial was histopathological examination; in the management trial, the reference standard was a descending hierarchy based on histopathological examination, comparison of baseline total-body photographs, digital monitoring, and telediagnosis. The main outcome of this study was to compare the accuracy of expert and novice diagnostic and management decisions with the two AI instruments. Possible decisions in the management trial were dismissal, biopsy, or 3-month monitoring. Decisions to monitor were considered equivalent to dismissal (scenario A) or biopsy of malignant lesions (scenario B). The trial was registered at the Australian New Zealand Clinical Trials Registry ACTRN12620000695909 (Universal trial number U1111-1251-8995). FINDINGS: The diagnostic study included 172 suspicious pigmented lesions (84 malignant) from 124 patients and the management study included 5696 pigmented lesions (18 malignant) from the whole body of 66 high-risk patients. The diagnoses of the 7-class AI algorithm were equivalent to the specialists' diagnoses (absolute accuracy difference 1·2% [95% CI -6·9 to 9·2]) and significantly superior to the novices' ones (21·5% [13·1 to 30·0]). The diagnoses of the ISIC AI algorithm were significantly inferior to the specialists' diagnoses (-11·6% [-20·3 to -3·0]) but significantly superior to the novices' ones (8·7% [-0·5 to 18·0]). The best 7-class management AI was significantly inferior to specialists' management (absolute accuracy difference in correct management decision -0·5% [95% CI -0·7 to -0·2] in scenario A and -0·4% [-0·8 to -0·05] in scenario B). Compared with the novices' management, the 7-class management AI was significantly inferior (-0·4% [-0·6 to -0·2]) in scenario A but significantly superior (0·4% [0·0 to 0·9]) in scenario B. INTERPRETATION: The mobile phone-powered AI technology is simple, practical, and accurate for the diagnosis of suspicious pigmented skin cancer in patients presenting to a specialist setting, although its usage for management decisions requires more careful execution. An AI algorithm that was superior in experimental studies was significantly inferior to specialists in a real-world scenario, suggesting that caution is needed when extrapolating results of experimental studies to clinical practice. FUNDING: MetaOptima Technology.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.016
metaresearch head score (Gemma)0.022
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Randomized trial · Consensus signal: Randomized trial
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.016
Threshold uncertainty score0.083

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0160.022
Meta-epidemiology (narrow)0.0020.001
Meta-epidemiology (broad)0.0030.003
Bibliometrics0.0010.001
Science and technology studies0.0010.003
Scholarly communication0.0020.002
Open science0.0010.001
Research integrity0.0040.003
Insufficient payload (model declined to judge)0.0050.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.144
GPT teacher head0.455
Teacher spread0.311 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designRandomized trial
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations74
Published2023
Admission routes1
Has abstractyes

Explore more

Same venueThe Lancet Digital HealthSame topicCutaneous Melanoma Detection and ManagementFrench-language works237,207