Image Generation of Common Dermatological Diagnoses by Artificial Intelligence: Evaluation Study of the Potential for Education and Training Purposes
Bibliographic record
Abstract
Background: The integration of artificial intelligence (AI) into dermatology holds promise for education and diagnostic purposes, particularly through image generation, which has not been well studied. Objective: This study aimed to assess whether AI image generation software can generate accurate images of classic dermatological conditions and whether they are recognizable as computer-generated. Methods: Images of 10 dermatological conditions were generated using DALLE-2 and DALLE-3 programs. These images were randomized among clinical photographs and distributed to dermatology residents and attending physicians. Participants were instructed to (1) identify AI-generated images and (2) provide their diagnosis. Results: AI-generated images were detected as computer-generated in 70.8% (85/120) of cases. Correct diagnoses were made based on all AI images 40.83% (49/120) of the time. This was significantly lower than the 72.0% (46/60) recognition rate for clinical photographs (P<.001). DALLE-2 images were diagnosed correctly less frequently (25.0%, 15/60) than DALLE-3 images (56.6%, 34/60; P<.001). Conclusions: AI-generated images of common dermatological conditions are becoming more accurate. This holds great implications for education but should be used with caution as further research is needed with more advanced, specific, and inclusive training data. Limitations include the use of AI image generators created by a single parent company as well as the use of a limited set of diagnoses.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".