Skin of Color Representation on Wikipedia: Cross-sectional Analysis (Preprint)
Bibliographic record
Abstract
BACKGROUND: Wikipedia is one of the most popular websites and may be a go-to source of health and dermatology education for the general population. Prior research indicates poor skin of color (SOC) photo representation in printed dermatology textbooks and online medical websites, but there has been no such assessment performed to determine whether this discrepancy also exists for Wikipedia. OBJECTIVE: The aim of this study was to investigate the number and quality of SOC photos included in Wikipedia's skin disease pages and to explore the possible ramifications of these findings. METHODS: Photos of skin diseases from Wikipedia's "List of Skin Conditions" were assigned by three independent raters as SOC or non-SOC according to the Fitzpatrick system, and were given a quality rating (1-3) based on sharpness, size/resolution, and lighting/exposure. RESULTS: We identified 421 skin disease Wikipedia pages and 949 images that met our inclusion criteria. Within these pages, 20.7% of images of skin diseases (196 of 949 images) were SOC and 79.3% (753 of 949 images) were non-SOC (P<.001). There was no difference in the average quality for SOC (2.05) and non-SOC (2.03) images (P=.81). However, the photo quality criteria utilized (sharpness, size/resolution, and lighting/exposure) did not capture all aspects of photo quality. Another limitation of this analysis is that the Fitzpatrick skin typing system is prone to subjectivity and was not originally intended to be utilized as a non-self SOC metric. CONCLUSIONS: There is SOC underrepresentation in the gross number of SOC images for dermatologic conditions on Wikipedia. Wikipedia pages should be updated to include more SOC photos to mend this divide to ameliorate access to accurate dermatology information for the general public and improve health equity within dermatology.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".