Advancing Artificial Intelligence For Accurate, Equitable, And Interpretable Skin Cancer Diagnosis And Management
Bibliographic record
Abstract
Skin cancer is one of the most common cancers worldwide, and its incidence has been rising over the years. Early diagnosis substantially contributes to enhancing patient outcomes and increasing survival rates. However, due to a lack of dermatologists, especially in rural areas, cancer cases may go undiagnosed or inaccurately diagnosed. Subsequently, the burden of early diagnosis falls on non-specialists, such as primary healthcare providers, who are typically not trained to deal with complex dermatological conditions. Given the increasing prevalence of skin cancer and the chronic shortage of dermatological expertise, there is a critical need to develop computer-aided skin cancer decision support systems that offer an accurate early diagnosis. These applications are crucial to ensuring that patients receive timely treatment and that their chances of survival are significantly increased. The recent advances in artificial intelligence (AI) have given rise to a new era of skin cancer diagnosis models that perform on par with dermatologists. Nevertheless, the current AI diagnostic applications are subject to critical limitations. These include the lack of racial data diversity that results in the development of inequitable diagnostic models. Additionally, the black-box nature of AI models poses interpretability challenges that diminish human understandability and trust thus limiting their application in a clinical workflow. Furthermore, the paucity of applications dedicated to disease management prediction, primarily caused by the dearth of labeled data for the purpose of managing skin cancers, presents a significant hurdle in advancing AI in treatment prediction. This thesis aims to harness the power of AI to overcome these limitations, thereby achieving equitable, interpretable skin cancer diagnosis, and enhanced disease management. To accomplish these objectives, this work comprised five phases. In Phase 1, a comprehensive and analytical review employing text mining techniques was conducted to study AI methods and applications in skin cancer diagnosis and treatment. This analysis sought to gain a deep understanding of the explored capabilities and challenges of AI within these fields. Phases 2 and 3 were dedicated to resolving the data diversity issue. Phase 2 focused on the development of an integrated tool that encompassed segmentation, pixel clustering and classification to quantitively assess representation disparities of dark skin tones in dermatological resources. Phase 3 was centred around augmenting the training data with the underrepresented skin tones and developing an inclusive malignancy detection model employing deep neural networks. Phase 4 focused on developing interpretable diagnosis models that capitalize on the incorporation of human knowledge into model design and training to create transparent diagnosis models. Finally, Phase 5 delved into disease management, where a comparison between human-centred and machine-centred approaches was conducted. The two approaches aimed to accurately predict skin cancer management options while overcoming the challenges posed by data size limitations.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.019 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.006 | 0.004 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.002 | 0.004 |
| Insufficient payload (model declined to judge) | 0.004 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".