A Prototype Machine Learning Pipeline for Assessing and Tracking Keloid Scars
Bibliographic record
Abstract
Abstract Dysregulated wound healing, marked by excessive collagen deposition, is the hallmark to keloid scar formation. Current methods for assessing keloids in clinical settings rely heavily on subjective measures, which are prone to interrater variability. This study introduces a machine learning (ML) pipeline prototype, designed to automate the detection, measurement, and colour analysis of keloid scars. Using a convolutional neural network (CNN), the pipeline segments keloid lesions from 2D images, applies fiducial markers for accurate size measurement, and utilizes K-Means clustering for colorimetry analysis. The CNN achieved a classification accuracy of 98% on a small test dataset. Segmentation was further refined using binary masks and contour-based detection. Colorimetry analysis revealed heterogeneity in pigmentation across keloid lesions, was varied by patient skin type, and tracked changes over time. The pipeline was validated on patients over a 5–6-month period, accurately detecting changes in keloid size and colour. While the algorithm was highly effective in most cases, challenges were noted in patients with nascent keloid or those with dark skin tones where the contrast between keloid and skin was insufficient for accurate segmentation. Additionally, early-stage keloid detection showed inconsistencies in defining lesion boundaries, particularly when keloids expanded rapidly. Despite these limitations, the ML pipeline presents a promising tool for objective keloid assessment, offering a practical, accessible, and accurate alternative to current clinical practices.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.003 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".