Systematic Review of Psychodermatologic Assessment Tools: Diagnostic Accuracy and Clinical Utility
Bibliographic record
Abstract
Primary psychodermatologic disorders such as body dysmorphic disorder, trichotillomania, and excoriation disorder present significant challenges in dermatological and psychiatric assessment due to their complex psychological and dermatological symptoms. Reliable and valid screening tools are essential for effective diagnosis and management, yet there is a lack of consensus on the most appropriate instruments. A systematic review was conducted, identifying 81 studies that employed 45 different psychodermatologic tools, of which 13 studies provided empirical data on their diagnostic accuracy. Tools were assessed for their psychometric properties, including sensitivity, specificity, reliability, and validity. The Body Dysmorphic Disorder Questionnaire (BDDQ) and its variants demonstrated high diagnostic accuracy, with the BDDQ showing a sensitivity of 0.97 [95% CI: 0.82-1.00] and specificity of 0.91 [95% CI: 0.86-0.95]. The Skin Picking Scale-Revised showed high diagnostic accuracy for excoriation disorder, with a sensitivity of 0.89 [95% CI: 0.84-0.94] and specificity of 0.95 [95% CI: 0.93-0.96]. Similarly, the Massachusetts General Hospital Hairpulling Scale, frequently utilized for trichotillomania, exhibited strong psychometric properties, with a sensitivity of 0.90 [95% CI: 0.81-0.96] and specificity of 0.72 [95% CI: 0.63-0.80]. Despite their frequent use, many tools lack a comprehensive assessment of the full range of symptoms, including social impairment and behavioural nuances. The review highlights the importance of developing standardized, multidimensional assessment tools that are valid, reliable, and easy to implement in daily practice. Further research is needed to establish the practical utility of these tools in routine dermatology settings, addressing gaps in effectiveness, referral and intervention limitations, and patient acceptability.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.034 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.010 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".