Clinical utility of ICD‐11 diagnostic guidelines for high‐burden mental disorders: results from mental health settings in 13 countries
Bibliographic record
Abstract
In this paper we report the clinical utility of the diagnostic guidelines for ICD-11 mental, behavioural and neurodevelopmental disorders as assessed by 339 clinicians in 1,806 patients in 28 mental health settings in 13 countries. Clinician raters applied the guidelines for schizophrenia and other primary psychotic disorders, mood disorders (depressive and bipolar disorders), anxiety and fear-related disorders, and disorders specifically associated with stress. Clinician ratings of the clinical utility of the proposed ICD-11 diagnostic guidelines were very positive overall. The guidelines were perceived as easy to use, corresponding accurately to patients' presentations (i.e., goodness of fit), clear and understandable, providing an appropriate level of detail, taking about the same or less time than clinicians' usual practice, and providing useful guidance about distinguishing disorder from normality and from other disorders. Clinicians evaluated the guidelines as less useful for treatment selection and assessing prognosis than for communicating with other health professionals, though the former ratings were still positive overall. Field studies that assess perceived clinical utility of the proposed ICD-11 diagnostic guidelines among their intended users have very important implications. Classification is the interface between health encounters and health information; if clinicians do not find that a new diagnostic system provides clinically useful information, they are unlikely to apply it consistently and faithfully. This would have a major impact on the validity of aggregated health encounter data used for health policy and decision making. Overall, the results of this study provide considerable reason to be optimistic about the perceived clinical utility of the ICD-11 among global clinicians.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".