A case-controlled field study evaluating ICD-11 proposals for relational problems and intimate partner violence
Bibliographic record
Abstract
Background/Objective: Intimate partner relationship problems and intimate partner abuse and neglect — referred to in this paper as “relational problems and maltreatment” — have substantial and well-documented impact on both physical and mental health. However, classification guidelines, such as those found in the International Classification of Diseases (ICD-10), are vague and unlikely to support consistent application. Revised guidelines proposed for ICD-11 are much more operationalized. We used standardized clinical vignette conditions with an international panel of clinicians to test if ICD-11 changes resulted in improved classification accuracy. Method: English-speaking mental health professionals (N = 738) from 65 nations applied ICD-10 or ICD-11 (proposed) guidelines with experimentally manipulated case presentations of presence or absence of (a) individual mental health diagnoses and (b) relational problems or maltreatment. Results: ICD-11, compared with ICD-10, guidelines resulted in significantly better classification accuracy, although only in the presence of co-morbid mental health problems. Clinician factors (e. g., gender, language, world region) largely did not impact classification performance. Conclusions: Despite being considerably more explicated, raters’ performance with ICD-11 guidelines reveals training issues that should be addressed prior to the release of ICD-11 in 2018 (e. g., overriding the guidelines with pre-existing archetypes for relationship problems and physical and psychological abuse). Antecedentes/Objetivo: Los problemas en la relación de pareja y relacionados con abuso y negligencia de pareja, referidos como “problemas relacionales y maltrato”, tienen un importante impacto en la salud física y mental. Sin embargo, guías de clasificación, como la Clasificación Internacional de Enfermedades (CIE-10), son vagas y su aplicación es inconsistente. Las guías propuestas por el CIE-11 son más operacionales. Junto con un panel de clínicos, utilizamos viñetas clínicas estandarizadas, para evaluar si los cambios propuestos por CIE-11 mejoraban la precisión de la clasificación. Método: Profesionales de la salud de habla inglesa (N=738) de 65 naciones compararon la aplicación del CIE-10 y CIE-11 en casos experimentales, estableciendo presencia o ausencia de (a) diagnósticos individuales de salud mental y (b) problemas de relaciones o maltrato. Resultados: CIE-11 tuvo resultados significativamente más precisos, aunque solo en presencia de comorbilidades de salud mental. Factores como género, idioma y región no presentaron mayor alteración. Conclusiones: Aunque el CIE-11 está mejor explicado, este estudio revela problemas de capacitación que deberían abordarse antes de su publicación en 2018.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.011 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".