Beyond Diagnosis and Comorbidities—A Scoping Review of the Best Tools to Measure Complexity for Populations with Mental Illness
Bibliographic record
Abstract
Beyond the challenges of diagnosis, complexity measurement in clients with mental illness is an important but under-recognized area. Accurate and appropriate psychiatric diagnoses are essential, and further complexity measurements could contribute to improving patient understanding, referral, and service matching and coordination, outcome evaluation, and system-level care planning. Myriad conceptualizations, frameworks, and definitions of patient complexity exist, which are operationalized by a variety of complexity measuring tools. A limited number of these tools are developed for people with mental illness, and they differ in the extent to which they capture clinical, psychosocial, economic, and environmental domains. Guided by the PRISMA Extension for Scoping Reviews, this review evaluates the tools best suited for different mental health settings. The search found 5345 articles published until November 2023 and screened 14 qualified papers and corresponding tools. For each of these, detailed data on their use of psychiatric diagnostic categories, definition of complexity, primary aim and purpose, context of use and settings for their validation, best target populations, historical references, extent of biopsychosocial information inclusion, database and input technology required, and performance assessments were extracted, analyzed, and presented for comparisons. Two tools-the INTERMED, a clinician-scored and multiple healthcare data-sourced tool, and the VCAT, a computer-based instrument that utilizes healthcare databases to generate a comprehensive picture of complexity-are exemplary among the tools reviewed. Information on these limited but suitable tools related to their unique characteristics and utilities, and specialized recommendations for their use in mental health settings could contribute to improved patient care.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".