Identifying individuals with undiagnosed post-traumatic stress disorder in a large United States civilian population – a machine learning approach
Bibliographic record
Abstract
BACKGROUND: The proportion of patients with post-traumatic stress disorder (PTSD) that remain undiagnosed may be substantial. Without an accurate diagnosis, these patients may lack PTSD-targeted treatments and experience adverse health outcomes. This study used a machine learning approach to identify and describe civilian patients likely to have undiagnosed PTSD in the US commercial population. METHODS: The IBM® MarketScan® Commercial Subset (10/01/2015-12/31/2018) was used. A random forest machine learning model was developed and trained to differentiate between patients with and without PTSD using non-trauma-based features. The model was applied to patients for whom PTSD status could not be confirmed to identify individuals likely and unlikely to have undiagnosed PTSD. Patient characteristics, symptoms and complications potentially related to PTSD, treatments received, healthcare costs, and healthcare resource utilization were described separately for patients with PTSD (Actual Positive PTSD cohort), patients likely to have PTSD (Likely PTSD cohort), and patients without PTSD (Without PTSD cohort). RESULTS: A total of 44,342 patients were classified in the Actual Positive PTSD cohort, 5683 in the Likely PTSD cohort, and 2,074,471 in the Without PTSD cohort. While several symptoms/comorbidities were similar between the Actual Positive and Likely PTSD cohorts, others, including depression and anxiety disorders, suicidal thoughts/actions, and substance use, were more common in the Likely PTSD cohort, suggesting that certain symptoms may be exacerbated among those without a formal diagnosis. Mean per-patient-per-6-month healthcare costs were similar between the Actual Positive and Likely PTSD cohorts ($11,156 and $11,723) and were higher than those of the Without PTSD cohort ($3616); however, cost drivers differed between cohorts, with the Likely PTSD cohort experiencing more inpatient admissions and less outpatient visits than the Actual Positive PTSD cohort. CONCLUSIONS: These findings suggest that the lack of a PTSD diagnosis and targeted management of PTSD may result in a greater burden among undiagnosed patients and highlights the need for increased awareness of PTSD in clinical practice and among the civilian population.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.003 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".