Systems Science, Data Science, and Machine Learning to Model the Dynamic Suicide Process
Bibliographic record
Abstract
Suicide and related behaviors, such as ideation, planning, and attempts, pose significant public health challenges globally. Accurately predicting these behaviors, which involve both observable (e.g., lethal attempts) and latent (e.g., suicidal ideation) aspects, is crucial for effective interventions. However, traditional models struggle to capture the complex dynamics of suicide-related behaviors, especially latent states. This dissertation begins with a systematic scoping review of existing Systems Science models for suicide-related behaviors, identifying gaps in latent state modeling and the use of stochastic methods. These gaps inform the development of three distinct modeling approaches: regular System Dynamics (SD) modeling, SD enhanced with particle filtering, and SD incorporating Particle Markov Chain Monte Carlo (PMCMC). Using time-series data on suicide-related deaths stratified by sex and method, the study develops an aggregated SD model followed by sex-stratified and sex-method-stratified models. To address the limitations of deterministic models, particle filtering and PMCMC are applied, introducing stochastic elements that improve the estimation of system states and underlying parameters, such as transition rates between suicidal behavior stages. PMCMC, in particular, refines parameter estimates, enhancing prediction accuracy. The performance of these models is rigorously evaluated using metrics like root mean square error (RMSE), acceptance ratio of MCMC iterations, and the plot of observed and estimated data, wherever it was available. Results show that stochastic models, especially those incorporating PMCMC, outperform regular SD models in estimating both observed and latent behaviors. This research advances Systems Science methodologies in public health by demonstrating the value of stochastic methods in dynamic models.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.009 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.004 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".