Idiographic Lapse Prediction With State Space Modeling: Algorithm Development and Validation Study
Bibliographic record
Abstract
BACKGROUND: Many mental health conditions (eg, substance use or panic disorders) involve long-term patient assessment and treatment. Growing evidence suggests that the progression and presentation of these conditions may be highly individualized. Digital sensing and predictive modeling can augment scarce clinician resources to expand and personalize patient care. We discuss techniques to process patient data into risk predictions, for instance, the lapse risk for a patient with alcohol use disorder (AUD). Of particular interest are idiographic approaches that fit personalized models to each patient. OBJECTIVE: This study bridges 2 active research areas in mental health: risk prediction and time-series idiographic modeling. Existing work in risk prediction has focused on machine learning (ML) classifier approaches, typically trained at the population level. In contrast, psychological explanatory modeling has relied on idiographic time-series techniques. We propose state space modeling, an idiographic time-series modeling framework, as an alternative to ML classifiers for patient risk prediction. METHODS: We used a 3-month observational study of participants (N=148) in early recovery from AUD. Using once-daily ecological momentary assessment (EMA), we trained idiographic state space models (SSMs) and compared their predictive performance to logistic regression and gradient-boosted ML classifiers. Performance was evaluated using the area under the receiver operating characteristic curve (AUROC) for 3 prediction tasks: same-day lapse, lapse within 3 days, and lapse within 7 days. To mimic real-world use, we evaluated changes in AUROC when models were given access to increasing amounts of a participant's EMA data (15, 30, 45, 60, and 75 days). We used Bayesian hierarchical modeling to compare SSMs to the benchmark ML techniques, specifically analyzing posterior estimates of mean model AUROC. RESULTS: Posterior estimates strongly suggested that SSMs had the best mean AUROC performance in all 3 prediction tasks with ≥30 days of participant EMA data. With 15 days of data, results varied by task. Median posterior probabilities that SSMs had the best performance with ≥30 days of participant data for same-day lapse, lapse within 3 days, and lapse within 7 days were 0.997 (IQR 0.877-0.999), 0.999 (IQR 0.992-0.999), and 0.998 (IQR 0.955-0.999), respectively. With 15 days of data, these median posterior probabilities were 0.732, <0.001, and <0.001, respectively. CONCLUSIONS: The study findings suggest that SSMs may be a compelling alternative to traditional ML approaches for risk prediction. SSMs support idiographic model fitting, even for rare outcomes, and can offer better predictive performance than existing ML approaches. Further, SSMs estimate a model for a patient's time-series behavior, making them ideal for stepping beyond risk prediction to frameworks for optimal treatment selection (eg, administered using a digital therapeutic platform). Although AUD was used as a case study, this SSM framework can be readily applied to risk prediction tasks for other mental health conditions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.019 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".