A sequential modeling approach for predicting clinical outcomes with repeated measures
Bibliographic record
Abstract
The increased availability of healthcare data has made predictive modeling popular in a clinical setting. If an expected patient-specific outcome can be estimated prior to a medical intervention the healthcare costs can be reduced for both patient and provider. The nature of data used to train such predictive models is frequently longitudinal, as interventions with convalescence times or chronic conditions contain outcome measures at intermediate follow-up points. Here we outline a predictive modeling approach that takes advantage of the longitudinal structure of the data by sequentially predicting the outcomes at intermediate time points and including them as predictors in models for later time points. This is done for continuous and threshold-dichotomized outcomes. The proposed method improves predictive accuracy as it takes advantage of the correlation in follow-up measures to distribute the estimation of coefficient effects over several models, making it advantageous for smaller datasets. This formulation also allows for effective screening of first-order interaction effects. The improved performance is illustrated using a simulation study and an applied example of predicting outcomes following surgery. The proposed approach is shown to be consistent for prediction, effective in modeling interactions and robust to presence of noise variables.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.014 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.010 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".