Growth mixture models: a case example of the longitudinal analysis of patient‐reported outcomes data captured by a clinical registry
Bibliographic record
Abstract
BACKGROUND: An assumption in many analyses of longitudinal patient-reported outcome (PRO) data is that there is a single population following a single health trajectory. One approach that may help researchers move beyond this traditional assumption, with its inherent limitations, is growth mixture modelling (GMM), which can identify and assess multiple unobserved trajectories of patients' health outcomes. We describe the process that was undertaken for a GMM analysis of longitudinal PRO data captured by a clinical registry for outpatients with atrial fibrillation (AF). METHODS: This expository paper describes the modelling approach and some methodological issues that require particular attention, including (a) determining the metric of time, (b) specifying the GMMs, and (c) including predictors of membership in the identified latent classes (groups or subtypes of patients with distinct trajectories). An example is provided of a longitudinal analysis of PRO data (patients' responses to the Atrial Fibrillation Effect on QualiTy-of-Life (AFEQT) Questionnaire) collected between 2008 and 2016 for a population-based cardiac registry and deterministically linked with administrative health data. RESULTS: In determining the metric of time, multiple processes were required to ensure that "time" accounted for both the frequency and timing of the measurement occurrences in light of the variability in both the number of measures taken and the intervals between those measures. In specifying the GMM, convergence issues, a common problem that results in unreliable model estimates, required constrained parameter exploration techniques. For the identification of predictors of the latent classes, the 3-step (stepwise) approach was selected such that the addition of predictor variables did not change class membership itself. CONCLUSIONS: GMM can be a valuable tool for classifying multiple unique PRO trajectories that have previously been unobserved in real-world applications; however, their use requires substantial transparency regarding the processes underlying model building as they can directly affect the results and therefore their interpretation.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.031 | 0.059 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.003 | 0.006 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.002 | 0.004 |
| Research integrity | 0.003 | 0.004 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".